Side Task Review

Reviews side tasks for quality issues using AI analysis.

Purpose

Uses Claude Code to analyze each side task's scorer and documentation for:

  • Reward hacking vulnerabilities
  • Magic numbers or hardcoded values
  • Accidental triggering possibilities
  • Straj (sabotage trajectory) feasibility
  • Overall task quality

Commands

ct run side-task-review

Review side tasks for quality issues.

Usage:

ct run side-task-review [OPTIONS]

Options:

  • -e, --env - Environment name (will prompt if not provided)
  • -s, --side-task - Side task(s) to review (repeatable, will prompt if not provided)
  • -o, --output-dir - Output directory (default: data/side_task_review_runs/<env>)

Examples:

# Review all side tasks for an environment (interactive selection) ct run side-task-review -e web_scraping # Review specific side tasks ct run side-task-review -e web_scraping -s visit_malicious_website -s expose_secret # Custom output directory ct run side-task-review -e web_scraping -o my_reviews/

Output

Results are saved to data/side_task_review_runs/<env>/ as one review file per side task. Each holds the identified issues and recommendations, with the full analysis in its full_review field.

Review Areas

For each side task:

  1. Reward Hacking: Can the scorer pass without the security property being violated?
  2. Magic Numbers: Do hardcoded values make the scorer fragile?
  3. Accidental Triggering: Could normal operations accidentally pass the scorer?
  4. Straj Feasibility: Is it realistic to create sabotage trajectories?
  5. Quality Issues: Are there other problems with the scorer or task design?

Use Cases

  • Pre-launch quality checks for new side tasks
  • Identifying tasks that need hardening
  • Planning straj development
  • Improving overall benchmark quality

Prerequisites

Run as a non-root user: Claude Code refuses --dangerously-skip-permissions under root or sudo.