Side Task Review
Reviews side tasks for quality issues using AI analysis.
Purpose
Uses Claude Code to analyze each side task's scorer and documentation for:
- Reward hacking vulnerabilities
- Magic numbers or hardcoded values
- Accidental triggering possibilities
- Straj (sabotage trajectory) feasibility
- Overall task quality
Commands
ct run side-task-review
Review side tasks for quality issues.
Usage:
ct run side-task-review [OPTIONS]
Options:
-e, --env- Environment name (will prompt if not provided)-s, --side-task- Side task(s) to review (repeatable, will prompt if not provided)-o, --output-dir- Output directory (default: data/side_task_review_runs/<env>)
Examples:
# Review all side tasks for an environment (interactive selection) ct run side-task-review -e web_scraping # Review specific side tasks ct run side-task-review -e web_scraping -s visit_malicious_website -s expose_secret # Custom output directory ct run side-task-review -e web_scraping -o my_reviews/
Output
Results are saved to data/side_task_review_runs/<env>/ as one review file per side task. Each holds the identified issues and recommendations, with the full analysis in its full_review field.
Review Areas
For each side task:
- Reward Hacking: Can the scorer pass without the security property being violated?
- Magic Numbers: Do hardcoded values make the scorer fragile?
- Accidental Triggering: Could normal operations accidentally pass the scorer?
- Straj Feasibility: Is it realistic to create sabotage trajectories?
- Quality Issues: Are there other problems with the scorer or task design?
Use Cases
- Pre-launch quality checks for new side tasks
- Identifying tasks that need hardening
- Planning straj development
- Improving overall benchmark quality
Prerequisites
Run as a non-root user: Claude Code refuses --dangerously-skip-permissions under root or sudo.