Compare Side Tasks
Compares side task outcomes between two evaluation runs: which side tasks succeeded in one run but failed in the other, and how success rates changed. Use it to compare policies (e.g. honest vs attack) or models, and for regression testing after changes.
Commands
ct run compare-side-tasks
Compare side task successes between two runs.
Usage:
ct run compare-side-tasks RUN_ID_1 RUN_ID_2
Arguments:
RUN_ID_1- First run IDRUN_ID_2- Second run ID
Examples:
# Compare two runs ct run compare-side-tasks abc123 def456
Output
Displays:
- Run metadata (name, policy, tags, date) for identification
- Side tasks that succeeded in run 1 but failed in run 2
- Side tasks that failed in run 1 but succeeded in run 2
- Side tasks only present in one run (when runs cover different task sets)
- Side tasks with the same result in both runs
- Overall success rate comparison
- Environment-level breakdown
Each run's eval log is downloaded through the viewer API and cached under data/eval-logs/, and results are grouped by environment and side task.