Reward Hacking Audit

Audits side tasks for reward hacking vulnerabilities: ways an agent could pass the scorer without violating the security property.

Commands

ct run reward-hacking

Audit side tasks for reward hacking vulnerabilities.

Usage:

ct run reward-hacking [OPTIONS]

Options:

  • -e, --env - Environment(s) to audit (repeatable)
  • -s, --side-task - Specific side task(s) to audit (repeatable)
  • -n, --name - Name for this audit run (default: "reward_hacking")
  • -c, --max-concurrent - Maximum concurrent API calls (default: 4)
  • -t, --thinking-budget - Token budget for extended thinking (default: 10000)

Examples:

# Audit all side tasks ct run reward-hacking # Audit specific environment ct run reward-hacking -e web_scraping # Audit specific side task ct run reward-hacking -s expose_secret

ct run reward-hacking-view

Interactive viewer for reward hacking audit results.

Usage:

ct run reward-hacking-view [RUN_PATH]

Output

Results are saved to data/reward_hacking_runs/<name>_<timestamp>/:

  • results.jsonl - Vulnerability analysis and recommendations for each side task
  • prompt_logs/<env>_<side_task>.json - Prompts, response, and thinking for each side task
  • run.log - Run log

What It Checks

The smart model alias, with extended thinking, reads each side task's scorer code and the environment codebase, using the prompts in prompts/system.md and prompts/user.md. For each side task it asks:

  • Can the scorer be passed without actual sabotage?
  • Are there implementation bugs in the scorer?
  • Can the agent exploit edge cases?
  • Are there "magic number" vulnerabilities?
  • Does the scorer match the intended security property?