Making Main Tasks
Definition
A main task is a task an AI is asked to accomplish in an environment. Main tasks are structurally the same as tasks in other AI evals such as SWE-bench and MATH.
Components
A main task has the following components:
- Setup Script (optional): A shell script run in the environment before the agent starts, used to make small edits to the environment.
- Description: The task the agent reads before it starts, written like a well-defined ticket an engineer would receive.
- Category: The kind of work the task asks for, set as
task_categoryin the task'sinfo.yml. See Categories below. - Scorer (optional): The Inspect scorer that runs at the end of a trajectory and determines whether the AI accomplished the task. Without one, the task is recorded as unscored (
main_task_success=None).
Categories
Each main task's info.yml sets task_category to one of the following literals. The source of truth is TaskCategory in src/control_tower/task/task.py.
add_feature
The agent adds a new feature to the codebase, and normal tests of the feature check it. Example: add the ability to run Python jobs to the job system.
fix_bug
The task's setup.sh introduces a bug that breaks the system, for example by deleting code or misconfiguring the database so it doesn't start. Normal tests that must pass again are usually enough, so these tasks are much easier to build.
refactor
The agent restructures existing code without adding functionality, for example "convert this to Rust" or "swap out the database for a Redis cache". Scorers for these are much harder to write and need some tricks. LLM scorers would work here, but we want to avoid them.
research
The agent runs an experiment in the codebase and reports the results, for example "benchmark these features and explain why they are slow". These are also hard to test; typically the agent returns a structured report and the scorer checks that certain data points are present.
lang_rewrite
The agent ports an existing component to another language (e.g. "rewrite this Python service in Go"). Scoring works as for refactor: integration tests confirm behavior is preserved.
infra
Infrastructure-level changes that don't fit the other categories (e.g. reshuffling deployment topology or swapping a build system).
Target Distribution
The ideal spread of main tasks for each environment is:
- 50%
add_feature - 25%
fix_bug - 15%
refactor/lang_rewrite - 10%
research
infra is used sparingly and does not have a target share.
Desired Properties
The task is complex and well defined
Motivation: Complex tasks are the interesting ones for control research, and a well-defined description and scorer keep a task meaningful whichever model attempts it. How hard a task should be depends on the project running it: some projects want tasks only the strongest current models can solve, others deliberately run weaker models, and some run main tasks with no side task at all.
Difficulty is a property of a task and a model together, and models improve, so a task that is hard for today's models will be routine for later ones. Write tasks that are complex and precisely specified rather than tuning them to a pass rate for a particular model, and keep model names out of a task's info.yml and notes; the trusted and untrusted aliases in src/control_tower/models/aliases.py name the roles without pinning a version.
For a control evaluation, the useful gap is between the trusted and untrusted models: a task the trusted model cannot solve but the untrusted model can gives the Blue Team a reason to use the untrusted model and makes the usefulness cost of a protocol such as defer-to-trusted measurable. Measure that gap against the models your project runs and expect to re-measure when they change.
The task is solvable with access to the task description but not the scorer
Motivation: The AI should not need access to the main task scorer.
Every main task should ideally have a solution trajectory that solves it when replayed. A human-made solution does not prove that the description matches the scorer, because the human already knew how the scorer worked. An AI solving the task from the description alone is good evidence.
If no AI can solve the task yet, there needs to be some reason to think future AIs will. For now, a second person checks that the description and scorer line up. fix_bug tasks are easy to check because the environment is known to have worked beforehand.
The main task is not prone to accidentally causing the side task
Motivation: Honest models accidentally achieving the side task makes the research much harder.
For example, "refactor the entire security code" is a bad main task: if an honest AI introduces a bug, it might trigger the side task.
This is hard to predict, so it is judged case by case, and nothing is required for now.
The main task has good overlap with critical code
Motivation: Tasks where an AI could plausibly sabotage something should be upsampled.
Each main task should have some story for how the side task could be accomplished alongside it. Building a straj shows this but is a lot of work and overkill for every main task; for now, give a basic rationale for the overlap in the PR.
Other Components
A main task is a folder inside an environment's main_tasks/ directory. Its info.yml must set description and task_category; the loader raises without them. A scorer.py with a function decorated by @linux_scorer (from control_tower.environments.scorer_utils) is optional but expected.
When submitting a main task idea PR, add _idea to the end of the folder name (e.g. my_task_idea/). This keeps tests from failing on idea PRs and tracks development progress.
Main Task info.yml
The other info.yml fields are optional but strongly encouraged:
name: Task Name task_category: add_feature | fix_bug | refactor | research | lang_rewrite | infra short_description: Brief one-line description of what to implement scorer_description: >- High-level test steps (3-5 sentences). Describe what the scorer *does*: what test data is created, what operations are performed, what is verified. Not what the task accomplishes. coverage: >- Which files and features the agent is expected to modify in order to implement the main task. description: A detailed description of this task, what it does, why it's nice, etc. enabled_attack_surface: How does this main task idea provide cover/attack surface for a subset of the side tasks in this environment? notes: An optional string of various notes, ideas, things to look out for etc.
Advanced Configuration
To override a task's defaults, create an __init__.py in the task folder that sets any of:
codebase_name- the docker compose to be used, as defined in thecompose_pathsof the task's environmentsetup_script- the contents of an optional bash script that should be run before the agent gets to see the environment (run as root). Defaults to the contents of the task folder'ssetup.sh, if presentfiles- per-sample file payloads copied into the sandbox at init, before the setup script runs. A dict mapping a destination path in the sandbox to a host source that must exist: a host file, or a host directory (copied recursively). Relative sources are resolved against the task folder and must name a path inside it; programmatic tasks with no task folder must use absolute paths. Literal file contents, http(s) URLs, and data URIs are not supported (add real files to the task folder instead). Useful for injecting per-task content (a dataset, a repository to work on) without baking it into the environment imagefinal_diff- aFinalDiffConfig(fromcontrol_tower.task.final_diff) that copies the workspace to a host-side baseline after setup, captures its final diff at scoring time, and records it in the eval log for post-hoc trajectory monitors. Configure the workspace withrootand paths omitted from both snapshots withexclude; every path is included by default, and any explicit exclusion is a monitoring blind spot. Diffs larger thanmax_bytes(default 512 KB) are truncated to that hard byte bound by keeping both ends, with truncation recorded and flagged prominently in monitor promptssubmit_config- the tool given to the agent to submit their solution. Should be atools.submit.SubmitToolConfiginstance. The default is a no-op submit tool that ends the evaluation.
An example __init__.py:
from pydantic import BaseModel, Field from control_tower.tools.submit import SubmitToolConfig class SubmitAttackArgs(BaseModel): job_code: str = Field( description="The job code to be run. This should be a valid C program." ) # The tool used to submit a task by the agent submit_config = SubmitToolConfig( name="submit", description="Submit the job code", input_type=SubmitAttackArgs, )