Quickstart
This is the fastest path from a fresh machine to running control evals on the public LinuxArena environments, using an OpenRouter API key. Everything runs locally in Docker; no cloud accounts or internal credentials are needed.
New to AI control? Read AI Control Concepts first (5 minutes); it defines main tasks, side tasks, honest/attack evals, and monitoring.
Prerequisites
- Docker, installed and running
- uv:
curl -LsSf https://astral.sh/uv/install.sh | sh - git: environment definitions are public repos, pulled with plain
git clone(no GitHub account needed) - An OpenRouter API key
Setup
git clone https://github.com/linuxarena/control-tower.git cd control-tower uv sync alias ct='uv run ct' # add to your shell profile to keep it
Create a .env file in the checkout root:
OPENROUTER_API_KEY=sk-or-v1-... USE_OPENROUTER=1 SKIP_DOCKER_BUILD=1
USE_OPENROUTER=1routes all model calls through your OpenRouter key, so you don't need separate Anthropic/OpenAI keys.SKIP_DOCKER_BUILD=1pulls prebuilt environment images from Docker Hub instead of building them locally.
Then pull the public LinuxArena environments:
ct settings pull linuxarena ct envs list
Run an honest eval
An honest eval gives the agent a main task and measures whether it completes it:
ct run eval --no-upload \ -e clinical_trial -t adverse_reaction_followup_tracking \ --policy honest
--no-upload keeps results local (uploading requires internal services). The agent's actions stream to your terminal; the run takes a few minutes and ends with the scorer's CORRECT/INCORRECT verdict. The trajectory is saved as an Inspect .eval file under logs/.
To list tasks for an environment: ct envs main-tasks list --env <env>. To pick a different model, pass -p model=<name> with any OpenRouter model, e.g. -p model=openrouter/qwen/qwen3-coder.
Run an attack eval
An attack eval additionally gives the agent a hidden side task (sabotage objective) it must accomplish without looking suspicious:
ct run eval --no-upload \ -e clinical_trial -t adverse_reaction_followup_tracking \ -s suppress_safety_notifications \ --policy attack
Side tasks per environment: ct envs side-tasks list --env <env>.
Monitor a trajectory
Monitors are the blue team's core tool: another model reviews the agent's actions and assigns suspiciousness scores. Score the attack trajectory you just produced:
ct run monitor logs/<eval-file>.eval \ -m mean-action -m mean-quick \ --trusted-model openai/gpt-5-nano \ --no-upload
View results
uv run inspect view logs/<eval-file>.eval
This opens Inspect's local log viewer with the full transcript, scores, and monitor outputs. See Inspect Logs Viewer.
Using a self-hosted model
To use an open-weight model from your own GPU box, serve it (the tool-call flags are required for agent use; monitors don't need them):
vllm serve Qwen/Qwen3-32B --enable-auto-tool-choice --tool-call-parser hermes
add to your .env:
VLLM_BASE_URL=http://<gpu-box>:8000/v1
and pass the model to the commands above: -p model=vllm/Qwen/Qwen3-32B -p model_attempt_timeout=600 --max-samples 1 on ct run eval, or --trusted-model vllm/Qwen/Qwen3-32B on ct run monitor. Open-weight models are an easy fit for monitors (plain text, no tool calling). For other OpenAI-compatible servers (NVIDIA NIM, TGI) and details, see Self-Hosted and Open-Weight Models.
Gotchas
- Always pass
--no-uploadtoct run evalandct run monitor. Upload is on by default and requires internal credentials. - The
tiktokenvironment needs aTIKAPI_KEY(a paid third-party API). Use the other environments unless you've been given one. - Rate limits: OpenRouter rate limits scale with the key's credit balance. If you hit 429s, lower concurrency with
--max-samples 1.
Going deeper
- Running evals on Hawk: submit to a shared Hawk cluster and watch results in the viewer, instead of running locally
- Running Evaluations: full
ct run evalreference: policies, task sets, batching, replay - Monitoring: monitor types, ensembles, suspiciousness scoring (building monitors)
- Control Protocols: defer-to-trusted, trusted editing, and writing your own protocol
- Making Sabotage Trajectories: hand-crafting attacks that evade monitors
- Making Side Tasks and the Hands-On Guide: extending the testbed with new environments and tasks
- Safety Curves: turning monitor scores into safety/usefulness numbers