EC2 Fleet (Distributed Evaluation)

--ec2 runs an evaluation on a fleet of EC2 workers instead of local Docker. Jobs are distributed over SQS; each worker runs Docker locally through the standard lb-docker sandbox. ct run mtgen uses the fleet by default (pass --no-ec2 to run locally).

Authentication

Use AWS SSO where you can, since it has no long-lived secrets; static IAM keys also work.

AWS SSO. You need an Identity Center assignment to the fleet account (326120904690) with the LinuxbenchFleet permission set; ask a fleet account admin in #devops. This permission set is enough to run evals. If you already have a profile for that account and permission set, reuse it. Otherwise run aws configure sso, using the organization's start URL from #devops and a profile name of your choice.

List your profiles with aws configure list-profiles, then replace your-profile below with the one you configured:

export AWS_PROFILE="your-profile" aws sso login aws sts get-caller-identity

If your default AWS profile already selects the fleet account and permission set, omit AWS_PROFILE. Re-run aws sso login when the selected profile's session expires. In-flight runs refresh credentials only while the SSO session is valid, so log in fresh before a long run.

Control Tower also reads AWS_PROFILE from .env if the shell does not set it; this only selects an existing local profile and does not configure SSO. The AWS CLI and the ct-box scripts do not read .env, so exporting the variable in your shell is the way to select the same profile for all three.

Remove any AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY lines from .env when using AWS_PROFILE. Environment-variable keys take precedence over the profile, so fleet commands refuse to start when both are set.

Static IAM keys. Request a scoped fleet IAM key pair from a fleet account admin (see Granting Fleet Access), then put it in .env:

AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=...

Fleet Mode

How It Works

  1. Your machine enqueues each sample as a message in SQS.
  2. It launches N EC2 workers from the pre-baked AMI, or stock Ubuntu if there is none.
  3. Each worker takes samples from the queue, runs them locally with Docker, and uploads the .eval results to S3.
  4. Workers post completion status to a done queue, which your machine uses to track progress.
  5. Your machine collects all results and merges them into one run.
  6. Workers run shutdown -h now when the queue is empty, which terminates the instances.

Prerequisites

One-time setup:

  1. S3 bucket: lb-fleet-results-migrated stores .eval files, manifests, and code artifacts.
  2. IAM role and instance profile: linuxbench-worker, with the inline policy in ops/docs/fleet-worker-iam-policy.json.
  3. API keys: set ANTHROPIC_API_KEY and/or OPENAI_API_KEY (and GH_TOKEN) in .env. The controller passes them to workers as per-instance SSM SecureStrings that only that instance can read, and deletes them at teardown.
  4. Docker registry auth for private env images: either
    • set DOCKER_ORG_NAME and DOCKER_ORG_PAT in .env, which reach workers like the API keys, or
    • store them in SSM under /linuxbench/docker-org-name and /linuxbench/docker-org-pat (overridable with DOCKER_ORG_NAME_SSM_PARAM / DOCKER_ORG_PAT_SSM_PARAM, staying under /linuxbench/docker-org-). The helper script stores the PAT as SecureString.

To store the Docker auth from your current .env into SSM:

./scripts/put_docker_org_auth_ssm.sh

Usage

# Fleet mode on EC2 ct run eval --ec2 --task-file medium --policy honest --epochs 10 # Scale up with more workers for faster runs ct run eval --ec2 --task-file medium --policy honest --epochs 10 --ec2-arg max_workers=20 # Cost estimate only (no launch) ct run eval --ec2 --task-file medium --policy honest --epochs 10 --ec2-arg estimate=true # Skip confirmation prompt (CI-friendly) ct run eval --ec2 --task-file medium --policy honest --epochs 10 --ec2-arg auto_confirm=true # Force a fresh fleet (skip auto-resume) ct run eval --ec2 --task-file medium --policy honest --epochs 10 --ec2-arg new_fleet=true

Every fleet knob is an --ec2-arg key=value (repeatable); the keys are the fields of Ec2Config in src/control_tower/sandbox/cli.py, listed with their defaults in --help. A sabotage-eval config sets the same keys in an ec2_args: mapping at run level or per attack; the attack's keys win.

Selecting Which Settings Workers Pull

Workers run ct settings pull <target> during bootstrap instead of using your local .settings/. By default they pull linuxarena. For an env from a different setting (e.g. basharena), prefix the -e value with the setting:

# Single explicit setting ct run eval --ec2 -e basharena:basharena -s ... # Multiple settings — repeat -e per env, deduplicates settings automatically ct run eval --ec2 \ -e basharena:basharena \ -e linuxarena:load_balancer \ -s ... # Bare form keeps working; falls back to the linuxarena default ct run eval --ec2 -e load_balancer ...

Only the fleet controller parses the setting:env form; local runs ignore the prefix. Either prefix every -e or leave them all bare, since mixing the two forms is rejected.

Shipping external code to workers (--extra-src)

Workers get only a git archive of the control-tower checkout, so an external --policy mypkg.foo:bar raises ModuleNotFoundError on the worker. --extra-src DIR also ships DIR as an import root, like PYTHONPATH=DIR:

ct run eval --ec2 --extra-src ~/experiments --policy mypkg.policies:my_policy -e load_balancer -s ...

The flag is repeatable, and each directory needs a distinct basename. It applies only to --ec2, since local runs resolve external refs from the working directory. Build artifacts, caches, VCS metadata (__pycache__, .git, node_modules, .venv/venv, build, dist), and lockfiles are not shipped, so source that lives in a directory named build or dist will be missing on the worker.

External sandboxes on the fleet (--sandbox with --ec2)

--sandbox module:type runs on the fleet too: each worker imports the module from the shipped extra-src trees and runs the sandbox for its own jobs. If the sandbox needs host-side installation (a microVM runtime, a CLI login), name a script inside an extra-src tree with --ec2-arg worker_setup=<basename>/<path>. The worker runs it as root after extracting extra-src and before taking jobs, with the venv at /opt/linuxbench/.venv and the worker's IAM role, so it can read its own SSM parameters. Put environment variables the sandbox needs at run time in /opt/linuxbench-worker-setup.env, which the worker sources after the script. Pick the hardware with --ec2-arg instance_type=...; a microVM sandbox needs /dev/kvm, which only *.metal types expose. The setup script and instance type are part of the fleet ID, so a run that differs only in them does not resume another's results.

ct run eval --ec2 --ec2-arg instance_type=c5.metal \ --ec2-arg worker_setup=agent-glovebox/evals/control_tower/ct-worker-setup.sh \ --extra-src ~/agent-glovebox --sandbox evals.control_tower.sbx_gb:sbx-gb -e linuxarena:file_indexing ...

Scaling

The main knob for throughput is --ec2-arg max_workers=N. Each worker starts at 10 concurrent evals and adapts to rate-limit pressure up to concurrency_per_worker. A fleet first launches up to 24 workers, fewer when max_workers or the pending work is smaller, and adds workers up to max_workers while it is CPU-bound.

WorkersConcurrent evalsEC2 cost/hrGood for
550~$3.50Small runs (<100 samples)
12120~$8.50Medium runs (100-300 samples)
24240~$17Large runs (300-600 samples)
40400~$28Very large runs (600+ samples)

Fleet Options (--ec2-arg key=value)

KeyDescription
max_workersMaximum EC2 worker instances (default: 64)
concurrency_per_workerMax concurrent samples per worker (default: 32)
instance_typeInstance type for workers (default: c7i.4xlarge)
regionAWS region (default: us-east-1)
spotUse spot instances, with on-demand fallback
worker_setupScript inside an --extra-src tree that workers run as root before taking jobs
worker_userAccount the worker process runs as (default: root)
amiAMI ID or SSM parameter naming one (default: the pre-baked worker AMI)
estimatePrint cost estimate and exit
auto_confirmSkip the confirmation wait
new_fleetForce fresh fleet, skip auto-resume

For worker_user, the bootstrap passes the name to the setup script as LINUXBENCH_WORKER_USER, adds the account to the docker group, and chowns /opt/linuxbench to it. For ami, an SSM parameter looks like /aws/service/canonical/ubuntu/server/24.04/stable/current/amd64/hvm/ebs-gp3/ami-id.

Resume

Fleet runs resume automatically. The fleet ID is a hash of the git SHA, policy, tasks, and epochs, so after a controller crash, credential expiry, or network failure, re-run the same command. The controller finds the existing manifest in S3, lists the completed .eval files, enqueues only the remaining work, and launches fresh workers for it.

What Gets Created

  • SQS queues: lb-fleet-{fleet_id}-work, lb-fleet-{fleet_id}-done, lb-fleet-{fleet_id}-dlq
  • S3 objects: {fleet_id}/manifest.json, {fleet_id}/{job_id}.eval, artifacts/{sha}.tar.gz
  • Security group: linuxbench-fleet-worker-sg (outbound only, no inbound)
  • EC2 instances: tagged with FleetId={fleet_id}, auto-terminate on shutdown

Failure Handling

FailureHandling
Worker crashMessage reappears after visibility timeout, another worker retries
Poison messageDLQ after 3 attempts
Controller crashRe-run same command to auto-resume
>20% failure rateController terminates workers, preserves state for resume
Missing API keysWorker fails fast with clear error

Worker AMI

Fleet workers boot from a pre-baked AMI with dependencies and Docker images already installed, which cuts worker bootstrap from ~4 min to ~65s. If no AMI exists in SSM, workers fall back to stock Ubuntu 22.04 and install everything at boot.

Build the AMI with:

# Via CLI (incremental by default -- only pulls changed images) ct cloud build-ami # Force full rebuild from stock Ubuntu ct cloud build-ami --fresh # Build in a different region ct cloud build-ami --region us-west-2 # Or directly ./scripts/build_worker_ami.sh --region us-east-1 --fresh

The build takes ~15-20 minutes, mostly EBS snapshot creation. It launches a builder instance, installs dependencies, pulls the Docker images for all environments (~74 images), snapshots the AMI, stores its ID in SSM at /linuxbench/worker-ami-id, deregisters the previous AMI (only one is kept), and terminates the builder.

The AMI contains:

  • Docker + docker-compose-plugin
  • Python 3.13 + uv (installed for both ubuntu and root users)
  • Git + AWS CLI
  • Docker images for all environments
  • Python wheel cache at /opt/uv-cache, which speeds up uv sync at boot

Workers fetch at boot:

  • Repo code from S3 at a pinned git SHA: HEAD when the worktree is clean, or a tree SHA of the dirty worktree, so uncommitted changes reach workers. Untracked files ship only after git add; the controller warns at launch if any are present.
  • Python dependencies (uv sync --frozen, using cached wheels)
  • Secrets (API keys, GH_TOKEN, Docker registry credentials) from SSM

Rebuild the AMI after environment Docker images are updated and pushed, or after adding environments. Code-only changes need no rebuild.

Granting Fleet Access (Admins)

Both flows below need an admin and use the scoped policy in ops/docs/fleet-iam-policy.json.

SSO (preferred). This needs Identity Center admin in the org management account. As a one-time setup, enable the ABAC session tag and create the permission set:

INSTANCE_ARN=$(aws sso-admin list-instances --query 'Instances[0].InstanceArn' --output text) STORE_ID=$(aws sso-admin list-instances --query 'Instances[0].IdentityStoreId' --output text) aws sso-admin create-instance-access-control-attribute-configuration \ --instance-arn "$INSTANCE_ARN" \ --instance-access-control-attribute-configuration \ 'AccessControlAttributes=[{Key=LaunchedBy,Value={Source=["${path:userName}"]}}]' PS_ARN=$(aws sso-admin create-permission-set --instance-arn "$INSTANCE_ARN" \ --name LinuxbenchFleet --session-duration PT12H \ --query 'PermissionSet.PermissionSetArn' --output text) aws sso-admin put-inline-policy-to-permission-set --instance-arn "$INSTANCE_ARN" \ --permission-set-arn "$PS_ARN" --inline-policy file://ops/docs/fleet-iam-policy.json

If the instance already has an ABAC configuration, use update-instance-access-control-attribute-configuration and merge the LaunchedBy attribute into the existing list. The mapping tags every SSO session with LaunchedBy=<userName>. The policy's *Sso statements match that tag, and the controller tags instances with the role session name, which Identity Center also sets to <userName>.

Per-teammate assignment:

USER_ID=$(aws identitystore list-users --identity-store-id "$STORE_ID" \ --filters AttributePath=UserName,AttributeValue=<userName> \ --query 'Users[0].UserId' --output text) aws sso-admin create-account-assignment --instance-arn "$INSTANCE_ARN" \ --target-id 326120904690 --target-type AWS_ACCOUNT \ --permission-set-arn "$PS_ARN" --principal-type USER --principal-id "$USER_ID"

Static IAM keys (fallback). This needs IAM admin on the fleet account (326120904690). Create an IAM user and attach the scoped linuxbench-fleet-access managed policy:

aws iam create-user --user-name teammate-name aws iam attach-user-policy --user-name teammate-name \ --policy-arn arn:aws:iam::326120904690:policy/linuxbench-fleet-access aws iam create-access-key --user-name teammate-name

Send the teammate the resulting AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY (no session token) for their .env. The policy must be a managed policy because it exceeds IAM's 2 KB inline limit; create it once from ops/docs/fleet-iam-policy.json if it doesn't exist yet.

Propagating policy changes. After editing ops/docs/fleet-iam-policy.json, re-run the put-inline-policy-to-permission-set command from the setup above, then:

aws sso-admin provision-permission-set --instance-arn "$INSTANCE_ARN" \ --permission-set-arn "$PS_ARN" --target-type ALL_PROVISIONED_ACCOUNTS aws iam create-policy-version \ --policy-arn arn:aws:iam::326120904690:policy/linuxbench-fleet-access \ --policy-document file://ops/docs/fleet-iam-policy.json --set-as-default

create-policy-version fails once 5 versions exist, so delete the oldest non-default version first.

Use the scoped policy rather than AdministratorAccess. It ensures that:

  • ec2:TerminateInstances only works on instances tagged LaunchedBy=<your identity> (IAM username, or SSO session name via the ABAC tag)
  • ec2:RunInstances requires the LaunchedBy and ManagedBy tags
  • SQS access is scoped to lb-fleet-* queues
  • S3 access is scoped to the lb-fleet-results-migrated bucket

Monitoring Fleet Resources

# Show your running instances and queues ct cloud status # Show all control-tower resources (all users) ct cloud status --all-users

Debugging

  • Bootstrap logs: uploaded to s3://lb-fleet-results-migrated/{fleet_id}/logs/{instance_id}-bootstrap.log when bootstrap exits, with timing at each step
  • EC2 console: link printed after launch
  • SSH: via EC2 Instance Connect (emergency debugging)
  • Stale warning: the controller warns after 10 minutes with no progress
  • Hard timeout: the controller terminates the fleet after 3 hours with no progress while workers are silent (or 1.5 × the sample time limit, if longer), and after 6 hours with no completed sample even while workers are active