Database

We persist evaluation artifacts in MongoDB so they can be replayed, inspected, and surfaced in dashboards.

CLI access

The ct CLI has no direct MongoDB access. Every read and write goes through the viewer's HTTP API, authenticated with a per-user token: run ct login to cache one, or set CONTROL_TOWER_API_TOKEN (the same token for reads and writes). MONGOURI lives only in a viewer deployment's server environment and in the ops/ operator tooling.

Uploads (ct run eval, ct runs make, ct run monitor) need only the API token, with no AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY, because the data host presigns an add-only S3 PUT for the client. Reading a .eval back is a direct S3 GET and uses your own AWS credentials, which must be able to GET the eval-log bucket. Runs uploaded before the S3 migration carry an eval_log_url on the eval-log viewer instead, and reading those needs LOGS_AUTH_PASSWORD; ask the linuxarena maintainers for the current value.

ct run eval --no-upload, ct run monitor <file> --no-upload, ct live commands, and ct run side-task-review operate on local files and touch no remote service. ct traj download and ct traj list read through the website API.

How uploads reach the store

The data host holds the S3 and Mongo credentials and mediates every write:

  • POST /api/eval-logs/upload plans an add-only upload for a .eval file as a single presigned PUT. It rejects with 409 if the key already exists, or 413 if size_bytes exceeds S3's 5 GB single-PUT limit; multipart upload isn't implemented.
  • POST /api/trajectories/register writes the run header and/or trajectory documents in one Mongo transaction; see HTTP API — register semantics for the full contract.

Connection Details

  • Collection names come from the _ct_ops record: MONGOURI and MONGO_DB (default linuxbench) are the only Mongo settings a viewer takes. The physical collection for each role (trajectories, runs, datasets, dataset_versions, sabotage_evals, pricing) is recorded in the _ct_ops state document, which ops/ writes on install and moves on every release. The viewer reads it from there and re-reads it every 30 seconds (packages/viewer/lib/mongo-client.ts). The viewer refuses a database with no record and reports that ct-ops setup install has not been run. The compose.yml stack runs it itself (ops-init) before the viewers start. A preview deployment sets PREVIEW_VERSION=<version> to serve that preview generation instead of the canonical one.
  • Replica set required: MONGOURI must point at a replica set, because register writes the run header and its trajectories in one Mongo transaction. A single-node replica set is fine. Atlas/production is already one, and the self-host compose.yml runs Mongo with --replSet rs0 and connects with ?replicaSet=rs0&directConnection=true. Against a standalone mongod, register silently degrades to sequential, non-atomic writes, so an interrupted register can leave trajectories without their run header.
  • Trajectories collection: Mongo stores one metadata document per trajectory, keyed by trajectory_id. Each document points to the source of truth (eval_log_url, eval_sample_id, and epoch) and carries the filter and display fields the viewer queries: environment, main-task, and side-task identifiers, success flags, monitor scores, tags, and timing. Full trajectory bodies (actions, prompts, tool output) live only in the .eval log in S3 and are read from there. control_tower.trajectories.mongo_docs.build_trajectory_document builds the documents.
  • run_id is a soft reference: a trajectory's run_id is a client-generated key that the store does not enforce, and a client may register trajectories with no run header (posting run=None). A trajectory whose run_id does not resolve renders and lists normally; its "Run" link points at /runs/<run_id>, which returns a 404. Runs are enumerated only from the runs collection, so a dangling run_id never produces a phantom run.
  • Pricing collection: model rates, one {model, pricing, manual} document per model, the same row shape as the JSON rate file. Only ct-ops pricing (set, del, flush, sync) writes it; the viewer serves it read-only at GET /api/pricing, which ct reads when CONTROL_TOWER_PRICING=api (see Cost Tracking).
  • Runs collection: agent_runs.io.upload_run (via the register API) produces a run document containing the run_id, optional human-readable name, timestamps, the URL of the uploaded .eval log, and a metadata blob summarizing the environment, tasks, policy, and model settings. Runs also store a top-level producer field (for example ai_eval, monitor, or sample) so the website can filter agent and monitor runs, the set of trajectory_ids created for that evaluation, and a created_by field. The register route sets created_by and created_at itself, taking created_by from the authenticated token's username and discarding any client-sent value. UPLOAD_USER_NAME only labels locally built trajectory objects.
  • Ingestion workflow: upload_run is the only client-side upload entry point. It uploads the run's .eval to S3 through the presigned-upload API, then calls POST /api/trajectories/register in bounded batches that stay under the data host's request-size limit, writing the run header and one metadata document per sample. The run header rides on the final batch, so an interrupted earlier batch leaves run-less trajectories rather than a header that promises missing members. Each batch is safe to retry. When no .eval exists yet (post-hoc monitoring, sampling, ad-hoc uploads), one is generated from the in-memory trajectories and uploaded first, so every Mongo document is backed by a .eval in S3.
  • Indexes: operators create indexes per cluster from the ops/ package; the app and CLI never do. Run ct-ops setup install on a new deployment and ct-ops setup indexes to repair drift; ct-ops release upgrade provisions them as its last step. Index specs are frozen per release in ops/ct_ops/index_sets/.
  • Local caches: trajectory bundles cache in ~/.cache/control-tower/trajectory-bundles/ and run metadata in ~/.cache/control-tower/runs/; downloaded trajectories default to data/traj/ and run exports to data/run_trajs/. Use the helpers in control_tower.trajectories.io and control_tower.agent_runs.io rather than editing these directories by hand, so the cache stays consistent with Mongo.

Publishing Trajectories to the Website

Uploaded runs appear in the website viewer, where the team can browse them.

Upload while a script runs

All eval entry points, such as ct run eval, upload the .eval log and create the run as soon as the evaluation finishes. Pass --no-upload to disable this.

Upload a saved .eval

For a completed run, ct runs make logs/<file>.eval uploads the eval file (best effort), creates a run document, and uploads the derived trajectories.

Add descriptions and tags in the UI

  • Each run has an optional name and description that you can edit on the website.
  • The upload's metadata (environment, main task, side task, models, scores, etc.) appears as filterable facets, the "tags" in the UI. You can extend them by editing metadata in the web interface.

Guidance on Assigning Tags

  • Each new main task needs a trajectory that passes its scorer; tag it qa.
  • Tag the sabotage trajectories built for each side task (at least one per side task) qa.
  • Add the ai or human tag to these according to who produced them.
  • Tag every sabotage trajectory (one that attempts a main and a side task) straj.

Uploading Sabotage Trajectories (strajs)

For creating and uploading human sabotage trajectories, see Making Strajs.

Terms

See Eval-Native Architecture for what is stored where, and Trajectories for how ct run, ct runs, and ct traj relate.