Database
We persist evaluation artifacts in MongoDB so they can be replayed, inspected, and surfaced in dashboards.
CLI access
The ct CLI has no direct MongoDB access. Every read and write goes through the viewer's HTTP API, authenticated with a per-user token: run ct login to cache one, or set CONTROL_TOWER_API_TOKEN (the same token for reads and writes). MONGOURI lives only in a viewer deployment's server environment and in the ops/ operator tooling.
Uploads (ct run eval, ct runs make, ct run monitor) need only the API token, with no AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY, because the data host presigns an add-only S3 PUT for the client. Reading a .eval back is a direct S3 GET and uses your own AWS credentials, which must be able to GET the eval-log bucket. Runs uploaded before the S3 migration carry an eval_log_url on the eval-log viewer instead, and reading those needs LOGS_AUTH_PASSWORD; ask the linuxarena maintainers for the current value.
ct run eval --no-upload, ct run monitor <file> --no-upload, ct live commands, and ct run side-task-review operate on local files and touch no remote service. ct traj download and ct traj list read through the website API.
How uploads reach the store
The data host holds the S3 and Mongo credentials and mediates every write:
POST /api/eval-logs/uploadplans an add-only upload for a.evalfile as a single presigned PUT. It rejects with 409 if the key already exists, or 413 ifsize_bytesexceeds S3's 5 GB single-PUT limit; multipart upload isn't implemented.POST /api/trajectories/registerwrites the run header and/or trajectory documents in one Mongo transaction; see HTTP API — register semantics for the full contract.
Connection Details
- Collection names come from the
_ct_opsrecord:MONGOURIandMONGO_DB(defaultlinuxbench) are the only Mongo settings a viewer takes. The physical collection for each role (trajectories,runs,datasets,dataset_versions,sabotage_evals,pricing) is recorded in the_ct_opsstate document, whichops/writes on install and moves on every release. The viewer reads it from there and re-reads it every 30 seconds (packages/viewer/lib/mongo-client.ts). The viewer refuses a database with no record and reports thatct-ops setup installhas not been run. Thecompose.ymlstack runs it itself (ops-init) before the viewers start. A preview deployment setsPREVIEW_VERSION=<version>to serve that preview generation instead of the canonical one. - Replica set required:
MONGOURImust point at a replica set, because register writes the run header and its trajectories in one Mongo transaction. A single-node replica set is fine. Atlas/production is already one, and the self-hostcompose.ymlruns Mongo with--replSet rs0and connects with?replicaSet=rs0&directConnection=true. Against a standalonemongod, register silently degrades to sequential, non-atomic writes, so an interrupted register can leave trajectories without their run header. - Trajectories collection: Mongo stores one metadata document per trajectory, keyed by
trajectory_id. Each document points to the source of truth (eval_log_url,eval_sample_id, andepoch) and carries the filter and display fields the viewer queries: environment, main-task, and side-task identifiers, success flags, monitor scores, tags, and timing. Full trajectory bodies (actions, prompts, tool output) live only in the.evallog in S3 and are read from there.control_tower.trajectories.mongo_docs.build_trajectory_documentbuilds the documents. run_idis a soft reference: a trajectory'srun_idis a client-generated key that the store does not enforce, and a client may register trajectories with no run header (postingrun=None). A trajectory whoserun_iddoes not resolve renders and lists normally; its "Run" link points at/runs/<run_id>, which returns a 404. Runs are enumerated only from the runs collection, so a danglingrun_idnever produces a phantom run.- Pricing collection: model rates, one
{model, pricing, manual}document per model, the same row shape as the JSON rate file. Onlyct-ops pricing(set,del,flush,sync) writes it; the viewer serves it read-only atGET /api/pricing, whichctreads whenCONTROL_TOWER_PRICING=api(see Cost Tracking). - Runs collection:
agent_runs.io.upload_run(via the register API) produces a run document containing therun_id, optional human-readablename, timestamps, the URL of the uploaded.evallog, and ametadatablob summarizing the environment, tasks, policy, and model settings. Runs also store a top-levelproducerfield (for exampleai_eval,monitor, orsample) so the website can filter agent and monitor runs, the set oftrajectory_ids created for that evaluation, and acreated_byfield. The register route setscreated_byandcreated_atitself, takingcreated_byfrom the authenticated token's username and discarding any client-sent value.UPLOAD_USER_NAMEonly labels locally built trajectory objects. - Ingestion workflow:
upload_runis the only client-side upload entry point. It uploads the run's.evalto S3 through the presigned-upload API, then callsPOST /api/trajectories/registerin bounded batches that stay under the data host's request-size limit, writing the run header and one metadata document per sample. The run header rides on the final batch, so an interrupted earlier batch leaves run-less trajectories rather than a header that promises missing members. Each batch is safe to retry. When no.evalexists yet (post-hoc monitoring, sampling, ad-hoc uploads), one is generated from the in-memory trajectories and uploaded first, so every Mongo document is backed by a.evalin S3. - Indexes: operators create indexes per cluster from the
ops/package; the app and CLI never do. Runct-ops setup installon a new deployment andct-ops setup indexesto repair drift;ct-ops release upgradeprovisions them as its last step. Index specs are frozen per release inops/ct_ops/index_sets/. - Local caches: trajectory bundles cache in
~/.cache/control-tower/trajectory-bundles/and run metadata in~/.cache/control-tower/runs/; downloaded trajectories default todata/traj/and run exports todata/run_trajs/. Use the helpers incontrol_tower.trajectories.ioandcontrol_tower.agent_runs.iorather than editing these directories by hand, so the cache stays consistent with Mongo.
Publishing Trajectories to the Website
Uploaded runs appear in the website viewer, where the team can browse them.
Upload while a script runs
All eval entry points, such as ct run eval, upload the .eval log and create the run as soon as the evaluation finishes. Pass --no-upload to disable this.
Upload a saved .eval
For a completed run, ct runs make logs/<file>.eval uploads the eval file (best effort), creates a run document, and uploads the derived trajectories.
Add descriptions and tags in the UI
- Each run has an optional
nameanddescriptionthat you can edit on the website. - The upload's metadata (environment, main task, side task, models, scores, etc.) appears as filterable facets, the "tags" in the UI. You can extend them by editing metadata in the web interface.
Guidance on Assigning Tags
- Each new main task needs a trajectory that passes its scorer; tag it
qa. - Tag the sabotage trajectories built for each side task (at least one per side task)
qa. - Add the
aiorhumantag to these according to who produced them. - Tag every sabotage trajectory (one that attempts a main and a side task)
straj.
Uploading Sabotage Trajectories (strajs)
For creating and uploading human sabotage trajectories, see Making Strajs.
Terms
See Eval-Native Architecture for what is stored where, and Trajectories for how ct run, ct runs, and ct traj relate.