Trajectories

trajectories reads agent behavior straight off Inspect eval samples; the .eval is the single source of truth.

A trajectory is the recorded behavior of one sample:

  • environment/main-task/side-task context (SampleMetadata on the sample)
  • action/tool-call history
  • timing/cost/monitor metadata
  • optional run linkage (run_id) and eval linkage (eval_log_url, sample ID)

This package handles:

  • sample accessors (sample.py): executed actions, task success, monitor scores, costs, diagnostics
  • extraction internals (from_eval_log.py, timing_utils.py, tool_calls.py)
  • viewer thin-doc assembly for uploads (mongo_docs.py)
  • storage and retrieval of metadata+sample bundles (io.py, api.py, viewer_api.py)
  • trajectory-focused CLI commands (cli.py)
  • TUI support for interactive trajectory review/edit (in control_tower.live.tui)

CLI Overview

The ct traj group is defined in control_tower.trajectories.cli:traj_cli. Its primary commands:

  • ct traj list [--fs] [-e ENV] [-t MAIN_TASK] [-s SIDE_TASK] [--tag TAG] [--limit N]

    • Lists trajectories from the API (default) or local bundle files (--fs).
  • ct traj download [<trajectory_id_or_viewer_url>] [--eval] [-e ENV] [-t MAIN_TASK] [-s SIDE_TASK] [--tag TAG] [-n N] [--dataset NAME] [--dataset-version V] [-o OUTPUT] [--overwrite]

    • Downloads one trajectory by ID from the viewer API, or by viewer URL (e.g. https://basharena-data.vercel.app/trajectories/<id>; see Downloading from the Viewer).
    • With filter options instead of an ID, bulk-downloads the latest N matching trajectories (default 20, max 500).
    • --dataset NAME downloads every trajectory in a dataset, with no count limit and no auth or Mongo URI required. It fetches the dataset's trajectory ID list and downloads by ID in batches. --dataset-version V picks a version (default: latest).
    • --eval downloads the trajectory's full .eval archive instead of the bundle JSON.
  • ct traj patch <traj> ...

    • Generates a patch from a trajectory using live tooling.
  • ct traj apply <eval_log_file> ...

    • Applies file edits from an eval log into an environment codebase.
  • ct traj tui <trajectory_path> ...

    • Opens the trajectory TUI for interactive review/editing.

Cost accessors on a sample

The functions below live in control_tower.trajectories.sample; each takes the EvalSample you want to read.

Every cost except cost_by_role is inspect's stamp on usage.total_cost. Sample totals come from sample.model_usage; per-call figures come from each ModelEvent. An Inspect response-cache read is left out of model_usage (actual spend) and marked on timing.model_calls[i].inspect_cache_hit. Provider prompt-cache discounts are already in the stamp.

  • Per-call: action.timing.model_calls[i].cost_usd.
  • Run total: sample.agent_cost(sample) is {model: usage.total_cost} from sample.model_usage, which counts every model call in the run, live monitors included; use cost_by_role for the agent alone. A null amount is an unpriced model.
  • Monitors: sample.monitor_cost(sample) -> Priced sums each monitor score's cost_usd metadata (see Cost Tracking in function-registry.md); a score with no cost_usd is a call nobody could price. It walks ensemble members without double-counting the parent.
  • By role: cost_by_role(sample) -> dict[str, Priced] prices each model call at the rate table's current rates, not inspect's stamp, and keys it by the call's role (names in models/roles.py; AGENT_ROLES is the agent). A response-cache hit counts at its original tokens, so the figure is what a fresh run costs. A call with no role raises, which it does on logs from before every Control Tower model carried one.
  • Grand total: sample.total_cost(sample) -> Priced is agent + monitors, per model. A None in its by_model means the total is short, and the viewer shows the known total as $N+. sample.total_cost_usd(sample) is known spend, or None only when nothing at all was priced.

Workspace diff on a sample

When a run's main task declared a final_diff config, control_tower.trajectories.sample.workspace_diff(sample) reads the final workspace diff captured at scoring time: the unified diff text (truncated to the task's byte budget), a short stat summary, the baseline it was computed against, and whether truncation applied. It is None for tasks without the config, for simulated and replay-sandbox runs, and for samples without captured diff metadata. See Monitoring for how monitors consume it.

How trajectories fits with other CLI groups

  • ct run ... runs experiments and can produce trajectories.
  • ct runs ... manages run-level containers around groups of trajectories.
  • ct traj ... is the direct trajectory manipulation and analysis surface.

Typical workflows

List trajectories from API:

  • ct traj list -n 20

Download one trajectory:

  • ct traj download <trajectory_id>
  • ct traj download https://basharena-data.vercel.app/trajectories/<id> (viewer URL)
  • ct traj download <trajectory_id> --eval (full .eval archive instead of bundle JSON)

Bulk-download the latest matching trajectories:

  • ct traj download -e web_scraping --tag baseline -n 100

Download every trajectory in a dataset (latest version, or a specific one):

  • ct traj download --dataset LaStraj
  • ct traj download --dataset LaStraj --dataset-version 2