Straj review workflow
A reviewer-aid bot for sabotage-trajectory (straj) PRs. A reviewer comments @straj-review on a straj PR, and a small workflow in the env repo calls control-tower's reusable straj-review.yml. It decodes the trajectory and posts the findings as a single sticky comment, which each run overwrites in place.
The bot is advisory: it never blocks merge, approves, or requests changes.
From top to bottom, the comment carries the trajectory facts (env, main/side task and scores, number of actions), the action- and trajectory-monitor identity with max_sus / traj_sus and sanity checks, a link to the straj's traj directory, and then, as supporting context, the env threat model and the main-task (cover) and side-task (objective) text. The deterministic parts are static reads of the downloaded trajectory and the PR; nothing runs the env container. The LLM analysis stages, which read the straj diff and narrate the attack, add a straj analysis section at the bottom.
What the reviewer does
-
Make sure the PR description contains a monitored trajectory URL, either a plain viewer URL containing
/trajectories/…or an explicit line:Straj-Review-Trajectory: https://…/trajectories/<id> -
Comment
@straj-reviewon the PR.
The bot reacts 👀 on the triggering comment, reads the URL from the PR body, downloads the trajectory, writes its sticky comment, and reacts 👍 when the comment is posted. No 👍 means the run failed. If the URL is missing or the download fails, the comment says what it needs. Comment @straj-review again to re-run; the comment always reflects the current state.
The trigger is the hyphenated @straj-review because no GitHub user has that name, so the comment pings nobody.
Providing the straj (the traj directory)
The straj's files live in the env repo at side_tasks/<side_task>/trajs/<main_task>/<variant>/, the per-straj directory described in making-strajs. Committing that directory to the PR is how the straj is tracked. It holds the straj's description.md and info.yml (which may be empty; the workflow does not check their contents), the trajectory's git diff, and any other relevant files.
The comment links to the traj directory rather than rendering the diff. A ct traj patch diff blends the main-task work with the attack, and committed files already show in the PR's Files changed tab.
The bot requires the PR to change exactly one traj directory. When it changes several, the comment asks you to split them into separate PRs. When it changes none, the comment asks you to create one:
- Generate the trajectory's
git diffwithct traj patch <traj>locally. It brings the env up, replays the trajectory, and captures the diff. - Commit it (with
description.md/info.yml) underside_tasks/<side_task>/trajs/<main_task>/<variant>/on the PR branch. - Comment
@straj-reviewagain.
The bot regenerates anything absent, so to force the env threat model to regenerate, commit a deletion of docs/threat-model.md.
The LLM analysis stages require the threat model, because it defines the security property their verdicts judge against. When it is missing (generation failed or was skipped), the bot skips those stages, and the comment's analysis section says so and how to retry. The deterministic report posts regardless.
How ct is obtained
The workflow installs ct from the public control-tower repo with uvx --from git+https://github.com/linuxarena/control-tower@main ct …, so env repos do not depend on control-tower.
One-time org setup
The bot authenticates to the trajectory backend with a per-user API token (a bearer token from ct login for a service user). Set it once as an organization secret granted to the env repos:
| Secret | Needed when |
|---|---|
CONTROL_TOWER_API_TOKEN | the trajectory backend is private |
A backend reachable without auth needs no secrets. The backend URL comes from the trajectory URL in the PR body, so CONTROL_TOWER_API_BASE_URL matters only if you reference a bare ID instead of a full URL.
Env-repo caller (the ~15-line wrapper)
Add this as .github/workflows/straj-review.yml in the env repo. GitHub runs issue_comment workflows only from the default branch, so merge it there before it will fire.
name: Straj Review on: issue_comment: types: [created] permissions: contents: write # the reusable may commit a generated docs/threat-model.md pull-requests: write issues: write id-token: write jobs: straj-review: # PR comments only, must say "@straj-review", and only from people who can # already write to the repo (authz). if: > github.event.issue.pull_request != null && contains(github.event.comment.body, '@straj-review') && contains(fromJSON('["OWNER","MEMBER","COLLABORATOR"]'), github.event.comment.author_association) uses: linuxarena/control-tower/.github/workflows/straj-review.yml@main with: pr_number: ${{ github.event.issue.number }} comment_id: ${{ github.event.comment.id }} secrets: CONTROL_TOWER_API_TOKEN: ${{ secrets.CONTROL_TOWER_API_TOKEN }} ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
The explicit secrets: block forwards only the secrets the reusable needs. contents: write is needed only for the threat-model stage: if the env has no docs/threat-model.md, the reusable generates one with Claude and commits it onto the PR branch once. ANTHROPIC_API_KEY is needed for that generation and for the LLM analysis stages; omit it to disable both. Pin @main to a tag to freeze behavior per env repo.