CI checks and reports

Run uv sync --locked --dev --all-extras and make check for the same Python checks as the PR matrix. make github-pr-checks tests a synthetic merge against current origin/main; it requires committed changes. Pre-commit uses the same lint/format commands and locked tools. CI uses Python 3.13, uv 0.11.29, and Bun 1.3.11 where applicable.

Gates and focused checks

PR gate succeeds only when every Python matrix check and the changed-line convention check succeed. It does not change approval requirements. Repository administrators can require this stable check name in the organization ruleset once the workflow is on the default branch; adding the workflow does not configure branch protection by itself. The docs, viewer, metrics, and sandbox checks run only when a PR changes their inputs, so do not configure them as unconditional required checks.

make conventions-check PR_BASE_REF=origin/main checks the Python comments and docstrings a PR adds. Issue IDs in those additions fail the check; long comment blocks and docstrings are reported as review hints. It posts no comments on PRs. Lint, import boundaries, and public type completeness keep their own checks.

make metrics-check reads four committed replay samples through the public saved-run/report pipeline, verifies sample/scorer/unscored accounting, and checks independently calculated full/partial AUROC and average precision cases. It makes no model calls. The metrics workflow runs the same check against the PR base and the proposed merge, each with its own locked dependencies and the same fixtures, and publishes JSON and a compact comparison. Mathematical or accounting violations on the head fail; numerical changes are shown for review. If the base fails the current check, the summary marks the comparison unavailable and still requires the head to pass. Browser parity, confidence interval, bootstrap, and configuration tests run in make check.

make sandbox-check runs the real Docker network-isolation integration checks, and its workflow runs on sandbox and dependency changes. The replay jobs exercise honest and attack trajectories; summaries distinguish first-try matches from retry recovery, and unexpected sample IDs fail. Environment baseline reports compare declared and observed task identities, including equal-sized but different sets.

Description reports and bot replies

One collapsed section in each PR description shows source and test changes, category totals, and current-head workflow results, with links to detailed metrics and performance summaries. The writer runs once per push, after the last watched workflow for that head completes. It checks out only the default branch, reads PR data through the API, does not execute fork code, preserves author prose outside its markers, and serializes updates with the preview-link writers. Counts are path-based size measurements, not quality scores.

Claude automatic reviews skip draft PRs; explicit maintainer mentions still run. If a comment invocation fails or GitHub marks it action_required, the feedback workflow posts at most one status reply in that comment's thread and updates it on reruns. The caller records the exact comment ID in its run name, and the comment's author and publication time must match the run (pending-review comments use the review submission time). --comment is only for read-only diagnosis of an older run.

For an approval-blocked inline question, a read-only Claude request answers in that same reply using the title and diff. It has no tools, executes no fork code, and verifies that the head did not move. The full available diff, PR description, and review thread are counted with the selected model's token-counting API, and the model's Models API limit, minus reserved output and thinking capacity, decides whether the complete request fits. An oversized request fails explicitly without discarding context. Binary or oversized patches that GitHub omits are marked unavailable. A repeated blocked event does not generate another answer, and provider failures fail the feedback job. Recovery never approves a fork workflow or applies a fix; approval stays a maintainer action. Mentions in a submitted review without a comment cannot currently be correlated by this workflow.

Docs cache

The environment catalog is built per setting. Each setting's records are cached by its resolved public commit, the generator code, pyproject.toml without its dependency groups and extras, the runtime dependencies locked for it (what the docs jobs install), and Python/platform (plus the live vLLM dataset revision for vLLM), so a run regenerates only the settings whose commit moved, cloning just those at depth 1. Isolated settings export in parallel with the in-process imports, each seeing only its own setting. The Hugging Face datasets the exports download are cached under a key naming the dataset snapshots they contain; the latest entry is restored by prefix, and a new one is saved only when a build adds snapshots. The merged catalog is cached under one exact key, and a restored artifact must match its input manifest. CLI and Python references take seconds and are regenerated on every run. The docs check selects the same source paths as the production docs deploy, which make workflow-check enforces.

The docs check also deploys the PR's docs preview, so the catalog is generated once. Production docs and the docs image use the same generation action. A cache hit skips settings installation and generation; docs lint and build still run. Settings changes appear on the next triggered docs build.

Timing and reliability

The metrics workflow interleaves three base and head measurements of fresh-process CLI startup and saved-run report generation on the same runner, after installing dependencies; operating-system file caches are not controlled. Reports include medians, ranges, absolute and relative deltas, and revision and lockfile identity. The small fixtures measure pipeline overhead, not production-scale or model latency. Timing changes are informational until a stable workload and regression budget are agreed.

The daily CI health workflow summarizes a bounded sample of recent repository runs: successful-run p50/p95, failures, cancellations, reruns, and expensive steps. Its JSON artifact keeps the underlying measurements for 30 days. Workflow elapsed time includes queue and cleanup time and is not billed runner time. It posts no issues or PR comments. Replay artifacts separately show failures recovered by retries, because a green final workflow does not show the first-attempt pass rate.

Repository-owned Claude feedback and straj reviews read the model from scripts/ci/models.json; change its claude value to change both. Straj review reads that file from Control Tower's default branch, including when invoked from an environment repository. The organization-wide @claude implementation sets its own model in linuxarena/.github/.github/workflows/claude-reusable.yml and has no caller model input.

Workflow validation and weekly coverage

make workflow-check runs actionlint 1.7.12 and zizmor 1.30.1 offline. Install the pinned actionlint binary on PATH; CI verifies the release archive checksum. Findings are compared with the pre-gate revision recorded in scripts/ci/workflow_check.py: new syntax errors and new high-confidence, high-severity findings fail, and existing or lower-confidence findings stay visible in the JSON artifact. The check also verifies that shared inputs and recorded viewer fixtures stay in the docs and viewer trigger paths. The two description writers have a narrow exception for actionlint's unsupported queue: max field. Shellcheck and pyflakes are not part of this check.

The weekly offline workflow runs Python/ops checks, installed-wheel consumers, saved-run metrics, real Docker isolation, both recorded replays, and frontend lint/build/contracts without path selection. Generated docs data is rebuilt from pinned public settings without restoring its cache. It runs on Mondays, on manual dispatch, and when its own workflow changes, with 15–20 minute job timeouts. It performs no deployment or model evaluation, posts no PR comments, and retains diagnostic artifacts for seven days.