Integration test quality

Claude-assisted review of environment integration tests.

Purpose

ct envs integration-test-quality uses Claude Code on the host to judge whether an environment's integration tests are truly end-to-end from a user’s perspective, with optional passes for task conflicts and coverage gaps. It writes integration_test_quality.md under the output directory.

It requires Claude Code (the claude CLI) on your PATH.

Commands

ct envs integration-test-quality

Analyze integration test quality for an environment.

Usage:

ct envs integration-test-quality [OPTIONS]

Options:

  • -e, --env — Environment name (prompts if omitted)
  • -o, --output-dir — Output directory (default: data/test_quality_runs/<env>)
  • --check-tasks — Whether completing main/side tasks would likely break integration tests
  • --coverage — Coverage / gap analysis for integration tests
  • --fix — After analysis, run Claude to edit integration tests using the saved report. In a non-interactive terminal the fix runs automatically; in an interactive terminal you are asked to confirm before any edits.

Examples:

ct envs integration-test-quality -e web_scraping ct envs integration-test-quality -e web_scraping --check-tasks ct envs integration-test-quality -e web_scraping --coverage ct envs integration-test-quality -e web_scraping --check-tasks --coverage ct envs integration-test-quality -e web_scraping --fix

Output

The command writes or appends to data/test_quality_runs/<env>/integration_test_quality.md (or the file under --output-dir), prints a short summary, then prints the report path.

The main analysis is plain text with no markdown headings beyond the top title:

  • A title line: # Integration test quality: <env>
  • Summary:, a blank line, then a 1–2 sentence overall assessment of the suite.
  • List of problems, a blank line, then one line per issue as problem: how to fix.

--check-tasks appends a Task conflicts section in the same shape: a header line, a blank line, then one line per task or area and why tests would break.

--coverage appends a Coverage gaps section in the same shape, with one line per gap and its suggested coverage.

Criteria (good vs bad integration tests)

Good tests:

  • Exercise the system through HTTP, CLI, or other stable external interfaces
  • Use inputs and read outputs the way a real user would
  • Survive internal refactors because they only depend on public behavior

Bad tests:

  • Import internal modules or query the DB directly instead of the product API
  • Assert on internal paths, logs, or implementation details
  • Over-mock so they no longer test real integration