September 15, 2026: A September 14 engineering case study details an Anthropic test selection redesign following a reported 25-fold increase in continuous integration (CI) jobs over six months. Anthropic connects that pressure to agent-assisted development and a growing test suite. See the official case study.
What changed
The deterministic service chooses tests using package relevance and past results. Its result-ingestion component fell behind, leaving selection decisions dependent on outdated history. Anthropic explicitly says this did not mean CI had skipped those pull requests or that untested code reached production.
The redesign moved state into an in-memory data store. Stateless workers append results to a journal; a separate consumer updates test histories. Anthropic reports higher running costs but easier scaling, with the service stable after tuning. This is an account of internal infrastructure work, rather than an announcement of a new customer-facing testing product.
Why this matters for QA engineers
QATechTools takeaway: Include the freshness of test-selection data in your CI health checks. A fast test run is less useful if the decision about which tests to execute relied on stale evidence.
For a pilot with coding agents, measure the time between a completed test and its result becoming available to the selector. Compare incoming and processed result counts during bursts. Check how quickly a newly added or repaired test becomes eligible to run, and sample selected suites against a fuller regression run to detect missed coverage.
These are suggested QA checks, not results from Anthropic’s study. Use your own workload and failure history to set thresholds before expanding agent-driven pull requests.
