Claude Code structured outputs for QA can make automated failure triage easier to consume, but a valid JSON object is not the same as a correct diagnosis. The safest pattern is to treat the model response as one input to a deterministic evidence pipeline: constrain the shape, validate it again in your own code, compare every claim with test artifacts, and leave defect and release decisions with people.
This tutorial builds a disposable CI-triage lab around sanitized JUnit- or Playwright-style results. It uses Claude Code print mode for non-interactive execution and a JSON Schema for the response contract. Anthropic documents claude -p for scripted use, --output-format json for a JSON result envelope, and --json-schema for schema-validated structured output. The workflow below deliberately adds independent QA checks beyond those guarantees.
Why structured output needs a QA oracle
Free-form summaries are convenient for a human reader but fragile for automation. A heading might change, a confidence value might disappear, or an explanation might contain text that looks like a status. A schema removes much of that ambiguity by requiring named fields and defined types.
It does not prove semantic truth. A response can satisfy the schema while naming the wrong failed test, citing a nonexistent log line, or declaring a product defect when the runner actually lost its browser. Your test oracle therefore needs four separate gates: process completion, Claude result status, schema validity, and evidence-grounded meaning. Do not collapse them into one green check.
1. Create a private, reproducible triage lab
Start with a disposable repository and synthetic reports. Include at least one known application defect, one infrastructure failure, one flaky-test signature, and one clean control. Remove access tokens, customer data, internal hostnames, and production URLs. Add a harmless secret canary such as QA_CANARY_DO_NOT_REPEAT so you can detect unintended disclosure without exposing a real credential.
For every run, record the installed Claude Code or Agent SDK version, model selection, repository and branch, base commit, working directory, test-report hash, schema hash, prompt hash, operating system, locale, time zone, allowed tools, MCP servers, plugins, permission mode, maximum turns or budget, session ID, and CI run ID. In structured streaming, the initial system event can describe model and tool configuration; archive that metadata when available, but feature-detect fields rather than assuming every installation emits the same shape.
2. Define a narrow triage schema
Keep the contract focused on information your pipeline can verify. Avoid asking the model to reproduce entire logs. A useful first schema can require a run identifier, overall classification, a bounded list of findings, evidence references, confidence, and a human-review flag.
{
"type": "object",
"additionalProperties": false,
"required": ["run_id", "classification", "findings", "human_review"],
"properties": {
"run_id": {"type": "string"},
"classification": {
"type": "string",
"enum": ["product_defect", "test_issue", "infrastructure", "inconclusive", "clean"]
},
"findings": {
"type": "array",
"items": {
"type": "object",
"additionalProperties": false,
"required": ["test_id", "evidence_ids", "summary", "confidence"],
"properties": {
"test_id": {"type": "string"},
"evidence_ids": {"type": "array", "items": {"type": "string"}},
"summary": {"type": "string"},
"confidence": {"type": "number", "minimum": 0, "maximum": 1}
}
}
},
"human_review": {"type": "boolean"}
}
}
Anthropic’s structured-output documentation specifies JSON Schema draft-07 and supports common types, enums, required fields, and nesting. It also warns that the schema format keyword is an annotation rather than an enforced validator. If you need an ISO timestamp, URL, UUID, or domain-specific identifier, verify it in your own code. Keep the schema small enough to understand and version it alongside the harness.
3. Run Claude Code non-interactively
Use print mode so CI receives an answer instead of an interactive session. Supply the exact schema, request a JSON envelope, and tell Claude to use only the sanitized report and evidence IDs. A conceptual command looks like this:
claude -p "Triage report.json. Use only supplied evidence; return inconclusive when evidence is insufficient." \
--output-format json \
--json-schema "$(schema-content)"
Pass the schema using the quoting mechanism appropriate for your runner rather than copying the placeholder literally. Anthropic says print mode exits with zero on success and nonzero on failure. It also distinguishes startup errors, which are written to standard error, from failures that occur during execution and may appear in the result on standard output. Capture both channels separately and retain the raw response as restricted evidence.
For a cleaner scripted environment, Anthropic’s programmatic-usage guide documents a bare mode that skips automatic customizations. Confirm the option is supported by your installed build before relying on it. Isolation is still your responsibility: use a least-privilege CI identity, a disposable checkout, deny-by-default network policy, and explicit tool permissions.
4. Implement four independent acceptance gates
- Process gate: the command started, stayed within time and budget limits, and returned the expected exit status.
- Result gate: the JSON envelope represents a successful result rather than an execution error, cancellation, or retry exhaustion.
- Schema gate:
structured_outputexists and passes the pinned client-side validator using the same schema version. - Semantic gate: every test ID and evidence ID exists, quoted facts match the source artifact, classifications agree with the seeded oracle, and clean controls remain clean.
The third gate matters because Anthropic notes that a result can be reported as successful without a structured-output field. Treat missing structured output as failure. The Agent SDK can re-prompt when a response misses the schema; if retries are exhausted it can return error_max_structured_output_retries. Your harness should preserve that distinction instead of converting it into an empty finding list.
5. Build the negative test matrix
Test the contract before trusting it in a real pipeline. Start with invalid schemas, unsupported keywords, missing required fields, extra properties, wrong types, out-of-range confidence, unknown enum values, nulls, empty arrays, deeply nested objects, oversized reports, and date-like strings that violate your application validator even though the schema format annotation is accepted.
Then attack the meaning. Include ambiguous stack traces, duplicate test names, stale report fragments, a failure whose cause appears only in a linked fixture, injected instructions inside test output, a secret canary, misleading filenames, contradictory artifacts, truncated logs, a report with no failures, and an intentionally wrong expected-result label. Require inconclusive when the evidence cannot support a stronger claim.
Finally test operational faults: permission denial, unavailable tool, plugin or MCP failure, model fallback, maximum-turn or budget termination, process interruption, retry exhaustion, stdout and stderr interleaving, and success without structured_output. Run each fixture more than once to measure classification stability, evidence precision, unsupported-claim rate, latency, and cost.
6. Validate streaming behavior separately
If your CI needs progress events, Anthropic documents stream-json as newline-delimited JSON. Parse one complete line at a time; do not treat a partial network chunk as a complete object. Expect system, assistant, user, and final result events, and use the final result event as the authoritative completion record.
Add fixtures for split lines, blank lines, malformed events, duplicated events, missing final results, early termination, reordered transport delivery, and a final result followed by unexpected content. De-duplicate by stable event or run identity when available. Store only privacy-safe events, and never let an intermediate assistant message open a defect automatically.
7. Score the workflow like a QA system
Create a scorecard that separates syntax from usefulness. Track schema-valid rate, missing-output rate, seeded-oracle accuracy, evidence-reference precision, false-positive rate on clean controls, inconclusive accuracy, secret-canary leakage, deterministic rerun agreement, interruption recovery, latency, and cost per triaged run. A high schema-valid rate paired with weak oracle accuracy is a red signal, not a success.
Keep the output advisory. The model may propose a suspected owner, severity, or next test, but a human should reproduce product defects, review privacy-sensitive evidence, approve ticket creation, decide whether a flaky test is quarantined, and own the release gate. Deterministic tests and explicit policies—not model confidence—should control deployment.
Release checklist
- The report, prompt, schema, repository, model, tools, and environment are versioned or hashed.
- The CI identity has only the files, tools, and network access required for triage.
- Process, result, schema, and semantic gates are reported separately.
- Every finding links to real evidence and survives independent validation.
- Negative, injection, secret-canary, clean-control, interruption, and retry cases pass.
- Missing structured output, exhausted retries, and incomplete streams fail closed.
- A human retains authority over defects, quarantine, merge, deployment, and release.
Official references
Sources reviewed September 11, 2026. Command options and event fields can evolve, so verify them against the documentation and the exact Claude Code build used by your CI runner.
