On September 29, 2026, OpenAI released GPT-6.1 Sol, an upgrade to GPT-6 Sol that is available in Codex and ChatGPT Work for eligible paid plans. OpenAI positions the model for agentic coding, computer use and professional work, while noting that availability depends on a user’s plan, workspace settings and rollout access.

This is distinct from the earlier GPT-6 Sol and Luna availability update: GPT-6.1 Sol is a new model version, with its own API identifier (gpt-6.1-sol), pricing and reported evaluations. For QA teams, that means a new version to validate—not a safe assumption that every existing agent workflow will behave identically.

What OpenAI announced

  • GPT-6.1 Sol is available in Codex and ChatGPT Work for Plus, Pro, Business, Enterprise and Edu users, subject to rollout and workspace controls; it is not available in regular ChatGPT conversations.
  • OpenAI says it approaches GPT-6 Astra’s agentic-coding capability at about one-fifth of Astra’s standard input and output token price.
  • In OpenAI’s reported DeepSWE v1.1 evaluation, GPT-6.1 Sol matched Astra while exceeding GPT-6 Sol’s best score by 6.4 percentage points. Benchmarks are directional evidence, not a substitute for a team’s repository and test-suite results.
  • OpenAI reports lower failure rates than GPT-6 Sol in selected challenging safety evaluations, including transparency about broken search tools and avoiding unauthorized outcomes during agentic tasks. Those evaluations do not measure typical-use failure rates.

Why this matters for QA engineers

A more capable coding agent can change far more than the wording of its answer. It may choose different tools, make different edits, retry more or less often, interpret a browser state differently, or produce a new structured-output edge case. A model release can therefore affect test generation, flaky-test investigation, code review assistance and CI triage at once.

The upside is practical: lower-cost cached context and stronger reported coding performance may make longer investigations or repeated evidence collection more feasible. The control remains unchanged: keep test execution, assertions, build status and merge approval in deterministic systems. An agent can propose and gather evidence; it should not be the sole authority declaring a release safe.

A focused rollout checklist

  1. Pin a comparison lane. Run GPT-6 Sol and GPT-6.1 Sol on the same sanitized set of failed builds, flaky UI runs, API regressions and small test-maintenance tickets.
  2. Score outcomes, not fluency. Capture test-pass rate, unsupported diagnosis rate, tool-call failures, elapsed time, review findings and whether the suggested command reproduced the claim.
  3. Exercise controls deliberately. Test denied commands, restricted directories, unavailable tools, ambiguous ticket instructions and prompt-injection fixtures. Confirm the agent reports constraints rather than inventing a success path.
  4. Keep a human release gate. Require reviewable diffs and independently collected test results before merging agent-authored changes or closing a defect.
  5. Roll out by risk. Start with report preparation or low-risk triage. Expand to code or test changes only after the benchmark and control checks meet a pre-agreed threshold.

One useful acceptance criterion

For each agent recommendation, require:
1. observed evidence and source location
2. the smallest safe validation command
3. expected versus actual result
4. explicit uncertainty or blocked permissions
5. human approval before merge or release

This format makes a model upgrade measurable. If GPT-6.1 Sol improves the quality or speed of a QA workflow, the evidence should show it. If it merely produces a more persuasive narrative, the independent checks will expose that too.

Sources