On September 22, 2026, OpenAI introduced GPT-6 Sol and Luna in Codex and ChatGPT Work. Sol is positioned for complex coding and agentic workflows; Luna for focused, high-volume work. Both are separate from models in ordinary ChatGPT conversations, and availability depends on plan, workspace settings, and rollout access.
For QA engineers, this is less a reason to swap models everywhere than an opportunity to make model choice explicit in the test process.
What OpenAI confirmed
- GPT-6 Sol and GPT-6 Luna are available in Codex and ChatGPT Work.
- OpenAI describes Sol as built for complex coding and agentic workflows; Luna as an efficient option for focused, repeatable work at scale.
- The API model pages list a 1,050,000-token context window and tool support for both models when using the Responses API.
- OpenAI’s current model guidance recommends selecting Astra, Sol, or Luna according to task complexity, latency, and cost.
Why this matters for QA engineers
A test organization often has two very different AI workloads. One is a high-stakes investigation: isolate a flaky end-to-end failure, trace a regression across a large repository, or propose a safe test-plan change. The other is repeatable throughput: classify failure logs, draft test-data variants, normalize reports, or triage many similar issues. GPT-6 Sol and Luna give teams a reason to evaluate those lanes separately.
Do not assume that the faster or cheaper lane is automatically suitable for release decisions. Keep assertions, test execution, result collection, and merge approval in deterministic systems. Treat the model as an assistant that prepares evidence or proposes changes.
A practical evaluation plan
- Freeze a benchmark set. Include representative flaky failures, failed API tests, accessibility defects, and one large regression-planning task. Remove secrets and production data.
- Define evidence before prompting. Require a cited failure signature, affected files, proposed test command, and a confidence or uncertainty statement.
- Run a paired comparison. Give Sol and Luna the same versioned context and prompt. Record elapsed time, cost or usage, factual errors, unsupported claims, and whether the suggested validation actually passes.
- Score outcomes, not prose. A useful answer produces a reproducible command, a reviewable patch, or a correctly prioritized defect—not merely plausible-looking analysis.
- Gate rollout. Start Luna on low-risk classification or report-preparation tasks. Use Sol for deeper investigations only with human review and repository protections.
Try this evaluation prompt
You are assisting a QA engineer. Analyze the attached failed test run.
Return: (1) observed evidence, (2) ranked hypotheses,
(3) the smallest safe validation command, and (4) what you cannot conclude.
Do not mark the defect fixed or propose a merge without test results.
Run that prompt against a fixed set of sanitized artifacts. The same acceptance rubric is more valuable than a one-off subjective comparison.
Watch the integration details
For API-based tooling, OpenAI’s model guidance notes that Sol and Luna support none reasoning effort, while model behavior and available options can vary by surface. Confirm the exact model picker, workspace policy, tool permissions, usage limits, and data-handling configuration in your own environment before standardizing a workflow.
GPT-6 Sol and Luna in Codex can help QA teams match agent capability to the job. The durable practice is still the same: evaluate with real failures, preserve independent checks, and make a human accountable for release decisions.
