Codex Computer Use for QA can help a tester exercise real user journeys, record what fails, and turn observations into a structured triage report. The useful pattern is not to say “test everything.” It is to define a controlled environment, select a small set of high-risk flows, specify evidence requirements, and independently reproduce every reported defect.
This tutorial shows a practical pre-release pass for QA engineers, SDETs, automation testers, and AI testing learners. The example covers signup, inviting a teammate, and changing a subscription in staging. Adapt the same structure to your application.
What Codex Computer Use can do in a QA pass
OpenAI’s official QA workflow describes Computer Use as able to see an interface, click through flows, type into fields, and record failures. It recommends naming the environment and important flows, then requesting a report with reproduction steps, expected behavior, actual behavior, severity, and a triage summary.
That makes it useful for a bounded exploratory pass. It does not make the output a test result by itself. Treat each finding as a hypothesis until a human tester or deterministic check reproduces it on the same build and data state.
Before you start: create a safe test boundary
Use a staging or local environment with synthetic accounts. Do not point an exploratory agent at production billing, real customer data, destructive admin controls, or irreversible workflows. Prepare:
- The exact build identifier or commit SHA.
- A dedicated test account and known starting state.
- Synthetic test data that can be reset.
- Relevant feature flags and browser details.
- Three priority journeys with clear success criteria.
- A rule for actions that require human approval.
For the sample pass, the release candidate is RC-142. The test account has no team members and uses a sandbox payment provider. The agent may navigate and enter synthetic data, but it must stop before any action that creates a real charge, deletes shared data, or changes account ownership.
Step 1: write a compact risk-based test charter
A charter keeps the run focused and screenshot-friendly. Write one sentence per journey and list the observable result:
- Signup: create an account, verify the email-confirmation state, and reach the dashboard.
- Invite: invite a teammate, confirm success feedback, and verify the member appears as pending.
- Plan change: select a higher plan in the sandbox checkout, confirm the price summary, and stop before final submission unless explicitly approved.
Add negative checks that matter to the release: invalid email validation, double-submit protection, expired-session handling, keyboard focus, and clear error recovery. Keep the scope small enough that every observation can be reviewed.
Step 2: give Codex the environment and stop conditions
Use a prompt with explicit setup, flow, evidence, and safety sections:
Test build RC-142 in the staging environment using the dedicated synthetic account.
Cover these journeys:
1. Signup and email-confirmation state
2. Invite one teammate
3. Review a sandbox plan upgrade
For every issue, record:
- concise title
- exact reproduction steps
- expected result
- actual result
- severity with one-sentence rationale
- page and relevant test-data state
- screenshot checkpoint
Continue past non-blocking issues. Stop before a real charge, deletion, ownership change, or any action outside staging. End with a short triage summary. Do not modify application code.
If your repository already contains a test-plan file, point Codex to it so the pass follows existing terminology and acceptance criteria. Remove secrets and customer data before sharing context.
Step 3: observe the run and capture checkpoints
Do not wait only for the final summary. Review the session at meaningful checkpoints:
- Initial account state and visible build marker.
- Input data before submission.
- The exact screen where behavior diverges.
- Error messages, disabled controls, and recovery paths.
- Final state after each journey.
A screenshot should establish context, not merely show a red message. Pair it with the page URL or route, build ID, account state, and the action immediately before the failure. Redact tokens, personal data, and payment details.
Step 4: normalize findings into defect candidates
Review the generated report and split combined observations. One defect candidate should describe one behavior. Rewrite vague statements such as “invite is broken” into an observable claim such as “Invite dialog remains open after a successful 201 response and shows no confirmation message.”
Use a simple evidence table:
| Field | Required evidence |
|---|---|
| Build | Release candidate or commit |
| Precondition | Account, data, flag, and browser state |
| Action | Minimal reproducible sequence |
| Oracle | Requirement, acceptance criterion, or established behavior |
| Observation | Actual UI, response, log, or screenshot |
| Impact | User and release risk |
Step 5: independently reproduce every finding
This is the most important control. Start from the documented precondition and rerun the steps manually or with an existing deterministic test. Repeat at least once after resetting data. If the result changes, investigate timing, test pollution, network behavior, or environment drift before filing it as a product defect.
Check whether the expected result comes from a real requirement. An AI agent can misread copy, overlook a feature flag, or infer an oracle that the product never promised. A screenshot proves what appeared; it does not prove what should have appeared.
For confirmed defects, attach the smallest useful evidence bundle. For unconfirmed observations, label them clearly as needing investigation rather than inflating severity.
Step 6: add a deterministic regression check
When a confirmed issue is stable and valuable to guard, add or update an automated test. Prefer observable behavior: a confirmation message, resulting record, navigation state, or API response. Use framework-native waiting and retrying assertions instead of fixed delays. Keep the exploratory session transcript as supporting context, not as the regression oracle.
Run the focused test first, then the relevant suite. Record the exact command, exit status, report link, and build. A passing rerun on one path does not replace broader accessibility, security, performance, compatibility, or integration coverage.
Step 7: finish with human release triage
Rank confirmed defects by user impact, reach, recoverability, and release risk. Separate blockers from non-blocking polish. The QA owner should decide whether evidence is sufficient, whether more environments are required, and whether the release gate changes.
A concise final summary can include: flows attempted, flows completed, confirmed defects by severity, unconfirmed observations, blocked steps, data cleanup status, and recommended next action.
Common mistakes to avoid
- Broad prompts: “Test my app” creates inconsistent coverage and weak evidence.
- Unsafe environments: never allow uncontrolled destructive or financial actions.
- Missing setup: account state, flags, data, and build version change the result.
- AI-defined severity: use team impact definitions and human triage.
- No independent rerun: an agent observation is not automatically a confirmed defect.
- Replacing regression suites: exploratory Computer Use complements deterministic automation.
Practical QA checklist
- Use staging or local with synthetic data.
- Record build, browser, account state, and feature flags.
- Select three risk-based journeys.
- Define stop conditions and approval boundaries.
- Request repro, expected, actual, severity, and screenshots.
- Redact sensitive evidence.
- Reproduce every finding independently.
- Validate the oracle against requirements.
- Add deterministic coverage for confirmed regressions.
- Keep final severity and release decisions with human QA.
Conclusion
Codex Computer Use for QA is most valuable as a structured exploratory assistant. Give it a safe environment, a narrow charter, realistic test state, explicit stop conditions, and a strict evidence format. Then apply the discipline that makes QA trustworthy: reproduce the behavior, verify the oracle, run deterministic checks, and make the release decision from reviewed evidence.
Sources
