OpenAI announced on September 29, 2026 that its Agents API now supports computer use. The update also brings Codex-style multi-agent coordination, tool search, programmatic tool calling and context compaction to developers building agents through the API.
The API is still in public beta. For QA and automation teams, the important change is scope: an agent can now combine browser-like interaction with tools, longer-running sessions and delegated sub-tasks. That makes outcome-based verification and action controls more important than a one-off happy-path demonstration.
What OpenAI announced
- Computer use in the Agents API: developers can build agents that interact with software through the API.
- Multi-agent capabilities: a main agent can delegate independent work to subagents, each with its own context.
- Tool orchestration: tool search, programmatic tool calling and support for MCP, custom functions and built-in tools are part of the platform’s agent workflow.
- Longer-running sessions: OpenAI says the platform can compact earlier context as a session approaches its context limit.
OpenAI’s DevDay recap says the computer-use capability is available through the API and in Codex and ChatGPT Work on Pro 500 and Enterprise. Availability, permissions and product-plan access should be confirmed in the environment a team actually intends to test.
Why this matters for QA engineers
Computer use changes an AI test assistant from a generator of suggestions into a system that can act across an interface. Multi-agent delegation increases the number of paths and hand-offs to verify. A credible QA pilot should therefore test the agent’s observable actions, permissions, evidence and recovery behavior—not only whether it reaches the desired screen once.
A compact pilot checklist
- Constrain the environment: start with a disposable test tenant, synthetic accounts and an allowlist of permitted applications and URLs.
- Trace every action: capture the agent session, tool calls, screenshots or UI observations, delegated-task outputs and the final evidence bundle.
- Test interruption paths: inject a changed selector, expired login, unexpected modal, denied permission and ambiguous page state. Verify the agent pauses or escalates instead of guessing.
- Validate subagent boundaries: give parallel tasks conflicting instructions and verify the coordinator preserves scope, avoids duplicate mutations and clearly identifies unfinished work.
- Measure outcomes: require a reproducible test result, a verifiable assertion and a clean environment after each run. Treat natural-language confidence as non-evidence.
Keep product claims and pilot evidence separate
OpenAI describes the Agents API as a managed way to use the Codex harness and supports different sandbox options. That does not by itself prove that a particular QA workflow is reliable, safe or suitable for production. Teams should establish their own acceptance thresholds for completion rate, unsafe-action rate, recovery, cost and human-review quality before widening access.
Sources
- OpenAI: DevDay 2026 Recap (September 29, 2026).
- OpenAI: Introducing the Agents API (public-beta platform details).
- OpenAI API documentation: Computer use (accessed October 5, 2026).
