GitHub Copilot dynamic workflows for QA are now available in public preview in Copilot CLI, the GitHub Copilot app, and the Copilot SDK. GitHub announced the feature on October 1, 2026. Instead of asking an agent to improvise a process each time, teams can define the stages, conditions, handoffs, and checks in code while leaving analysis and judgment to one or more agents.

For QA engineers, the appeal is repeatability: a release gate can collect build results, run deterministic checks, assign focused analysis to agents, require structured evidence, and pause for a human decision before a rollout continues. The feature remains a public preview, so it deserves a limited-scope pilot rather than immediate use as the only deployment control.

What GitHub announced

A dynamic workflow is a program inside a Copilot extension. It can run commands, tools, or services; pass results between stages; run independent tasks in parallel; ask agents to verify each other’s findings; and stop at a checkpoint for review. The workflow author defines the process in code, whereas Copilot autopilot and /fleet decide their plan for each task.

GitHub lists release checks, parallel pull-request review, and investigations that report a result only when two models agree among the intended use cases. In Copilot CLI, experimental features must be enabled with --experimental or /experimental on. GitHub says dynamic workflows are available on all Copilot plans, subject to preview changes.

Why this matters for QA engineers

AI-assisted checks are more useful when their orchestration is inspectable. A workflow can preserve the same required stages for every run while allowing agents to interpret logs, summarize test failures, or compare evidence against acceptance criteria. That gives teams a clearer boundary between deterministic automation and probabilistic analysis.

  • Repeatable gates: keep build, test, and evidence-collection steps consistent across releases.
  • Parallel triage: assign independent agents to logs, changed files, or test failures, then combine structured findings.
  • Human checkpoints: pause before a risky action and require a reviewer to inspect the evidence.
  • Controlled spend: set limits for concurrent agents, total agents, runtime, and approximate AI-credit usage.

A practical QA pilot

Start with a read-only release-readiness workflow in a non-production repository. Have it collect CI status and recent test artifacts, run the existing deterministic suite, ask one agent to classify failures, ask a second agent to challenge that classification, and emit a fixed JSON evidence bundle. Add a checkpoint that requires a human tester to approve any conclusion before the workflow creates an issue, changes a label, or starts another system.

Scope: current pull request only
Run: existing unit and UI regression checks
Agents: one failure triage + one independent verifier
Output: structured evidence with links to logs
Stop: before any write action
Limits: 2 concurrent agents, 10 total, 20 minutes

Measure whether the result is complete, correctly structured, traceable to the source artifacts, and stable across reruns. Deliberately test missing logs, flaky tests, contradictory agent conclusions, timeouts, and a rejected checkpoint. A passing workflow should not be treated as a passing product unless the underlying test evidence supports it.

Permissions and limits need explicit testing

GitHub notes that subagents inherit the permission grants of the session that started them. A new permission request can surface in an interactive CLI session, and a session-wide grant can apply to other subagents that need that permission. The non-interactive copilot workflow run command does not show approval prompts, so permissions must be granted ahead of time; requests that cannot be approved automatically are denied.

  • Use least-privilege tool and URL permissions for the pilot.
  • Set an AI-credit ceiling and inspect actual usage after a small run before scaling up.
  • Keep write actions behind a checkpoint until reliability is demonstrated.
  • Version the workflow definition and test its changes like production automation.
  • Store returned evidence separately from the agent narrative so reviewers can audit it.

Dynamic workflows give QA teams a way to turn a useful agent pattern into a reviewable process. The best first use is not an autonomous release decision: it is a narrow, evidence-producing workflow with deterministic checks, bounded permissions, and a visible human approval step.

Sources