Cursor Cloud Agent Hooks for QA give test teams a way to observe and control what an AI coding agent does during an automated test-maintenance task. Cursor’s official 3.11 changelog says cloud agents support hooks for prompts, responses, thinking, subagents, compaction, turn completion, tool execution, file changes, and shell work. That makes hooks useful for audit trails and policy checks, but a hook event is not proof that the application works.

This tutorial builds a practical, screenshot-friendly workflow for a QA team that lets a Cursor cloud agent investigate a failing Playwright test. The goal is to capture useful evidence, stop clearly unsafe operations, and preserve deterministic CI and human review as the release gates.

What Cursor officially supports

Cursor documents hooks as scripts that run at defined points in an agent workflow. The July 10, 2026 changelog specifically names beforeSubmitPrompt, afterAgentResponse, afterAgentThought, stop, and subagentStart, while also describing coverage around tools, files, shell commands, compaction, and completion. Use the current Hooks documentation as the source of truth for event names, payload fields, configuration, and supported behavior in your environment.

For QA, these events can answer operational questions: What task was submitted? Which files did the agent touch? Which test command ran? Did a subagent start? Did the workflow complete? They cannot answer the product question by themselves: Did the user journey behave correctly under the required conditions?

QA scenario: audit a flaky checkout-test investigation

Assume a cloud agent receives a task to investigate an intermittent checkout test. The repository contains Playwright tests, a staging configuration, and CI workflows. Your policy allows read-only inspection and focused test execution. It does not allow production access, secret reads, dependency publishing, force pushes, or automatic merge approval.

Define the expected outcome before configuring hooks:

  • a traceable task identifier;
  • a list of files read or changed;
  • the exact focused test commands and exit codes;
  • a clear record of blocked operations;
  • a diff that a person can review;
  • fresh Playwright artifacts from CI; and
  • a human decision to accept, revise, or reject the change.

Step 1: choose the minimum useful events

Do not log every possible event merely because it exists. Start with events tied to a real QA control. A prompt event can associate the task with a ticket. File and shell events can show which test assets changed and which commands ran. A subagent event can reveal delegation. A response or stop event can mark the end of the investigation.

Create an event-to-control map before writing scripts:

  • Prompt: require a ticket ID, target environment, and explicit stop conditions.
  • File activity: flag changes outside approved test and fixture folders.
  • Shell activity: allow focused test and lint commands; block destructive or publishing operations.
  • Subagent activity: record the purpose and prevent unexpected scope expansion.
  • Completion: require a summary of changed files, commands, results, and unresolved risks.

This mapping keeps hook configuration understandable during review. It also prevents a large stream of low-value telemetry from hiding the events that matter.

Step 2: define a privacy-safe audit record

Hook payloads can contain prompts, paths, command text, outputs, or other project context. Treat the audit stream as sensitive engineering data. Store only fields needed for the control, redact tokens and personal data, restrict access, encrypt transport and storage, and set a retention period.

A compact QA record can include timestamp, task ID, repository, branch, commit, event type, approved command category, affected path category, result, and policy decision. Avoid storing full prompts or full model reasoning by default. If a team has a legitimate need to retain detailed content, document that need and apply security, privacy, and legal review.

Step 3: add a narrow pre-action policy

Use pre-action hooks for high-confidence rules. A good rule is easy to explain and unlikely to block safe work accidentally. For this example, allow reads from the repository and focused tests against staging. Require approval for changes outside test folders. Block commands that publish packages, modify shared infrastructure, access production, or expose secrets.

Prefer structured checks over fragile substring matching. Normalize paths, resolve the repository root, validate command arguments, and use an allowlist for approved test runners. Fail safely when the hook cannot parse input, but test that behavior in a disposable branch before enabling it for the team.

Do not let a hook silently rewrite test intent. If the agent proposes removing an assertion, adding broad retries, skipping a scenario, or weakening a CI gate, the change should remain visible in the diff and require human review.

Step 4: capture post-action evidence

After a file edit or shell command, record the observable result: path category, command category, exit code, duration, and artifact location. For Playwright, preserve the report, trace, screenshot, video when appropriate, and the exact configuration used. For API tests, preserve the sanitized request conditions, assertions, response category, and test report.

A successful exit code only proves that the command reported success. Review whether the correct tests ran, whether assertions measure the business outcome, whether retries hid instability, and whether the environment matched the release target.

Step 5: test the hooks like production code

Create a small validation matrix before rollout:

  1. an approved read-only investigation is allowed;
  2. a focused test command is allowed and logged;
  3. a command targeting production is blocked;
  4. a path traversal attempt is rejected;
  5. a malformed payload follows the documented failure policy;
  6. a hook timeout does not create an unobserved unsafe action;
  7. secrets in sample input are redacted from stored records;
  8. duplicate or retried events do not corrupt the audit trail; and
  9. completion without test evidence is marked incomplete.

Run these cases in an isolated repository or branch. Review both false positives and false negatives. A policy that blocks normal QA work will be bypassed; a policy that records everything but enforces nothing may create false confidence.

Step 6: run one bounded pilot

Choose a low-risk test-maintenance task with a known expected result. Give the agent the ticket, current commit, staging-only boundary, permitted commands, required artifacts, and stop conditions. Watch the first run closely and compare the hook record with source control, terminal output, and CI artifacts.

For the flaky checkout test, the agent might identify a missing readiness assertion and propose waiting for the blocking overlay to disappear. Review the trace to confirm the overlay caused the failure. Inspect the diff to ensure the change waits for meaningful UI state instead of introducing a fixed delay or force-click. Then run the focused test repeatedly and the surrounding checkout suite in CI.

Step 7: make the human review package explicit

The final package should contain the task ID, starting commit, agent summary, files changed, policy decisions, exact commands, exit codes, artifact links, current diff, reproduced root cause, regression evidence, and unresolved risks. A reviewer should be able to verify the result without trusting the agent’s narrative.

Keep merge approval outside the hook workflow. The QA engineer or code owner decides whether the change preserves test intent, covers the real defect, follows framework conventions, protects data, and meets release criteria.

Screenshot plan for the tutorial

  1. Cursor Hooks documentation with relevant lifecycle categories visible.
  2. A simple event-to-QA-control table.
  3. The hook configuration open in the repository.
  4. A sanitized allowed focused-test event.
  5. A blocked production-targeting command with its policy reason.
  6. A file-change record limited to the test folder.
  7. Playwright trace and deterministic CI result.
  8. The final human review package beside the source diff.

Redact internal repository names, user data, tokens, URLs, prompts, and proprietary code before publishing screenshots.

QA checklist

  • Use only current event names from official Cursor documentation.
  • Map each event to a specific QA control.
  • Minimize, redact, protect, and expire audit data.
  • Test allow, block, malformed-input, timeout, and retry paths.
  • Keep the permitted environment and command set narrow.
  • Preserve source-controlled diffs and test artifacts.
  • Verify that the intended tests and assertions actually ran.
  • Treat agent responses and reasoning as context, not proof.
  • Require deterministic CI and human approval.
  • Review hook behavior after Cursor or repository workflow changes.

Common mistakes

  • Logging everything: excessive payload retention creates privacy and security risk.
  • Blocking by loose text matching: fragile checks create bypasses and false positives.
  • Trusting a success event: completion does not prove correct coverage or behavior.
  • Ignoring hook failures: timeouts and malformed payloads need an explicit safe policy.
  • Auto-approving test changes: AI can weaken assertions while making a suite appear greener.
  • Skipping independent evidence: hook logs do not replace traces, reports, diffs, or CI.

References

FAQ

What are Cursor Cloud Agent Hooks useful for in QA?

They can record and control defined workflow events such as prompts, tool actions, file changes, subagents, responses, and completion, helping teams audit automated test work.

Do hook logs prove that a test fix is correct?

No. They record workflow events. Correctness still requires meaningful assertions, reproducible artifacts, deterministic test execution, diff review, and human judgment.

Should QA teams store agent prompts and reasoning?

Not by default. Retain only what a documented control needs, redact sensitive content, restrict access, and apply a clear retention policy.

Can a hook automatically approve a test pull request?

It should not be the final authority. Hooks can enforce narrow policy and collect evidence, while code owners and QA reviewers retain merge and release approval.

Conclusion

Cursor Cloud Agent Hooks for QA are most valuable as a thin control and evidence layer around bounded automation. Start with a few events, connect each to a clear policy, protect the resulting data, test failure paths, and compare the audit record with real diffs and test artifacts. The dependable pattern is hooks for visibility, deterministic tests for evidence, and people for judgment.