Antigravity artifact review for QA gives test engineers a structured way to inspect what an AI agent plans, changes, and demonstrates. Google describes Antigravity artifacts as deliverables such as implementation plans, code diffs, diagrams, screenshots, and browser recordings. Those assets make agent work easier to review, but they are not automatically proof that a feature is correct.
This tutorial uses a staging-only checkout-test repair. The workflow checks scope before edits, reviews the resulting diff, compares the walkthrough with independent test results, and keeps the QA engineer responsible for release approval.
What Antigravity artifacts provide
Official Antigravity documentation says implementation plans explain proposed technical changes and are designed for user review. Reviewers can comment on plans before work proceeds. Walkthroughs summarize completed work and, for browser tasks, can include screenshots or recordings. Screenshot artifacts can also receive comments.
- Plan: Does the work match the ticket and risk?
- Diff: Did the agent change only intended files and preserve test intent?
- Walkthrough: Does the claimed result match the implementation?
- Independent evidence: Do fresh tests, logs, traces, and manual reproduction support the claim?
The final checkpoint matters most. A polished walkthrough can omit an edge case, use stale data, or demonstrate a path that differs from the release configuration.
QA scenario: repair a checkout regression test
Assume a checkout test intermittently fails after a UI change. Keep the task on a disposable branch and staging environment, with synthetic data. Define completion before starting: reproduce the failure, preserve the business assertion, change only approved test assets, produce a reviewable diff, run focused and surrounding regression tests, capture sanitized evidence, and list unresolved risks.
Step 1: request a reviewable implementation plan
Ask the agent to plan before modifying files. Include the ticket, starting commit, affected journey, permitted directories, staging boundary, approved commands, data rules, and stop conditions. Require expected files, assertions, validation commands, and evidence.
Goal: investigate the intermittent staging checkout failure.
Before editing, create a plan listing:
- evidence needed to reproduce the failure
- files likely to change and why
- assertions that must remain intact
- focused and regression test commands
- screenshot or trace evidence
- risks, assumptions, and stop conditions
Do not access production, expose secrets, skip tests,
weaken assertions, or edit outside approved folders.
Step 2: comment on scope and test risk
Read the plan against the requirement and existing suite. Comment where scope is broad, an assumption lacks evidence, or validation is weak. Request missing negative paths, forbid fixed delays, require existing selectors to be checked, and add the surrounding payment regression pack where needed.
Do not approve a plan merely because it looks detailed. A test can become stable for the wrong reason if an agent removes an assertion, catches an exception, adds broad retries, or replaces a meaningful readiness check with a long pause.
Step 3: review the complete code diff
Inspect the source-controlled diff rather than trusting the walkthrough summary. Check every changed file, dependency update, fixture, configuration value, selector, assertion, retry, timeout, and cleanup path.
- Reject unrelated production-code changes unless explicitly in scope.
- Confirm selectors express stable user meaning.
- Keep assertions tied to checkout completion, not simple visibility.
- Verify the fix cannot hide a real product defect.
- Ensure secrets, internal URLs, and customer data are absent.
- Explain every deviation from the approved plan.
Step 4: examine walkthroughs and visual artifacts
Use the walkthrough to locate the demonstrated state quickly. Check environment, account state, test data, viewport, timestamps, and path. A screenshot is strong evidence for one visible state but weak evidence for the sequence around it. A recording improves sequence visibility but can still miss network failures, accessibility defects, data corruption, or another browser.
Comment on ambiguous artifacts and request a clearer capture. Before sharing, crop or redact account names, tokens, internal domains, source code, personal data, and unrelated tabs.
Step 5: run deterministic checks independently
Run the focused test from a clean checkout and then execute the relevant regression pack in CI. Preserve the commit SHA, command, configuration, exit code, report, trace, screenshots, and logs. Repeat the intermittent scenario enough times to challenge the proposed fix using your team’s risk-based threshold.
Manually reproduce the user journey when the defect is visible. If your result differs from the walkthrough, record the environment and data differences and keep the task open. Agent artifacts explain what the agent observed; independent execution determines whether it is reproducible.
Step 6: build a review package
- Ticket and acceptance criteria
- Starting and ending commit identifiers
- Approved plan and QA comments
- Final diff and deviations
- Walkthrough, screenshots, and recording links
- Focused and regression test results
- Manual reproduction notes
- Known gaps, redactions, and rollback plan
- Named human approver and decision
Choose artifact review settings deliberately
Antigravity review policies influence when the agent requests review. Choose approval boundaries that match risk. Do not use an always-proceed approach simply to save clicks when a task can modify tests, run privileged commands, or touch sensitive environments. Test rejection, requested changes, comments, and missing-review behavior in a disposable project.
Screenshot plan
- Implementation plan with scope, tests, and stop conditions.
- Inline QA comment requesting stronger negative-path coverage.
- Final code diff beside the approved plan.
- Sanitized walkthrough with browser screenshot evidence.
- Independent CI report and trace for the same commit.
- Final review package with a human approval field.
QA checklist
- Use staging and synthetic test data.
- Record the starting commit and task boundary.
- Review the implementation plan before edits.
- Challenge assumptions, missing risks, and weak validation.
- Inspect the complete diff independently.
- Verify visual artifacts show the intended environment.
- Run focused and regression checks from a clean state.
- Preserve reports, traces, logs, and commit identifiers.
- Redact sensitive content before sharing.
- Keep merge and release approval with accountable humans.
Common mistakes
- Approving a detailed-looking plan: formatting does not guarantee correct scope.
- Trusting the walkthrough: compare it with the diff and independent results.
- Treating screenshots as full proof: they cover only the captured state.
- Skipping clean reruns: cached state can mislead.
- Sharing raw artifacts: recordings may reveal sensitive data.
- Automating final approval: artifact quality does not remove accountability.
References
- Antigravity Artifacts overview
- Antigravity Implementation Plan
- Antigravity Walkthrough
- Antigravity Screenshots
- Antigravity Artifact Review
FAQ
Are Antigravity artifacts test evidence?
They are useful agent-produced review evidence. Pair them with fresh deterministic tests, diffs, logs, traces, and independent reproduction.
Should QA approve the plan before edits?
For meaningful changes, plan review catches scope errors, weak assertions, missing risks, and unsafe environment assumptions.
Can a walkthrough replace CI?
No. A walkthrough summarizes agent work; CI provides repeatable execution evidence for a known commit and configuration.
What should be redacted?
Remove tokens, account details, personal data, internal domains, proprietary code, unrelated tabs, and restricted information.
Conclusion
Antigravity artifact review for QA works best as evidence gates: review the plan, challenge assumptions, inspect the diff, examine the walkthrough, rerun deterministic tests, and document a human decision. Artifacts improve traceability; independent testing and accountable review keep the release decision defensible.
