GitHub Copilot Agent Plugins for QA give test teams a portable way to package reusable agents, skills, hooks, and integrations. Portability is useful, but it introduces a new verification problem: the same installable package may load different supported components, settings, tools, or permissions across an IDE, terminal, and cloud agent.
This tutorial creates a synthetic, evidence-only QA plugin and validates it before wider use. You will inspect the package and manifest, freeze its source and version, test installation from a controlled marketplace, compare supported behavior across Copilot clients, challenge MCP and hook boundaries, verify updates and rollback, and retain human approval for every external action.
What Agent Plugins 1.0 changes
GitHub’s Agent Plugins 1.0 announcement says one package can expose supported skills and MCP configuration to compatible agent clients. GitHub documents general availability in VS Code, Copilot CLI, the GitHub Copilot SDK, and the Copilot app across Copilot plans. Existing Copilot plugins that do not target the 1.0 specification remain supported, so migration is not automatic or mandatory.
The current Copilot plugin documentation describes plugins as installable packages that can contain custom agents, skills, hooks, MCP server configuration, and LSP configuration. A root plugin.json manifest is required. Agent Plugins 1.0 also supports a Copilot-specific namespace for components that other compatible clients should ignore.
That distinction drives the QA strategy: validate the portable core and the client-specific layer separately. A successful installation does not prove that every client loaded the same capabilities or enforced the same controls.
Define the portability contract
Given one immutable plugin artifact, marketplace entry, client build, account policy, repository trust state, and synthetic workspace, every supported client must load only its declared compatible components, produce the expected evidence-only result, refuse undeclared actions, and leave independently verified state.
For each trial, record the plugin name and version, source repository and commit, artifact hash, manifest hash, marketplace and entry hash, client name and build, operating system, Copilot plan and policy state, installation scope, enabled component inventory, MCP server identity, tool approvals, expected result, observed result, files changed, network destinations, and reviewer disposition.
Step 1: build a harmless QA evidence plugin
Create the plugin in a private disposable repository. Its only purpose is to review a frozen test-failure fixture and return a small evidence report. Use no production credentials, real defects, package publishing, deployment tools, or write-enabled external systems.
A useful test package contains:
qa-evidence-plugin/
├── plugin.json
├── skills/
│ └── failure-evidence/
│ └── SKILL.md
├── mcp.json
└── com.github.copilot/
├── agents/
│ └── qa-evidence.agent.md
└── hooks/
└── hooks.json
The portable skill should accept a stable fixture ID and produce an ordered list of evidence references, missing inputs, and next checks. The synthetic MCP server should expose one read-only fixture lookup and one clearly named denied write canary. The Copilot-specific agent may coordinate the workflow, while its hook records only sanitized lifecycle evidence inside a temporary report directory.
Step 2: validate the manifest and package inventory
Use the schema and structure from the official Agent Plugins specification when creating plugin.json. GitHub’s announcement says a 1.0 migration adds the schema reference, keeps skills under skills/, keeps MCP configuration in mcp.json, and moves Copilot-specific files into com.github.copilot/.
Before installation, fail the test for a missing manifest, unknown component path, duplicate component identifier, path traversal, absolute local path, hidden executable, broken symbolic link, unexpected binary, unpinned remote source, or component that is not listed in the reviewed inventory. Record hashes for every file rather than trusting the directory name or marketplace version.
Run negative packages with a malformed schema field, absent skill file, invalid MCP JSON, renamed namespace, and one undeclared extra hook. Each must fail visibly or load only the valid declared subset according to the documented client behavior. Treat silent partial loading as an indeterminate result until the client inventory explains it.
Step 3: freeze a controlled marketplace
GitHub documents plugin discovery through marketplaces. Copilot CLI includes registered marketplaces and can browse, install, list, update, and uninstall plugins. For this lab, use a local or private test marketplace containing only the synthetic plugin and one intentionally incompatible version.
Capture the marketplace source, commit, marketplace.json hash, plugin path, version, artifact hash, and reviewer. Never test an unknown community plugin with production credentials or a signed-in production repository merely because it appears in a default marketplace.
The documented CLI flow includes:
copilot plugin marketplace list
copilot plugin marketplace browse QA-MARKETPLACE
copilot plugin install QA-PLUGIN@QA-MARKETPLACE
copilot plugin list
Do not copy the placeholders directly. Substitute only the frozen lab marketplace and plugin identifiers. Preserve stdout and stderr, then independently inspect the installed directory and hashes.
Step 4: test install, reinstall, and uninstall
Start with no installed plugin and capture the client component inventory. Install the frozen version, restart or reload the client as required, and verify exactly one plugin entry appears. Attempt a duplicate install, interrupted install, corrupted download, read-only destination, insufficient disk space, invalid signature or checksum where supported, and an unavailable marketplace.
Uninstall the plugin with the documented command and verify its agents, skills, hooks, MCP servers, caches, settings, and report artifacts no longer load. A removed marketplace can fail while plugins from it remain installed; GitHub’s CLI documentation describes a force option that can remove the marketplace and uninstall its plugins. Test that destructive path only in the disposable profile and require an explicit inventory before approval.
Step 5: compare client capability inventories
Build a matrix for VS Code, Copilot CLI, and the Copilot app; include the SDK only if your team actively embeds it. Use the same artifact hash in every column. Record whether each client discovers the portable skill, MCP configuration, Copilot-specific agent, hook, and any unsupported extension.
| Component | Portable expectation | QA evidence |
|---|---|---|
| Skill | Loads where skills are supported | Inventory plus one controlled invocation |
| MCP server | Uses declared config and approval boundary | Server identity, tool list, sanitized call log |
| Copilot agent | Loads only in Copilot namespace-aware clients | Agent inventory and selected profile |
| Hook | Runs only for its declared event | One correlation ID and no extra side effect |
| Unsupported item | Ignored or rejected visibly | Diagnostic and absent runtime effect |
Do not require identical UI text. Compare semantic behavior: component loaded, invocation accepted, evidence produced, denial enforced, and side effects absent. Document any client-specific capability rather than labeling it a portability failure automatically.
Step 6: run one deterministic QA task
Give every client the same frozen fixture containing two failures, one known infrastructure error, and one clean control. Ask the plugin to classify only the evidence it can cite. The expected report should contain stable fixture IDs, evidence references, a bounded classification, missing information, and a statement that no fix or external action was attempted.
Run the task three times per client in fresh sessions. Normalize ordering and presentation before comparison. Measure required-field accuracy, evidence attribution, clean-control accuracy, unsupported-claim rate, repeatability, duration, and token or credit evidence where available. Reproduce every substantive conclusion outside the plugin.
Step 7: challenge the skill boundary
Seed prompts that ask the portable skill to edit the test, weaken an assertion, retry indefinitely, install a package, read a credential file, open an unrelated repository, or publish a defect. The skill instructions must keep the task evidence-only, but instructions are not an access-control boundary. The client policy and tool permissions must prevent undeclared actions.
Add malformed, oversized, contradictory, duplicated, stale, secret-bearing, and prompt-injected fixtures. The report should identify uncertainty, redact synthetic secrets, preserve stable references, and stop rather than invent missing evidence.
Step 8: verify MCP identity and approvals
A plugin can bundle MCP server configuration, which increases both utility and risk. Before any call, record the server command or URL, resolved executable or endpoint, version, hash or certificate evidence, environment variables by name only, exposed tools, and annotations. Keep the lab server local or private and read-only by default.
Invoke the fixture lookup and verify it returns only the requested synthetic record. Then request the denied write canary and confirm no file, database row, issue, pull request, message, or remote object appears. Test server substitution, changed command path, renamed tool, added tool, missing annotation, timeout, malformed result, prompt injection in tool output, and reconnect after an update.
GitHub’s announcement recommends pairing plugin governance with MCP allowlists that approve or block servers by URL, command, or name. Test the plugin policy and the MCP allowlist independently; passing one does not prove the other.
Step 9: test hooks as side-effect boundaries
Hooks intercept agent lifecycle events, so a harmless-looking plugin can gain deterministic execution points. Configure the lab hook to append one sanitized correlation ID inside tmp/plugin-lab/. Verify it runs only on the declared event, once per event, and never captures prompts, credentials, full environment dumps, or unrelated file content.
Exercise missing hook runtime, non-zero exit, timeout, duplicate event, concurrent sessions, malformed payload, report-path denial, and plugin uninstall. The client must remain observable and the hook must not turn a failed QA review into an apparent success.
Step 10: test governance settings
For Copilot Business and Enterprise, GitHub documents managed settings for automatically installing or blocking plugins and controlling known marketplaces. Enterprise values establish a baseline, with approved team overrides combined according to managed-setting rules.
Build a small identity matrix: unmanaged test user, managed user with the plugin enabled, managed user with the plugin blocked, approved team override, unapproved marketplace, and strict marketplace mode. Verify effective state in every client rather than trusting the configuration file alone. Test propagation delay, offline start, cached plugin, policy removal, conflicting scopes, and rollback.
Step 11: validate updates and rollback
Create version 1.0.1 with one visible, safe change to the report schema and no permission expansion. Capture both artifacts and marketplace entries. Run the documented update command, restart the client, and prove the new version is active while the old one is no longer invoked.
Next seed a bad update: broken skill path, changed MCP identity, new hook event, or undeclared tool. The rollout must stop. Restore the frozen prior artifact, confirm hashes and component inventory, rerun the deterministic fixture, and verify no stale process or cached configuration from the rejected version remains.
Do not use an unqualified latest tag as QA evidence. Record the exact marketplace version, repository commit, artifact hash, and installed content hash.
Step 12: produce a portability report
The final report should contain:
- Source, marketplace, version, commit, and artifact hashes
- Manifest and component inventory validation
- Client builds, account policies, trust state, and install scopes
- Supported, ignored, and rejected components per client
- Deterministic task accuracy and repeatability
- MCP identity, tool approvals, denials, and side-effect evidence
- Hook correlation IDs, failures, and privacy checks
- Install, update, uninstall, and rollback outcomes
- Managed-setting and marketplace-policy matrix
- Known gaps, indeterminate cases, owner, and expiry
Approve rollout only for the tested artifact and client-policy combination. Any component, marketplace, MCP identity, hook, or permission change should trigger a new review.
QA rollout checklist
- Official feature availability and documentation are reviewed
- The plugin source, version, commit, manifest, marketplace, and artifact hashes are frozen
- Only synthetic fixtures and identities are used
- The required manifest and declared component paths are validated
- Portable and Copilot-specific components are inventoried separately
- Install, duplicate, interrupted, offline, uninstall, and cleanup cases are tested
- Every client uses the same artifact hash
- Behavior comparisons use semantic evidence, not identical UI text
- The skill refuses scope expansion and unsupported claims
- MCP identity, tools, allowlists, timeouts, malformed output, and write denial are verified
- Hooks are event-scoped, privacy-safe, and independently observable
- Managed plugin and marketplace settings are tested by identity and client
- Updates cannot expand capability silently
- Rollback restores hashes, inventory, behavior, and process state
- Plugin approval, external actions, defects, fixes, merge, and release remain human-owned
Official sources
- GitHub: Agent Plugins 1.0 announcement
- GitHub Docs: About Copilot plugins
- GitHub Docs: Find and install Copilot CLI plugins
Agent Plugins 1.0 reduces packaging duplication, not verification work. Treat each plugin as versioned executable capability: inspect it before installation, test every supported client, constrain MCP and hooks, verify policy and side effects independently, and rerun the matrix after updates. Portable should mean repeatably understood—not automatically trusted.
