Coding

code-review

Try it

Multi-agent deep review for code PRs in any repo. Use when asked to "deep review this PR," "multi-agent review," "review

What it does

Multi-agent deep review for code PRs in any repo. Use when asked to "deep review this PR," "multi-agent review," "review

The skill document

Code Review — Multi-Agent PR Review

User-invokable orchestrator for code-PR review across any repo: Copilot for line-level findings, parallel subagents for the architectural and operational misses line-level review can't see. (Validation evidence and origin story: README.)

Fallbacks: no gh or not a GitHub repo → review the local diff (git diff ...HEAD) and skip the Copilot lane entirely; no subagent tool → run the review angles sequentially in the main thread. Detect what's available before assuming Copilot access.


Invocation: deliberately model-invocable — "review this PR" phrasing is the trigger. Pushes and issue-creation inside the flow are confirmation-gated.

Step 1: Assess PR Scope

The review target is $ARGUMENTS (a PR number, branch, or URL). If empty, infer it from the conversation or ask.

gh pr view N --json files,additions,deletions,title

Categorize:

TierExamplesTreatment
TrivialTypo, copy edit, single-line deps bump, README polishSkip multi-agent. Copilot alone is enough.
StandardComponent refactor, single-feature bug fix, contained logic changeCopilot + 1 subagent (pick the angle that matches the risk shape)
High-stakesPayment, auth, crypto/wallet, RPC, deployment configs, scale-relevant changes, new external SDK integrationsCopilot + 2 parallel subagents minimum (always include adversarial + operational)

If trivial → stop reading this skill, just request Copilot review and ship.


Step 2: Spawn Subagents in Parallel

Three angles that empirically work. Pick by risk shape:

  • Adversarial — "Try to break it. What's the worst attack?" Finds security holes, abuse vectors, replay/race conditions, malformed-input handling.
  • Operational — "What fails in production at scale? Latency, cost, observability, timeouts, env config, retries, cold-start?" Finds the architectural P0s line-level review misses (platform execution limits are the canonical class).
  • Reference-comparison — "Does this match the upstream reference impl exactly? Where does it diverge and why?" Use for third-party SDK / spec integration (Phantom deep-link, Stripe, OAuth providers, RPC clients).

Spawn all chosen subagents in a single tool-call batch so they run in parallel. Prompt templates live in patterns/subagent-prompts.md.

Every subagent prompt must specify:

  1. Repo path (absolute) + branch + commit hash at HEAD
  2. Files to focus on (paste the changed-files list from Step 1)
  3. Confidence rating + critical findings first
  4. Cap report length at ~500 words

Step 3: Request Copilot Review (in parallel with subagents)

gh pr edit N --add-reviewer copilot-pull-request-reviewer

Don't wait for Copilot to start before spawning subagents. They run in parallel.


Step 4: Triage Findings

Three buckets:

Real findings — fix now. Push fix to the PR branch. Pushing changes external state: unless the user already asked you to fix or babysit this PR, confirm before the first push of the session.

Stale re-flags — Copilot will sometimes re-flag findings against unmoved diff lines after fixes ship at HEAD. Always read the file at HEAD before re-fixing. If the fix is already there, drop a one-line reply on the comment ("Fixed at ") and move on. Do not re-fix.

Deferred (filed) — tests, refactors, or work outside scope. File as a GitHub issue (confirm with the user before creating it), link the issue number in the merge commit message, move on. Don't expand the PR.


Step 5: Iteration Cap

Maximum 2 Copilot review rounds. Empirically, by round 3 the signal-to-noise ratio collapses (~50% stale re-flags). Ship after round 2. If subagents flag genuinely new issues after round 2, those are issue-tracker items, not PR-blockers.


Step 6: When NOT to Use This Skill

  • Already-merged PRs → use Claude Code's built-in /security-review if you suspect a problem.
  • Trivial-tier PRs (per Step 1) → just request Copilot, skip the multi-agent overhead.
  • PRs where a paid multi-agent service was already run → don't double-spend.
  • Skill / prompt / instruction-file repos (where the "code" is markdown) → multi-agent is overkill; a single review pass against the trigger-quality and instruction-quality dimensions is enough.

Empirical Evidence

Validation case (5 rounds on a high-stakes PR):

  • Copilot caught ~10 real line-level issues (null checks, error handling, type mismatches).
  • Subagents caught what Copilot missed across all 5 rounds:
    • P0: Vercel maxDuration not set — function would time out under realistic load. Architectural miss invisible at the line level.
    • Persistent-replay vector in the auth flow.
    • Transient poll tolerance (no retry budget on flaky upstream).
    • Env-configurable timeout (hardcoded value).
    • Structured logging (string-concat logs, unparseable in production).
    • bodyParser size limit (DoS surface).
  • Round 3+ pattern: Copilot re-flagged the same already-fixed lines. Confirmed the round-2 cap.

Subagents are not redundant with Copilot. They cover a different layer.


Cross-references

  • patterns/subagent-prompts.md — paste-ready prompt templates for the three angles.

Related skills

Delegate coding, repository analysis, file edits, test runs, or code review to the local Codex CLI without embedding an OpenAI API key. This skill was created to work with ChatGPT/Codex enterprise accounts that may not have the same level of API access. Invokes codex exec with existing ChatGPT/Codex CLI authentication and returns Codex's final output.

6 installs1 stars

Conducts multi-axis code review. Use before merging any change. Use when reviewing code written by yourself, another agent, or a human. Use when you need to assess code quality across multiple dimensi Use when 需要Development领域自动化处理、数据分析和流程编排时使用。不适用于无明确需求的模糊场景。

Get senior-engineer-level code reviews with severity ratings, security checks, and ready-to-paste PR comments.

25 installs1 stars

Structured code reviews with severity-ranked findings and deep multi-agent mode. Use when performing a code review, auditing code quality, or critiquing PRs, MRs, or diffs. For the full multi-agent workflow, use the ia-review command (/ia-review in Claude Code).

35 installs

Review code for bugs, security, architecture, smells, patterns, performance, tests, and refactor plans

2 installs

Parallel code review — dispatches two subagents to audit **runtime safety** (resource leaks, null paths, race conditions) and **architecture consistency** simultaneously, then merges and deduplicates into a single report. One pass covers two orthogonal bug dimensions.

2 stars