No Python, no Node.js, no custom plugins. Just ~2,900 lines of prompt engineering and a fan-out/fan-in architecture that actually catches real bugs.

Code review is one of the most important — and most time-consuming — parts of any engineering workflow. AI tools can help, but a single "review this code" prompt almost always generates noise: false positives, style nitpicks, vague suggestions that waste more time than they save.
I decided to approach it differently. Instead of one universal prompt, I built a multi-agent system using Claude Code and a fan-out / fan-in architecture. Nine specialized agents analyze the same merge request in parallel, each through a different lens. The whole thing runs as a single slash command.
Here's how it works under the hood.
The entire system is ~2,900 lines of Markdown instruction files. No Python, no
Node.js, no build step — just prompt engineering. One command (/codereview)
triggers the Orchestrator, which fans out to nine agents and fans back in with a
structured report.
| Step | Component | Role |
|---|---|---|
| 1 | /codereview | Entry point — triggers the whole pipeline |
| 2 | Orchestrator | Fetches diff, routes files, spawns all agents in parallel |
| 3 | 9 Agents | Each analyzes the MR through its own lens, simultaneously |
| 4 | Orchestrator | Collects results, deduplicates, filters, resolves conflicts |
| 5 | MR Report | Structured comment posted back to the merge request |
The Orchestrator doesn't review code itself. Its job is to run the show:
Context gathering — fetches the diff and full file contents for every changed file in the merge request.
File routing — routes files to the right agents. A .tsx component goes to
the React and TypeScript agents simultaneously. A .module.css file goes only
to the Styles agent.
Aggregation — collects results from all nine agents and merges them into one report.
Filtering — this is where the real work happens. Deduplication, conflict resolution, and dropping low-quality findings before anything reaches the developer.
Specialization is the core idea. A narrow agent is more accurate than a generalist because it's not trying to think about everything at once.
I also made a deliberate choice on model tiering — not every agent needs the most powerful model:
| Agent | Model | Domain |
|---|---|---|
| React Agent | Sonnet | Component architecture, hook anti-patterns, state management |
| TypeScript Standards | Haiku | Type safety, naming conventions, import patterns |
| Security Agent | Sonnet | XSS vectors, hardcoded secrets, unsafe URLs and iframes |
| Algorithm & Performance | Sonnet | O(n²) complexity, memory leaks, unnecessary re-renders |
| Code Reuse | Sonnet | Detects reimplemented logic that already exists in shared utilities |
| Styles & CSS | Haiku | Design tokens, semantic HTML, accessibility (ARIA, alt text) |
| Storybook Agent | Haiku | Variant coverage, story typing, component registration |
| UX/UI Design System | Sonnet | Atomic design hierarchy, Figma links, interactive state coverage |
Haiku handles pattern matching — fast, cheap, accurate enough. Sonnet handles reasoning-heavy analysis — slower, but with real depth.
The biggest problem with AI code review is noise. Developers stop trusting a tool the moment it starts flagging missing spaces or suggesting "you might want to refactor this." So I built a strict Zero-Nitpicks policy.
Before any finding reaches the MR comment, it passes through three gates:
1. Existence check — the Orchestrator verifies that the code flagged by the agent actually exists in the diff. Agents can hallucinate changes that nobody touched.
2. Hard evidence — every finding must include an exact quote (1–3 lines) from the changed code. No evidence, no finding.
3. Language filter — findings containing hedging language ("might", "could potentially", "consider") or purely stylistic suggestions are automatically dropped.
The system also has a deduplication hierarchy. If the React Agent and the Security Agent both flag the same line, Security wins. There's a defined priority chain so conflicts are resolved deterministically, not randomly.
The system is designed defensively. If one agent fails, the other eight continue. The Orchestrator always delivers partial results — it never blocks the entire review because of one agent's error.
I also added telemetry that tracks the effectiveness of each agent over time. This tells me which prompts need refinement and where the system generates the most false positives.
If I had to name the two places where I spent the most time — it was here.
Tuning the Quality Gates and severity levels was a constant iteration cycle. It's easy to say "filter the noise" — it's hard to define precisely where a nitpick ends and a real problem begins. Filters that were too strict missed important findings. Filters that were too loose let the noise back in. Finding that balance took many test cycles and a lot of manual result review.
Formatting the final output turned out to be just as important as the accuracy of the findings themselves. If a developer receives a wall of text with no structure, they either ignore it or lose time parsing it. I spent a long time making the report scannable at a glance — what's a merge blocker, what's a suggestion, which line it refers to, and why it matters.
A good report is one where a developer knows in ten seconds what needs to be fixed before merge.
The architecture — fan-out/fan-in, model tiering, specialized agents — was actually the easy part. The hard part was answering a different question:
What does a developer actually want to see?
Technically correct and practically useful are not the same thing. Building this system taught me to think about the output as seriously as the logic that generates it.
If you're building something similar — spend at least as much time on your filtering layer as on your agents. The system is only as good as what it chooses not to say.
Was this helpful?