How I built an AI code review system with 9 agents — using only Markdown

No Python, no Node.js, no custom plugins. Just ~2,900 lines of prompt engineering and a fan-out/fan-in architecture that actually catches real bugs.

Mar 31, 2026~6 min read
How I built an AI code review system with 9 agents — using only Markdown

Code review is one of the most important — and most time-consuming — parts of any engineering workflow. AI tools can help, but a single "review this code" prompt almost always generates noise: false positives, style nitpicks, vague suggestions that waste more time than they save.

I decided to approach it differently. Instead of one universal prompt, I built a multi-agent system using Claude Code and a fan-out / fan-in architecture. Nine specialized agents analyze the same merge request in parallel, each through a different lens. The whole thing runs as a single slash command.

Here's how it works under the hood.

The architecture: fan-out / fan-in

The entire system is ~2,900 lines of Markdown instruction files. No Python, no Node.js, no build step — just prompt engineering. One command (/codereview) triggers the Orchestrator, which fans out to nine agents and fans back in with a structured report.

StepComponentRole
1/codereviewEntry point — triggers the whole pipeline
2OrchestratorFetches diff, routes files, spawns all agents in parallel
39 AgentsEach analyzes the MR through its own lens, simultaneously
4OrchestratorCollects results, deduplicates, filters, resolves conflicts
5MR ReportStructured comment posted back to the merge request

The Orchestrator — coordination, not review

The Orchestrator doesn't review code itself. Its job is to run the show:

Context gathering — fetches the diff and full file contents for every changed file in the merge request.

File routing — routes files to the right agents. A .tsx component goes to the React and TypeScript agents simultaneously. A .module.css file goes only to the Styles agent.

Aggregation — collects results from all nine agents and merges them into one report.

Filtering — this is where the real work happens. Deduplication, conflict resolution, and dropping low-quality findings before anything reaches the developer.

Nine specialized agents

Specialization is the core idea. A narrow agent is more accurate than a generalist because it's not trying to think about everything at once.

I also made a deliberate choice on model tiering — not every agent needs the most powerful model:

AgentModelDomain
React AgentSonnetComponent architecture, hook anti-patterns, state management
TypeScript StandardsHaikuType safety, naming conventions, import patterns
Security AgentSonnetXSS vectors, hardcoded secrets, unsafe URLs and iframes
Algorithm & PerformanceSonnetO(n²) complexity, memory leaks, unnecessary re-renders
Code ReuseSonnetDetects reimplemented logic that already exists in shared utilities
Styles & CSSHaikuDesign tokens, semantic HTML, accessibility (ARIA, alt text)
Storybook AgentHaikuVariant coverage, story typing, component registration
UX/UI Design SystemSonnetAtomic design hierarchy, Figma links, interactive state coverage

Haiku handles pattern matching — fast, cheap, accurate enough. Sonnet handles reasoning-heavy analysis — slower, but with real depth.

The secret to usefulness: Quality Gates

The biggest problem with AI code review is noise. Developers stop trusting a tool the moment it starts flagging missing spaces or suggesting "you might want to refactor this." So I built a strict Zero-Nitpicks policy.

Before any finding reaches the MR comment, it passes through three gates:

1. Existence check — the Orchestrator verifies that the code flagged by the agent actually exists in the diff. Agents can hallucinate changes that nobody touched.

2. Hard evidence — every finding must include an exact quote (1–3 lines) from the changed code. No evidence, no finding.

3. Language filter — findings containing hedging language ("might", "could potentially", "consider") or purely stylistic suggestions are automatically dropped.

The system also has a deduplication hierarchy. If the React Agent and the Security Agent both flag the same line, Security wins. There's a defined priority chain so conflicts are resolved deterministically, not randomly.

Resilience and telemetry

The system is designed defensively. If one agent fails, the other eight continue. The Orchestrator always delivers partial results — it never blocks the entire review because of one agent's error.

I also added telemetry that tracks the effectiveness of each agent over time. This tells me which prompts need refinement and where the system generates the most false positives.

What was actually hard

If I had to name the two places where I spent the most time — it was here.

Tuning the Quality Gates and severity levels was a constant iteration cycle. It's easy to say "filter the noise" — it's hard to define precisely where a nitpick ends and a real problem begins. Filters that were too strict missed important findings. Filters that were too loose let the noise back in. Finding that balance took many test cycles and a lot of manual result review.

Formatting the final output turned out to be just as important as the accuracy of the findings themselves. If a developer receives a wall of text with no structure, they either ignore it or lose time parsing it. I spent a long time making the report scannable at a glance — what's a merge blocker, what's a suggestion, which line it refers to, and why it matters.

A good report is one where a developer knows in ten seconds what needs to be fixed before merge.

What I learned

The architecture — fan-out/fan-in, model tiering, specialized agents — was actually the easy part. The hard part was answering a different question:

What does a developer actually want to see?

Technically correct and practically useful are not the same thing. Building this system taught me to think about the output as seriously as the logic that generates it.

If you're building something similar — spend at least as much time on your filtering layer as on your agents. The system is only as good as what it chooses not to say.

Was this helpful?

How I Built a Multi-Agent AI Code Review System in Markdown | Code Nomad