AI code isn't dangerous at random

I almost shipped a vulnerability to production because the model handed it to me with full confidence. Then I checked the numbers — the 2026 ones, not the two-year-old ones. Models hallucinate less, but picking a model stopped helping, and the attack surface is now shared by all of them.

Jul 3, 2026~7 min read
AI code isn't dangerous at random

I was writing a component to render posts on code-nomad.com and asked AI for help. The first thing I got back was a component with dangerouslySetInnerHTML.

tsx
function Post({ html }) {
  return <div dangerouslySetInnerHTML={{ __html: html }} />;
}

I told the model this opens the door to XSS. If the content isn't sanitized, someone can inject an arbitrary script. The model replied: you're right, my mistake, let me change the approach.

And that's the whole problem right there. I caught it because I happened to know to look for it. If I'd been in a hurry, if I hadn't known this particular trap, I would have shipped that code with a clear conscience. It would have worked. The bomb would have ticked quietly until someone eventually set it off.

I've written before that the hardest AI work is the kind that never shows up in any metric. This post goes one step further and asks the question I left open back then: if I caught dangerouslySetInnerHTML because I happened to know about it, how many things am I letting through because I don't know they exist at all?

I checked the numbers

Let me start with a baseline, because without it you can't see what changed. A USENIX Security study, run on models released between mid-2023 and mid-2024, found that 19.7 percent of packages recommended by AI didn't exist in any registry. Almost one in five. (16 models, 576,000 code samples.)

Sounds like a different world, and rightly so — that was two years ago. Security data always has this problem: by the time someone collects and publishes it properly, a year has passed and it describes outdated models. Which is why what matters more is what came out when someone repeated the same measurement on fresh hardware.

In May 2026, that methodology was replicated on five models from late 2025 and early 2026: Claude Sonnet 4.6, Claude Haiku 4.5, GPT-5.4-mini, Gemini 2.5 Pro, DeepSeek V3.2. Package name hallucinations dropped to a range of 4.62 to 6.10 percent. Roughly one in twenty. The models made enormous progress and they deserve credit for it. But "less often" is not "never", and at the scale we generate code today, one in twenty is still a lot.

This isn't "the model gets it wrong sometimes". This is a systemic property of a tool I use every day, and it doesn't disappear with the next version — it just shrinks.

Picking a model stopped helping

I have a soft spot for thinking about this through tiering: match the model to the task and half your problem is solved. With package hallucinations, that no longer works, and you can see it straight in the numbers.

In 2024, the spread between models was a chasm. Commercial models hallucinated packages at 5.2 percent, open-source models run locally at 21.7. A fourfold difference. Back then, model choice genuinely changed the risk — reaching for a better model was a security decision.

In 2026, that chasm closed. From 4.62 percent (Claude Haiku 4.5) to 6.10 (GPT-5.4-mini). A point and a half separates the best from the worst. There is no model left to escape to. The lever that used to work two years ago has nothing left to move.

Why this is attackable at all

For a while I comforted myself with the thought that since it's a hallucination, it must be random. Next time the model invents a different name, the attack doesn't connect, the whole thing fizzles out. Except the hallucination is stable — on two axes at once.

First axis: time. Same model, same prompt, ten runs. In the 2024 study, 43 percent of invented names came back every single time. Not random error — repeatable model behavior.

Second axis, and this is the one that stopped me: models. Those five 2026 models invented 127 identical non-existent package names: 109 on PyPI, 18 on npm. Not different hallucinations in different models. The same ones, in all five. That's a model-independent attack surface, and no single-model study would ever have shown it.

If the same name comes back across runs and shows up in five different models at once, that's not noise. That's something you can predict. And a predictable name is one an attacker registers in advance on npm or PyPI, with their own code inside, and then waits for anyone, on any model, to get that suggestion. That's slopsquatting, and it's why the model choice from the previous section saves nothing here.

And the mechanism has already worked in practice. A package called react-codeshift — a name the model glued together from two real ones (jscodeshift and react-codemod) — spread through 237 repositories. Driven not by humans copy-pasting an install command, but by agents executing their own generated output. With no human ever glancing at the name and noticing something was off. The only reason there was no payload is that a researcher at Aikido registered the name first, before someone with worse intentions did. The distribution channel was ready and running. The only thing missing was the charge.

Where my review ends

I run an agent-based code review system in production. The first thing people say when I talk about it: just add another agent, a security one, and let it catch this. I tried thinking that way. But the problem isn't that my prompt is too weak.

My agent reads the diff. The diff contains import package-name. The agent will check whether the code uses it sensibly. It won't read what that package does after installation, because the package's code isn't in the diff. The best AGENTS.md I could ever write describes how I think about my code. It doesn't describe the forty thousand lines in node_modules that nobody has ever read. Review covers authorship. It doesn't cover the trust I handed out at install time.

Say I build that agent anyway. It won't help. The most dangerous variant — a prompt injection hidden in a package or in a tool's response — doesn't live in the text you can see during review. The model checks a tool's description once, at connection time. That tool's responses flow straight into the model's context with no equivalent control, and that runtime channel is exactly what the attacker uses. At the moment of review, the instruction isn't there yet. It appears once the code is already running.

This is the difference between a boundary of scope and a boundary of capability: scope I can keep extending with more agents and better prompts, but capability I can't, because an instruction that only appears at runtime is something no prompt at review time will ever see.

Why I can't solve this with stubbornness

And only now can you see why this isn't a matter of trying harder. I can run Opus at maximum effort. I can make an agent audit every dependency, line by line. I'll burn the whole context window, hit the session limit, wait another four hours.

And all of it for nothing, because the prompt injection only fires at runtime, and at the moment I was looking at the diff, it wasn't there yet. All that friction would buy me not security, but the feeling that I checked something.

AI code isn't dangerous at random. It's dangerous in predictable ways. And the thing about predictability is that the attacker reads the same studies I do and gets more out of them, because all they have to do is register a name and wait. My review catches what's visible in the diff. The cheapest thing to attack is exactly the part that isn't in it.


Sources:

Was this helpful?

Slopsquatting: AI code is dangerous in predictable ways | Code Nomad