Vibe coding and the hidden costs of AI-generated code.

Development
13 min read

Last week I merged a pull request and realised I didn't really understand the code. Andrej Karpathy has a word for this: vibe coding. The research data from the past year shows a pattern that is worth examining honestly.

Development Opinion

Last week I merged a pull request and, halfway through the review, realised I didn't really understand the code. Not because it was too complex, but because I hadn't written it. An AI assistant had done most of the work, and I was reviewing for "does it work?" rather than "do I understand it?"

It was a small moment, but it stayed with me. Andrej Karpathy has since given it a name: vibe coding. And when you put together the research data from the past year, it turns out moments like this aren't harmless. A pattern is emerging that is worth examining honestly.

Developer behind a large screen with code

Where the term comes from

On 2 February 2025, Andrej Karpathy, co-founder of OpenAI, shared a tweet that clearly struck a chord: "There's a new kind of coding I call 'vibe coding', where you fully give in to the vibes, embrace exponentials, and forget that the code even exists." Over 4.5 million views. Collins Dictionary Word of the Year 2025. The word clearly existed before the awareness did.

Programmer Simon Willison later drew a distinction I find important. If an LLM writes your code but you review, test and understand it thoroughly, that is simply AI-assisted development. Vibe coding begins the moment you accept code without understanding it. A subtle difference, but it shapes everything that follows.

How widespread is it? Y Combinator's managing partner Jared Friedman reported that a quarter of their W25 batch had codebases that were 95% AI-generated. YC CEO Garry Tan then asked the question everyone should be asking: "Let's say a startup with 95% AI-generated code goes out and gets 100 million users. Does it fall over or not?"

Nobody has a definitive answer yet. But the first data is trickling in.

What 470 pull requests reveal

In December 2025, CodeRabbit published an analysis of 470 open-source GitHub pull requests (320 AI-written, 150 human-written). They used a standardised issue taxonomy and statistical testing. The results are not dramatic, but they are consistent.

1.7x more issues per PR in AI code
3x more readability errors
2.74x more XSS vulnerabilities

Source: CodeRabbit State of AI vs Human Code Generation, December 2025

Concretely: 10.83 issues per AI PR versus 6.45 for human PRs. Logic and correctness errors were 1.75x higher. Excessive I/O operations occurred roughly 8x as often. The 3x more readability errors are striking, because readability is exactly what you need when someone else (or your future self) has to maintain the code.

Caveat Human code scored worse on two points: 1.76x more spelling errors and 1.32x more testability problems. And CodeRabbit is itself an AI code review company, which means their motivation to find quality differences is not neutral. The sample of 470 PRs is also relatively small. So treat it as a signal, not a verdict.

The gap between feeling and measurement

The research that has got me thinking the most comes from METR. They had 16 experienced open-source developers (5+ years' experience, 1,500+ commits) complete 246 real tasks, half with AI tools and half without. The result: they were 19% slower with AI.

But here it gets interesting. Beforehand, these developers expected to be 24% faster. Afterwards, they still believed they had been about 20% faster. A gap of nearly 40 percentage points between perception and reality.

I think this reveals something about how we experience productivity. AI tools give you a sense of flow: you type less, code appears on your screen faster, you feel less "stuck". But if the generated code is subtly wrong, or simply does not fit the existing architecture, you spend more time debugging and adjusting than you realise. The speed lies in the writing. The slowdown lies in everything afterwards. And apparently we do not register that second part well.

Glasses reflecting code on screens

Other signals pointing the same way

GitClear analysed 211 million lines of code from 2020 to 2024. Refactoring fell from 24.1% to 9.5% of all changes. Copy-pasted code rose from 8.3% to 12.3% and overtook refactoring for the first time ever. Duplicated code blocks increased eightfold. This is exactly what you would expect when code is produced by a model that does not think about the existing codebase.

Google's DORA Report 2024 found that a 25% increase in AI use led to a 7.2% drop in delivery stability. More output, but also more failed deployments. The original Copilot study from 2023 reported a 55.8% speed gain, but that was on a single, controlled HTTP server task. Later field experiments on more complex work found only 7.5-21.8% more pull requests per week.

Security: where the pattern is clearest

45% Insecure choices by LLMs When LLMs had to choose between a secure and an insecure approach, they picked the insecure option in 45% of cases. Veracode, 2025 (100+ LLMs, 80 tasks)
40% Vulnerable Copilot code Of 1,689 programs generated by Copilot, roughly 40% were vulnerable to CWE Top 25 vulnerabilities. NYU "Asleep at the Keyboard", IEEE 2022
62% Formally verified as insecure At least 62% of 331,000 C programs generated by LLMs were formally verified as vulnerable. FormAI study, 2024

The Stanford study (ACM CCS 2023) added a psychological layer to this: users with AI access wrote significantly less secure code, yet were more convinced that their code was secure. That combination of worse code and higher confidence is a dangerous recipe.

And it's not only about the generated code itself. In 2025, the official Amazon Q Developer VS Code extension was hacked. An attacker injected prompts that deleted user files and disrupted AWS infrastructure. The compromised version got past Amazon's verification and was publicly available for two days. It makes you think about how much trust we place in tools that can themselves become an attack vector.

The debt you don't see building up

"I don't think I have ever seen so much technical debt being created in such a short period of time during my 35-year career in technology."

Kin Lane, API evangelist

Forrester predicted that by 2026, 75% of technology decision-makers would face technical debt of "moderate to high severity" from AI-generated code. Gartner went further, predicting that prompt-to-app approaches would increase software defects by 2,500% by 2028. That sounds hyperbolic, but the underlying logic is not far-fetched: if you produce more code with less understanding, maintenance burden piles up.

The Replit incident of July 2025 made this concrete in a way that's hard to ignore. SaaStr CEO Jason Lemkin used Replit's AI agent to build a web app. On day 8, the agent deleted the entire production database containing 1,206 records, during an explicit code freeze. What happened next was perhaps even more telling: the agent fabricated 4,000 fake records and lied about unit test results.

That is not simply a bug. It is what happens when a system has no understanding of what it is doing, but is optimised to give the impression that everything is fine. And that is precisely the dynamic that vibe coding also creates on a smaller scale among human developers: the feeling that it works, without the certainty that it is right.

What developers themselves say The Stack Overflow Developer Survey 2025 shows that trust in AI accuracy fell from 40% to 29%. The biggest frustration for 66% of developers: "AI solutions that are almost right, but not quite." Meanwhile, 72% say vibe coding is not part of their professional work. Awareness is there, it seems.

What does work

This is not an argument against AI coding tools. I use them every day. GitHub Copilot has over 15 million users, and 90% of Fortune 100 companies have adopted it. At Duolingo it led to 25% faster onboarding. AI is excellent for boilerplate, test generation, documentation and prototyping.

The teams that get the most out of it treat AI output no differently from code written by a junior developer: always review, always understand, always test. Shopify made AI proficiency part of performance reviews. Thoughtworks placed "complacency with AI-generated code" as a warning on their Technology Radar. Without good AI prompt training, teams see 60% lower productivity gains.

It comes down to this

The data points to a pattern that is hard to ignore. AI-generated code contains 1.5x to 2x more bugs, between 40% and 62% security vulnerabilities, and an eightfold higher rate of code duplication. Experienced developers who believe they work faster with AI are in fact working more slowly. It is not that the tools are poor. It is that they tempt us into a way of working we are not used to scrutinising critically.

The difference between productive AI use and vibe coding is ultimately not a technical question. It is a question of ownership. Do you understand what is in your codebase? Can you explain why a particular choice was made? If the answer is no, it does not matter how quickly the code came into being.

Would you like to discuss this further?

We work with AI tools every day and are happy to help you think through how to use them responsibly. No sales pitch, just a conversation.

Get in touch

Edit content