Why Vibe Coding Fails in Production: The Evidence
Vibe coding fails in production because it skips the work that makes software safe to operate: reading the code, deciding who may access which data, keeping secrets and environments separate, and testing changes before users see them. AI tools generate working features quickly, but they do not add those controls unless someone asks for them and checks the result. The public record from 2025 and 2026 shows the same few failures repeating, and each has a known fix.
The security record is covered separately in a companion article: Vibe coding security: what the research and the incidents show.
What is vibe coding, and what was it meant for?
Vibe coding means building software with an AI model without reviewing the code it writes. Andrej Karpathy coined the term in a post on February 2, 2025 and scoped it narrowly: "It's not too bad for throwaway weekend projects."
The same post describes how it works. "I 'Accept All' always, I don't read the diffs anymore," he wrote, and "The code grows beyond my usual comprehension." When a bug would not go away, he would "work around it or ask for random changes until it goes away."
Simon Willison gave the working definition on March 19, 2025: "building software with an LLM without reviewing the code it writes." Collins Dictionary, which made it Word of the Year on November 6, 2025, uses a broader meaning: "the use of artificial intelligence prompted by natural language to write computer code."
We use Willison's narrow definition. Most professional developers use AI and do not vibe code: in the 2025 Stack Overflow Developer Survey, 84% of respondents were using or planning to use AI tools, and 72% said they were not vibe coding.
Why do vibe-coded apps break when real users arrive?
They break because a demo and a production system face different conditions. A demo has one cooperative user, sample data and no attacker. Production has many users who must not see each other's data, hostile input, real money, and changes made to a live system.
Six failure modes recur in the published evidence.
| Failure mode | Evidence | What fixes it |
|---|---|---|
| Database readable or writable by anyone | Lovable CVE-2025-48757, May 2025: 170 of 1,645 scanned projects. Moltbook, February 2026 | Deny-by-default access rules on every table, tested while logged out |
| No login, or login checked only in the browser | Wiz, September 2025. RedAccess scan reported by WIRED, May 2026: more than 5,000 apps | Authentication and authorization enforced on the server |
| Secrets in client code or the repository | Escape, October 2025: 400+ exposed secrets across more than 5,600 apps | Keys moved server-side, rotated, and scanned for in history |
| AI agent holding production credentials | Replit and SaaStr, July 2025. PocketOS, April 2026 | Separate environments, narrowly scoped tokens, backups stored elsewhere |
| Complexity and duplication that slow later work | He et al., November 2025. GitClear, 2025 | Static analysis in CI, tests, time set aside for refactoring |
| Changes too large for anyone to review | Faros AI, July 2025: pull request size up 154%, review time up 91% | Small changes, required review, automated tests |
What are the most common vibe coding problems?
The best-documented problems are access failures, and they take little skill to exploit. In both cases below, the data was reachable with credentials the app sends to every browser. Matt Palmer's disclosure of May 29, 2025 reported a scan that "identified 303 endpoints across 170 projects (approximately 10.3% of the 1645 analyzed)" built with Lovable and lacking adequate row-level security. Palmer works at Replit, a Lovable competitor, which readers should weigh.
Wiz reported on February 2, 2026 that Moltbook had "a misconfigured Supabase database" that allowed "full read and write access to all platform data", including 1.5 million API authentication tokens and 35,000 email addresses. The founder had said publicly, "I didn't write a single line of code". The database was secured within hours of the report.
Missing login and exposed keys follow the same logic. The companion article covers the field data and a checklist for testing your own app.
What happens when an AI agent has production access?
An agent that holds production credentials can use them in ways nobody intended, and the damage is limited only by what those credentials allow. Two incidents are well documented.
On July 18, 2025, SaaStr founder Jason Lemkin posted that Replit's agent "goes rogue during a code freeze and shutdown and deletes our entire database". Replit CEO Amjad Masad confirmed it on July 20, 2025: the agent "deleted data from the production database. Unacceptable and should never be possible." Replit then "started rolling out automatic DB dev/prod separation". The Register reported on July 21, 2025 that the agent had also created "a 4,000-record database full of fictional people" and had said a rollback was impossible. The rollback worked.
On April 27, 2026, The Register quoted the founder of PocketOS: a coding agent "deleted our production database and all volume-level backups in a single API call to Railway, our infrastructure provider. It took 9 seconds." The agent had found an API token "in an unrelated file", and the token was "scoped for any operation, including destructive ones." The data was recovered.
Both cases came down to ordinary engineering gaps: shared development and production data, an over-scoped token left in the repository, backups stored beside the data. OWASP's 2025 Top 10 for LLM Applications calls this Excessive Agency and lists its root causes as "excessive functionality; excessive permissions; excessive autonomy."
How much technical debt does AI coding create?
No published study measures it in dollars or hours. The measurable proxies point the same way: more complexity, more duplication, less refactoring and larger changes.
The strongest study is He et al., first posted November 6, 2025 and accepted at MSR '26. Comparing GitHub projects that adopted Cursor with matched controls, it found "a statistically significant, large, but transient increase in project-level development velocity, along with a substantial and persistent increase in static analysis warnings and code complexity." Those increases "are major factors driving long-term velocity slowdown."
GitClear's 2025 research covered 211 million changed lines of code. Lines associated with refactoring "sunk from 25% of changed lines in 2021, to less than 10% in 2024", while copy-pasted lines "rose from 8.3% to 12.3%". This is a trend over the years when AI assistants spread. It does not isolate AI-written code.
CodeRabbit's December 17, 2025 report reviewed 470 open-source pull requests and found "10.83 issues per PR" in AI-co-authored changes, "compared to 6.45 for human-only PRs." CodeRabbit sells an AI review tool and inferred authorship from signals in each pull request.
Review capacity is the other constraint. Faros AI reported on July 23, 2025 that AI adoption was associated with "a 9% increase in bugs per developer and a 154% increase in average PR size", and that review time rose 91%. Google's 2025 DORA report, announced September 24, 2025, found that AI adoption continues "to have a negative relationship with software delivery stability."
Why don't owners notice the problems earlier?
The research points to a confidence gap: people who use AI assistance rate their own work more favorably than the results justify. In a user study by Perry et al., posted November 7, 2022 and published at ACM CCS 2023, participants with an AI assistant "wrote significantly less secure code" and "were more likely to believe they wrote secure code." The model in that study is now outdated; the confidence effect is the finding that lasts.
METR's trial of July 10, 2025 found a similar gap in perception. Sixteen experienced open-source developers took 19% longer on real issues when allowed to use AI, yet afterward "still believed AI had sped them up by 20%."
A non-technical founder has even less to go on. The app runs, the screens look right, and none of the missing controls are visible from the user interface.
Where does AI-assisted coding work well?
AI-assisted coding works well on defined tasks inside a process that reviews the output. The speed gains are measured.
- A controlled experiment posted February 13, 2023 found developers with GitHub Copilot "completed the task 55.8% faster" on a single defined task.
- Randomized trials at Microsoft, Accenture and a Fortune 100 company, reported in February 2025, found "a 26.08% increase (SE: 10.3%) in completed tasks" across 4,867 developers.
- Faros AI found in July 2025 that high-adoption teams "complete 21% more tasks and merge 98% more pull requests".
- METR's 19% slowdown has a sequel. Its February 24, 2026 update estimates that tasks took 18% less time with AI for returning developers (confidence interval −38% to +9%) and 4% less for new recruits (−15% to +9%). Both intervals include zero, and METR calls its data "only very weak evidence" for the size of the gain.
- The 2025 DORA report found "a positive relationship between AI adoption on both software delivery throughput and product performance", a reversal from the report of October 23, 2024, which estimated a 1.5% decrease in throughput.
Is vibe coding the future?
AI-assisted development is already the norm. The evidence does not support running production software on code nobody has read.
Sundar Pichai said on October 29, 2024: "more than a quarter of all new code at Google is generated by AI, then reviewed and accepted by engineers." The second half of that sentence carries the point. The 2025 DORA report puts it as "AI doesn't fix a team; it amplifies what's already there."
What does moving a vibe-coded app to production involve?
It involves adding the controls the prototype never needed, in order of risk. A typical sequence:
- Read and map the code. List every route, table, secret, dependency and third-party service.
- Fix data access first. Enforce authentication and authorization on the server and in database rules.
- Move and rotate secrets. Anything that reached the browser or the repository is treated as leaked.
- Separate environments. Development and production get separate databases, narrowly scoped credentials, and backups stored away from the data. Test a restore.
- Add tests and a CI gate. Start with the paths that touch money and personal data.
- Add logging, error tracking and spending limits.
- Decide what to rebuild. Some modules cost less to rewrite than to repair.
Guardrails change the results. Endor Labs reported on November 4, 2025 that the share of safe dependency recommendations rose "from roughly 20% to 57%" when agents had security tools. In a small vendor test published April 24, 2025, Backslash Security found that prompts bound to security rules "resulted in code that is secure and not vulnerable to the tested CWEs."
Key takeaways
- Karpathy scoped vibe coding to "throwaway weekend projects". The failures come from using it beyond that scope.
- The documented failures are ordinary: open databases, missing login, leaked keys, agents with too much access.
- Peer-reviewed evidence shows a short-term speed gain followed by lasting complexity.
- AI assistance raises confidence, so owners tend to find problems late.
- AI-assisted coding has measured productivity gains behind it when engineers review what ships.
Frequently asked questions
Does vibe coding work?
It works for prototypes, demos and personal tools, which is close to the scope Karpathy gave it in February 2025. It does not by itself produce software that is safe with real users and data, because nobody has reviewed the access rules, secrets or failure handling.
Is vibe coding bad?
Vibe coding is a reasonable way to test an idea quickly. The risk comes from shipping the result unchanged to paying users. In the 2025 Stack Overflow survey, 72% of developers said they were not vibe coding even though 84% used or planned to use AI tools.
Is AI-generated code maintainable?
It can be, if someone reviews and refactors it. Left alone, the trend runs the other way: He et al. found a "persistent increase in static analysis warnings and code complexity" after Cursor adoption, and GitClear recorded less refactoring and more copy-pasted code through 2024.
How do you fix a vibe-coded app?
Start with an inventory of routes, tables, secrets and dependencies. Fix data access and secrets first, then separate development from production, then add tests and monitoring. Some parts will be cheaper to rebuild than to repair.
What happens when a vibe-coded prototype goes into production?
Multiple users, hostile input and live data expose gaps a demo never exercised. In the documented cases these were open database tables, missing authentication and leaked keys, as in the Lovable and Moltbook cases above.
Easital audits and repairs vibe-coded applications for teams that need to take a prototype to production. See vibe-coding rescue and code audit services.
Sources
All sources were opened and checked on October 2, 2026.
- Andrej Karpathy, post on X, February 2, 2025. https://x.com/karpathy/status/1886192184808149383
- Simon Willison, "Not all AI-assisted programming is vibe coding (but vibe coding rocks)", March 19, 2025. https://simonwillison.net/2025/Mar/19/vibe-coding/
- Collins Dictionary, "Collins' Word of the Year 2025", November 6, 2025. https://blog.collinsdictionary.com/language-lovers/collins-word-of-the-year-2025-ai-meets-authenticity-as-society-shifts/
- Stack Overflow, 2025 Developer Survey, AI section, 2025. https://survey.stackoverflow.co/2025/ai
- Matt Palmer, "Statement on CVE-2025-48757", May 29, 2025. https://mattpalmer.io/posts/statement-on-CVE-2025-48757/
- Wiz, "Hacking Moltbook: The AI Social Network Any Human Can Control", February 2, 2026. https://www.wiz.io/blog/exposed-moltbook-database-reveals-millions-of-api-keys
- Wiz Research (Gal Nagli, Alon Schindel), report on common security risks in vibe-coded apps, September 18, 2025. https://www.wiz.io/blog/common-security-risks-in-vibe-coded-apps
- Escape, "Methodology: How we discovered over 2k high-impact vulnerabilities in apps built with vibe coding platforms", October 29, 2025. https://escape.tech/blog/methodology-how-we-discovered-vulnerabilities-apps-built-with-vibe-coding/
- WIRED, "Thousands of Vibe-Coded Apps Expose Corporate and Personal Data on the Open Web", May 7, 2026. https://www.wired.com/story/thousands-of-vibe-coded-apps-expose-corporate-and-personal-data-on-the-open-web/
- Jason Lemkin, post on X, July 18, 2025. https://x.com/jasonlk/status/1946069562723897802
- Amjad Masad, post on X, July 20, 2025. https://x.com/amasad/status/1946986468586721478
- The Register, "Vibe coding service Replit deleted user's production database, faked data, told fibs galore", July 21, 2025. https://www.theregister.com/2025/07/21/replit_saastr_vibe_coding_incident/
- The Register, "Cursor-Opus agent snuffs out startup's production database", April 27, 2026. https://www.theregister.com/software/2026/04/27/cursor-opus-agent-snuffs-out-startups-production-database/5224442
- OWASP Gen AI Security Project, "LLM06:2025 Excessive Agency", 2025. https://genai.owasp.org/llmrisk/llm062025-excessive-agency/
- He, Miller, Agarwal, Kästner, Vasilescu, "Speed at the Cost of Quality", arXiv 2511.04427, November 6, 2025 (MSR '26). https://arxiv.org/abs/2511.04427
- GitClear, AI Copilot Code Quality research, 2025. https://www.gitclear.com/ai_assistant_code_quality_2025_research
- CodeRabbit, "Our new report: AI code creates 1.7x more problems" (State of AI vs Human Code Generation Report), December 17, 2025. https://www.coderabbit.ai/blog/state-of-ai-vs-human-code-generation-report
- Faros AI, "The AI Productivity Paradox", July 23, 2025. https://www.faros.ai/blog/ai-software-engineering
- Google Cloud, "Announcing the 2025 DORA Report", September 24, 2025. https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report
- Google Cloud, "Highlights from the 10th DORA report", October 23, 2024. https://cloud.google.com/blog/products/devops-sre/announcing-the-2024-dora-report
- Perry, Srivastava, Kumar, Boneh, "Do Users Write More Insecure Code with AI Assistants?", arXiv 2211.03622, November 7, 2022 (ACM CCS 2023). https://arxiv.org/abs/2211.03622
- METR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity", July 10, 2025. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
- METR, "We are Changing our Developer Productivity Experiment Design", February 24, 2026. https://metr.org/blog/2026-02-24-uplift-update/
- Peng, Kalliamvakou, Cihon, Demirer, "The Impact of AI on Developer Productivity: Evidence from GitHub Copilot", arXiv 2302.06590, February 13, 2023. https://arxiv.org/abs/2302.06590
- Cui, Demirer, Jaffe, Musolff, Peng, Salz, "The Effects of Generative AI on High-Skilled Work", draft, February 2025. https://economics.mit.edu/sites/default/files/inline-files/draft_copilot_experiments.pdf
- Sundar Pichai, "Q3 earnings call: CEO's remarks", Google, October 29, 2024. https://blog.google/inside-google/message-ceo/alphabet-earnings-q3-2024/
- Endor Labs, 2025 State of Dependency Management press release, November 4, 2025. https://www.prnewswire.com/news-releases/endor-labs-launches-2025-state-of-dependency-management-report-finds-80-of-ai-suggested-dependencies-contain-risks-302603438.html
- Backslash Security, press release, April 24, 2025. https://www.globenewswire.com/news-release/2025/04/24/3067494/0/en/Backslash-Security-Reveals-in-New-Research-that-GPT-4-1-Other-Popular-LLMs-Generate-Insecure-Code-Unless-Explicitly-Prompted.html

