Back to Blog

The Dark Side of AI Coding Tools: What the 2026 Data Shows

82% of surveyed tech leaders traced a production failure to AI code in 2026. What AI coding tools really cost: outages, deleted databases, review overload and burnout.

The Dark Side of AI Coding Tools: What the 2026 Data Shows

The subscription is the cheapest part of an AI coding tool. Whether you pay $20, $100 or $200 a month, the bigger costs don't show up on the invoice. They show up later as code that passes review and breaks in production, apps that leak their users' data, more code than anyone can review, agents with too much access deleting production databases, and engineers burning out.

I'm a software engineer at a big tech company and I use these tools every day. I'm not an AI hater. But I've run into many of the problems in this post myself, and in 2026 there is finally enough data to put numbers on them.

TL;DR

  • AI code looks better at review than it behaves in production. In New Relic's June 2026 survey of 200 US tech leaders, 94% rated AI code higher quality than human-written code at review. 82% could point to a production failure caused by AI code in the previous six months (New Relic, 2026).
  • Agents with broad credentials destroy things fast. A Cursor agent running Claude deleted PocketOS's production database in about 9 seconds, and a terraform destroy run by Claude wiped 2.5 years of DataTalks.Club course data. In both cases, scoped permissions would have stopped the agent where its instructions didn't.
  • Review is the new bottleneck. As AI adoption rose across about 22,000 developers, median time in PR review went up 441% and bugs per developer went up 54% (Faros AI, 2026).
  • Security holes are rising. Georgia Tech researchers linked 35 CVEs to AI-written code in March 2026 alone, against about 18 from May to December 2025 (Georgia Tech, 2026).
  • People pay too. Juniors who learned with AI scored 50% on a comprehension quiz against 67% for those who coded by hand (Anthropic, 2026), and a survey of 442 developers found that AI adoption raises burnout through higher job demands (Feng et al., 2026).

Why does AI code pass review and break in production?

AI code passes review because it looks clean, and it breaks in production because the agent that wrote it doesn't know how your system behaves under real traffic or why the existing code is written the way it is.

In June 2026 New Relic published its State of AI Coding survey of 200 US tech leaders (New Relic, 2026). At the review stage, 94% rated AI-generated code higher quality than human-written code, and 33% said it was much higher. Once that code shipped, the answers flipped:

  • 78% reported more production incidents over the past 12 months
  • 86% said senior engineers spend more time fixing code
  • 74% said at least a quarter of their AI-generated code needed significant rework
  • 82% could point to a production failure in the past six months caused by AI code

Bar chart from New Relic's June 2026 survey: 94% of tech leaders rated AI code higher quality at review, yet 78% saw more production incidents, 86% said seniors spend more time fixing code, 74% said a quarter or more of AI code needed rework, and 82% traced a production failure to AI code

If the code is better at review, why does it keep breaking? I see four reasons in my own work.

The agent doesn't know the system. It can't see how much traffic a service handles, which days are extra busy, or which downstream dependency falls over first. None of that is in the diff.

It doesn't know why the code is written that way. Every mature codebase has strange conditions. Maybe the payment flow has an odd check because a previous developer hit a provider bug and worked around it. A human reading that code would ask the person who wrote it. The agent doesn't ask. It makes a guess and keeps going, and the guess is often wrong.

Reviewers go blind on large diffs. When an agent produces hundreds of lines at once, the reviewer feels overwhelmed. After a while you stop reading closely, and code gets approved because it looks fine. The edge cases that an engineer who has worked in the code for years would handle without thinking are exactly the ones nobody checks.

Agent debt piles up. This is sloppy AI code that ships and never gets cleaned up. Like any technical debt it grows, and eventually it turns into a production incident.

What happens when the AI tool itself gets worse?

Your team's output drops, and it's hard to notice. In March and April 2026 developers spent weeks complaining that Claude Code had gotten worse. Stella Laurenzo, director of the AI group at AMD, analyzed 6,852 Claude Code sessions and found the agent was reading far less code before making edits. It went from 6.6 file reads per edit to 2.0 (GitHub issue #42796, 2026).

Anthropic's postmortem traced the drop to three changes it had shipped, including a caching bug that wiped the model's thinking on every turn. That bug "made it past multiple human and automated code reviews, as well as unit tests, end-to-end tests, automated verification, and dogfooding" (Anthropic, 2026).

Think about what that means for a company with hundreds of engineers who lean on these tools. For weeks, the whole engineering org runs at reduced output, and almost nobody notices, because when an agent performs badly you shrug and think "it's AI, that happens sometimes."

Anthropic says none of this was intentional, and I believe them. It still made me uneasy. When your team depends on a model from Anthropic, OpenAI or anyone else, the vendor decides how good that model is today. If a provider ever chose to reduce quality, the developers who depend most on the tool would be hit hardest, and a handful of companies now shape the quality of a large share of the world's new code.

Can an AI agent delete your production database?

Yes. It has happened in at least three publicly reported cases in the past year, and each time the agent had far more access than its task needed.

PocketOS: production gone in 9 seconds

In April 2026 an AI agent in Cursor, running Anthropic's Claude Opus 4.6, was working on a task in PocketOS's staging environment and hit a credential problem. It went looking for credentials and found a Railway API token in an unrelated file. The token had been created to manage custom domains, but it had no scope limits. Anyone holding it could do anything in any environment.

The agent used it to delete a storage volume through Railway's API, believing the deletion would only affect staging. The volume held the production database. The backups lived on the same volume, so they went too, and the newest copy stored anywhere else was about three months old. The whole thing took about 9 seconds (The Register, 2026).

Diagram of the PocketOS incident: a Cursor agent on a staging task found an unscoped Railway API token and deleted a volume that held both the production database and its backups, in about 9 seconds

The agent's rules told it never to guess and never to run destructive commands unless asked. When founder Jer Crane asked why it did it anyway, it answered: "I guessed that deleting a staging volume via the API would be scoped to staging only. I didn't verify." Railway later restored the data from its own backups (The Register, 2026).

It could have been prevented without touching the agent. Production and staging shouldn't share infrastructure, each environment should have its own token, and the agent should never have had access to a token it didn't need. I still blame the agent, because it ignored instructions it had been given. But a rule you write for an agent is only a request, and this one got ignored.

DataTalks.Club: 2.5 years of data and one terraform destroy

Alexey Grigorev runs DataTalks.Club, a platform for free data engineering courses, and manages its infrastructure with Terraform (Alexey Grigorev, 2026). Terraform works with two files. The configuration describes what you want, such as servers, a database and a network. The state file records what Terraform has actually built. Terraform compares the two and builds whatever the state file doesn't list.

His state file lived on his laptop instead of in shared remote storage. When he switched laptops, the configuration came with him but the state didn't. Terraform concluded nothing had been built yet and started building everything a second time. He stopped it partway, which left a few duplicate resources running next to the real ones.

Then Claude, cleaning up the duplicates, unpacked an old Terraform archive from the previous laptop. In his words: "I didn't notice Claude unpacking my Terraform archive. It replaced my current state file with an older one that had all the info about the DataTalks.Club course management platform." That older state described the whole production platform. Claude ran terraform destroy, and Terraform deleted the database, the network, the servers and the automated snapshots, since Terraform had created those too. The database held 2.5 years of course submissions, with about 1.9 million rows in one table.

At midnight he paid to upgrade his AWS support plan so he could get someone from AWS on the phone. AWS had a snapshot on its side that he couldn't see from his own console, and the platform was restored 24 hours after the deletion.

A human would normally review the plan before running terraform destroy against production. Claude ran it without questioning what the plan would delete.

Amazon Kiro: 13 hours of AWS Cost Explorer

This doesn't only happen to small companies. In December 2025, Amazon engineers let Kiro, Amazon's own AI coding agent, fix a problem in AWS Cost Explorer, the tool customers use to track their AWS spending. Kiro decided the best fix was to delete the environment and recreate it. Cost Explorer was down for about 13 hours in one mainland China region, according to the Financial Times, as summarized by The Decoder (2026).

Two safeguards should have stopped it. By default Kiro asks before it acts, and production changes need a second person to sign off. But the engineer running Kiro had broader permissions than expected, Kiro inherited them, and no second approval was required. Amazon told the FT it was a coincidence that AI tools were involved. Its public response called the incident "user error, specifically misconfigured access controls" (Amazon, 2026). Either way, the agent inherited more access than the task needed.

What would have prevented these incidents?

Ordinary infrastructure rules would have stopped all three:

  • Give each environment its own credentials, so a staging task can only reach staging.
  • Give an agent the narrowest token that does the job, never an engineer's full access.
  • Keep production and staging on separate infrastructure. That goes double for storage, and backups should never sit on the same volume as the data.
  • Keep Terraform state in locked remote storage. A state file on a laptop is one lost laptop away from a rebuild or a destroy.
  • Have a person approve every destructive command against production, including deletes, destroy, force pushes and migrations.

How are attackers targeting AI coding tools?

Attackers now use the AI agent on your machine as the tool that finds and steals your secrets.

In August 2025 someone published eight malicious versions of Nx, a JavaScript build tool with about 6 million weekly downloads (Nx, 2025). The bad versions were live for roughly four to five hours. A postinstall script checked whether the machine had the Claude, Gemini or Amazon Q command-line tools installed. If it found one, it ran it with its safety checks turned off (claude --dangerously-skip-permissions, gemini --yolo or q --trust-all-tools) and told the agent to search the disk for GitHub tokens, npm tokens, SSH keys, .env files and crypto wallets. It then used the victim's GitHub token to create a public repository in their own account and uploaded the loot there (StepSecurity, 2025). Over 1,700 developers had their secrets published this way (Wiz, 2025).

Six-step diagram of the Nx malware: npm install, the postinstall script runs, it finds the Claude, Gemini or Amazon Q CLI, runs it with safety checks off, the agent searches the disk for secrets, and the loot is uploaded to a public repo in the victim's own GitHub account

In February 2026 a worm went after AI coding tools directly. At least 19 npm packages from two accounts used names close to Claude Code and OpenClaw, so anyone who mistyped a package name installed the wrong one. Once installed, it collected API keys for AI providers, added its own MCP server to the developer's AI tools with hidden instructions telling the assistant to gather SSH keys and AWS credentials, and spread by publishing more infected packages with stolen npm tokens and committing itself into repositories through the GitHub API (Socket, 2026).

npm supply-chain attacks are older than AI agents, but an agent with broad permissions widens the damage. It can read and send anything you can on your laptop. Both attacks also relied on npm running a package's install scripts automatically. Go modules have no install scripts, so go get never runs code from a dependency while downloading it. If you want the ecosystem side of this, Go vs Node.js supply chain security compares how the two package managers handle it.

Is AI breaking code review?

Yes. AI writes code faster than people can review it, and the data shows reviewers checking less as a result.

Faros AI tracked about 22,000 developers across more than 4,000 teams. As AI adoption went up, so did the problems (Faros AI, 2026):

  • Bugs per developer: +54%
  • Incidents per pull request: +243%
  • Median time in PR review: +441%
  • Pull request size: +51%
  • PRs merged with no review: +31%

Bar chart of Faros AI data: as AI adoption rose, median PR review time went up 441%, incidents per PR 243%, bugs per developer 54%, PR size 51% and PRs merged with no review 31%

A 2026 study followed 400 reviewers across 11,429 reviews of AI agent PRs. Reviewers who had already seen the most agent PRs approved them 14.5 percentage points more often than those who had seen the fewest, and inline review comments dropped by 22% over the study. The researchers call it habituation. The more AI code reviewers saw, the less closely they checked it (Yu et al., 2026).

Sonar's January 2026 State of Code survey of more than 1,100 developers found that 38% say reviewing AI code takes more effort than reviewing a colleague's, and 53% have seen AI code that looked correct but wasn't reliable. 96% don't fully trust AI-generated code, yet only 48% always check it before committing (Sonar, 2026). I think the gap between those two numbers comes from how tiring it is to review AI code carefully, day after day. I've caught myself skimming large agent diffs too.

Microsoft published ten months of data on the Copilot coding agent in the dotnet/runtime repository, where engineers assigned it tasks and it opened 878 pull requests (Microsoft .NET Blog, 2026). In the first month only 41.7% were merged, partly because the agent lacked build instructions. After the team fixed that, the rate climbed, and across all ten months 67.9% of agent PRs were merged. Engineers' own PRs merged at 87.1%. Agent PRs also cost more review. Merged agent PRs averaged 16.5 review comments against 12.4 for engineers' PRs, and humans pushed their own commits to 396 of the 878 agent PRs, about 45%, compared with 10.3% of merged human PRs.

StudySampleWhat changed with AI
Faros AI (2026)~22,000 developers, 4,000+ teamsMedian PR review time +441%, bugs per developer +54%
Habituation study (2026)400 reviewers, 11,429 reviewsApprovals +14.5 points for the most-exposed reviewers, inline comments -22%
Sonar (Jan 2026)1,100+ developers96% don't fully trust AI code, 48% always check it
Microsoft dotnet/runtime (2026)878 agent PRs over 10 months67.9% merged vs 87.1% for engineers' PRs

Is AI-generated code less secure?

The vulnerabilities traced to AI code are growing fast, and the published numbers are likely an undercount.

A lab at Georgia Tech runs the Vibe Security Radar. It goes through publicly reported vulnerabilities (CVEs), finds the commit that introduced each bug, and checks for signs that AI wrote it, such as an AI co-author tag. In March 2026 alone it found 35 CVEs linked to AI code, against about 18 from May to December 2025, after tracking began (Georgia Tech, 2026). The project's founder, Hanqing Zhao, estimates the real number is five to ten times what they can detect, because most AI-written code leaves no trace that it came from an AI (Infosecurity Magazine, 2026).

Bar chart from Georgia Tech's Vibe Security Radar: about 18 CVEs linked to AI-written code across May to December 2025, and 35 in March 2026 alone

GitGuardian found 28.65 million new hard-coded secrets in public GitHub commits in 2025, up 34% on 2024 and the largest yearly jump it has ever recorded. Commits made with Claude Code leaked secrets at 3.2%, about twice the 1.5% baseline, though a developer still approved each of those commits (GitGuardian, 2026).

In May 2026 researchers at RedAccess found around 380,000 vibe-coded apps, built on platforms like Lovable, Base44, Netlify and Replit, that were open to anyone on the web. About 5,000 of them were leaking sensitive data, including medical records, financial data and internal company strategy (Axios, 2026).

CodeScene adds a warning for anyone who has stopped reading the code. A peer-reviewed study it published found that AI coding assistants raise defect risk by at least 30% in unhealthy code (CodeScene, 2026), and the company's own white paper puts it at 60% or more. The messier a codebase gets, the worse the agent performs in it. Every unreviewed change makes the code a little messier, so teams that let the agent run unchecked are making its future work harder.

How are open source maintainers responding to AI?

Several major projects have restricted or banned AI contributions because maintainers can't keep up with low-quality submissions.

  • curl ran a bug bounty from 2019 and paid out more than $100,000 across 87 confirmed vulnerabilities. In 2025 about 20% of submissions were AI slop (Daniel Stenberg, 2025), and the maintainers couldn't keep up with the volume of bad reports. The bounty ended on January 31, 2026 (Daniel Stenberg, 2026).
  • Ghostty, the terminal maintained by Mitchell Hashimoto, now accepts AI-generated code from outside contributors only for work the maintainers have already agreed on. Anything else is closed, and people who submit bad AI contributions are banned. Hashimoto wrote that AI had "increased the 'bad' count by 10x if not more" (Ghostty, 2026).
  • Godot banned autonomous agents, vibe coding and substantial AI-generated code in June 2026, leaving only small assisted tasks like code completion: "AI cannot take responsibility, and we can't trust heavy users of AI to understand their code enough to fix it" (Godot Foundation, 2026).
  • Rust adopted an LLM policy for the rust-lang/rust repository in August 2026: "It's fine to use LLMs to answer questions, analyze, distill, refine, check, suggest, review. But not to create" (Rust Blog, 2026).
  • GitHub shipped repository settings that let maintainers turn off pull requests or restrict them to collaborators (GitHub, 2026), and later a cap on how many open PRs a contributor can have (GitHub, 2026).

What does AI do to junior developers' skills?

The cost for juniors is skill. In my experience, juniors who lean on AI often struggle with the fundamentals, and now there's research on it, some of it from the AI companies themselves.

Anthropic ran a randomized trial with 52 mostly junior engineers learning Trio, a Python async library they hadn't used before. The group that used AI scored 50% on a comprehension quiz afterwards. The group that coded by hand scored 67%. The AI group was only about two minutes faster, a difference that wasn't statistically significant (Anthropic, 2026). They lost understanding and got no measurable speed in return. I went through that study in detail in Should You Still Write Code by Hand?, including which ways of using AI kept scores high. For experienced engineers who want to avoid the same slide, How to Keep Your Coding Skills Sharp When AI Writes the Code covers which skills fade first and how to practice them.

AI can write the code for you, but everything above shows why someone still has to understand it. I think the fundamentals matter more now than they did before. I built LevelUpGo around writing Go yourself, from your first line to production code, with exercises that only pass when your code compiles and the tests pass. Language choice helps too. Why Go is the best language for AI-written code covers how Go's compiler and tooling catch more of an agent's mistakes before they ship.

Do AI coding tools cause burnout?

The research says they can, mostly through the higher workload and expectations that come with them.

Researchers at Oregon State University surveyed 442 professional developers about AI use and burnout. They found that generative AI adoption raises burnout, mainly because it raises job demands: more workload and more organizational pressure (Feng et al., 2026). Two participants put it plainly. One said: "the tools are fine, what they've done to leadership expectations is absolutely awful." Another said: "I move fast with AI and move mountains of work, but I am losing my passion for the work."

Researchers from UC Berkeley spent April to December 2025 inside a US tech company of about 200 employees, watching how people actually worked with AI (Harvard Business Review, 2026). The work grew in three ways:

  • Task expansion. Because AI could fill in what people didn't know, they started taking on work that had never been theirs.
  • Blurred boundaries. Starting a task got so cheap that work slid into lunch, meetings and evenings.
  • More multitasking. People wrote code by hand while AI drafted another version, ran several agents at once, and reopened old tasks because AI could now handle them, without anyone asking them to.

The researchers describe a loop. AI speeds up some tasks, which raises expectations for speed, which makes workers rely on AI even more. One engineer in the study summed it up: "You had thought that maybe... you can work less. But then really, you don't work less. You just work the same amount or even more."

Loop diagram from the UC Berkeley study: AI speeds up some tasks, expectations for speed rise, workers rely on AI even more, and the loop repeats

My take

AI can bring real value and make us more productive. I use it every day. But it isn't the solution to everything.

What struck me about these studies is how much of them matches what I see at work. In all my years as a software engineer, I've never felt that software in general was as low quality as it is now. To me that says we haven't found the balance between speed and quality yet, and I think quality matters more than speed.

So I treat AI as a tool that makes a lot of mistakes. I review what it writes instead of trusting it, and I give it only the permissions the task needs, because you can't secure AI with prompts.

After looking at all this data, I keep asking myself whether the cost is higher than the value these tools return. I don't have a final answer yet. I'm fairly sure it's a lot more than the subscription price, though.

FAQ

What are the biggest risks of AI coding tools?

The biggest risks are code that passes review but fails in production, agents with too much access deleting data, a review backlog that leads to less careful checking, security vulnerabilities, and lost skills for developers who delegate their learning. In New Relic's 2026 survey, 82% of tech leaders could name a production failure caused by AI code in the previous six months (New Relic, 2026).

Why does AI-generated code break in production?

The agent doesn't know how your system behaves in production. It can't see traffic patterns, busy days or why existing code handles an edge case the way it does, and it rarely stops to ask. Large AI diffs also make reviewers skim, so those gaps get approved.

How do you stop an AI agent from deleting production data?

Limit what it can reach. Give agents separate, narrowly scoped credentials per environment, never share storage or backups between production and staging, keep Terraform state in locked remote storage, and require a human to approve every destructive command. Instructions in a prompt are not a security boundary.

Is AI-generated code secure?

Not by default. Georgia Tech's Vibe Security Radar linked 35 CVEs to AI-written code in March 2026 alone, and the project's founder estimates the true number is five to ten times higher (Infosecurity Magazine, 2026). Review authentication, authorization, payments and anything that handles user data line by line.

Should junior developers use AI coding tools?

Use them to explain concepts and errors, not to write code you don't understand yet. In Anthropic's 2026 trial, juniors who learned a new library with AI scored 17 points lower on comprehension and weren't meaningfully faster (Anthropic, 2026).

Sources

Write Go like a senior engineer

Interactive lessons in your browser. The first ones are free.

Try a free lessonOr create a free account