Skip to content
Softronic

Vibe Coding in Production: What Breaks First

An AI-built app works in the demo and fails in production. What usually breaks, the order I check it in, and what to fix this week if you already shipped.

Founder, Softronic
7 min read Updated

Andrej Karpathy coined “vibe coding” in February 2025 to describe a way of working where you “fully give in to the vibes” and “forget that the code even exists.” In the same post he said it was “not too bad for throwaway weekend projects,” and he was right about that.

The trouble starts when the code that supposedly doesn’t exist is taking payments, storing customer data or running a business. That’s the version I see: someone built an app with Lovable, Bolt, Cursor or Replit, it works in the demo, real users arrived, and now something is off. I don’t read those codebases top to bottom. I check a short list of things in a specific order, because some failures cost a bad afternoon and others cost your customers’ data. This post is that list.

What breaks first is usually access control: one user can read another user’s data. After that, in the order I check them, come secrets shipped to the browser, payment webhooks that run twice or accept fake events, an AI tool with production access, queries that collapse under real data, failures nobody hears about, and code only its original author can safely change.

What the public data shows

Veracode tested code from more than 100 language models and found that 45% of the samples failed security tests and introduced OWASP Top 10 vulnerabilities. Newer, bigger models didn’t do meaningfully better on security. Two other public sources point the same way:

  • In Stack Overflow’s 2025 Developer Survey, 84% of respondents use or plan to use AI tools, but 46% distrust the accuracy of the output and only 33% trust it. The most common frustration, cited by 66%, was “AI solutions that are almost right, but not quite.”
  • CodeRabbit compared 470 open-source pull requests and found about 1.7 times more issues in AI-coauthored changes than in human-only ones, with excessive I/O showing up roughly eight times as often.

“Almost right” is the key phrase. Code that is obviously broken gets fixed. Code that is almost right goes to production.

Where vibe coding is fine

Prototypes you’ll throw away. One-off scripts. Internal tools with no sensitive data. UI scaffolding, where a broken layout is visible the moment you look at it. In all of these, failure is cheap and loud. The problems come from failures that are silent, and those cluster in a few predictable places.

The order I check things in

Here’s the whole list on one screen. Each check gets its own section below.

# Check What breaks How to verify
1 Data isolation One user can read or change another user’s records Run the RLS query below, then request another user’s ID while logged in
2 Secrets Keys like sk_live or service_role reach the browser or the git history Search the build output and the git history for anything that looks like a key
3 Money and side effects Webhooks process the same payment twice or accept a fake one Confirm processed event IDs are logged and signatures are verified
4 AI tool access An agent with production credentials deletes or rewrites data Separate dev and prod databases, no prod keys in the tool, a backup that has been restored
5 Data volume Whole-table loads, one query per row, a new connection per request Load realistic volumes and time the slowest screens
6 Failure handling Endless spinners, half-saved records, stack traces shown to users Cut the network to an external API and check that errors get reported
7 Maintainability Nobody but the original author can change the code safely Tests on auth, payments and core data rules, plus a reproducible deploy

1. Can one user see another user’s data?

This comes first because it’s the most expensive mistake and one of the most common. Many AI-built apps talk to the database directly from the browser, usually through Supabase or Firebase, and rely on database rules to decide who can read what. If those rules are missing, anyone with the public key embedded in your frontend can query your tables.

This isn’t hypothetical. CVE-2025-48757 describes insufficient row-level security in sites generated by Lovable that let unauthenticated attackers read or write arbitrary tables. Lovable disputed the CVE, arguing that each customer is responsible for protecting their own application’s data. Whatever you think of that argument, it tells you who ends up holding the problem: you.

If you’re on Supabase, this query lists tables without row-level security:

select tablename
from pg_tables
where schemaname = 'public' and rowsecurity = false;

Every table in that list is open unless you’ve revoked the grants. Supabase’s own RLS guide says to enable it on every table in an exposed schema. After that, log in as one user and request another user’s record by changing the ID in the URL or the API call. If it comes back, you have a problem regardless of the stack.

2. Are there secrets in the frontend?

Build the app and search the output for anything that looks like a key: sk_live, service_role, secret, API tokens for OpenAI or other paid services. Also check the git history, because a key that was committed and later deleted is still there. Any key that ever reached the browser or the repository should be rotated, not just removed.

3. What happens with money and side effects?

Payments, emails, stock updates, anything that happens once in the real world. Stripe’s documentation states plainly that webhook endpoints might receive the same event more than once, and recommends logging processed event IDs so you don’t process them twice. Generated webhook handlers often skip that step, and they sometimes skip signature verification too, which means anyone can send your endpoint a fake “payment succeeded.” Both are quick to check and quick to fix.

The same idea applies to background jobs: if two workers can pick the same task, something will eventually happen twice.

4. Can the AI tool touch production?

In July 2025, Replit’s agent deleted the production database of SaaStr founder Jason Lemkin’s project during a code freeze he had explicitly declared. The same thing can happen with any agent that holds production credentials, whatever the vendor. I check whether development and production use separate databases, whether the coding tool has access to production keys, and whether there’s a backup that someone has restored at least once.

5. Does it survive real data volumes?

Generated code loves loading an entire table into memory, running one query per row inside a loop, and opening a new database connection per request. None of that shows up with ten test records. It shows up when your biggest customer opens a dashboard. I load the database with realistic volumes and time the slowest screens before anything else in this category.

6. What happens when something fails?

Turn off the network to an external API and see what the app does. Common results: an infinite spinner, a half-saved record, an error message showing a stack trace to the user. Then check whether errors get reported anywhere. If the only way you learn about a failure is a customer complaint, add error tracking before touching anything else.

7. Can a second person change it safely?

The last check is whether anyone other than the original author (human or AI) can modify the code without fear. That means tests on the boundaries that matter (auth, payments, the core data rules), a reproducible deploy, and code a new engineer can follow. When I evaluate engineers, this is exactly the judgment I’m looking for: someone who reads a diff and asks what happens when it fails.

Using AI without the vibe part

I use AI tools every day, and I’m not asking anyone to stop. What changes is who holds the mental model of the system. My rules are simple:

  • Nobody merges code they can’t explain line by line.
  • Coding agents don’t get production credentials, can’t run migrations against production and can’t deploy.
  • Tests go where failures are expensive (permissions, money, data integrity), not where they raise a coverage number.
  • A senior person owns the architecture. The AI speeds up the typing; it doesn’t decide how data flows.

If you already shipped a vibe-coded app

Don’t panic, and don’t rewrite it yet. Most of these apps can be hardened. This week:

  1. Run the row-level security check and fix every open table.
  2. Rotate every key that ever touched the frontend or the repository.
  3. Add idempotency and signature verification to payment webhooks.
  4. Take a backup and restore it somewhere to prove it works.
  5. Turn on error tracking.

A rewrite only makes sense when the data model itself is wrong, and that’s a separate conversation. If you want a second pair of eyes on the list above, get in touch.

Updated September 23, 2026: rewritten, and figures I couldn’t back up were removed.

View all
Engineering

Corrective RAG for Billing Questions on WhatsApp

Billing is where RAG gets checked: the customer is holding the invoice. How corrective RAG, account lookups and a clean human handoff fit on WhatsApp.

7 min read

Ship the next thing. Today.

Book a 30-minute call. We tell you within the call if we can help — including an honest "no" when we can't.