The AI Wrote the Code. Nobody Can Change It.

Most conversations about AI-generated code are about whether it is correct.

In my experience that is not the most interesting problem. Modern coding models write correct code most of the time, and the mistakes they do make are the ordinary kind that tests catch. Something else goes wrong, and it goes wrong later, which is why it keeps catching people out.

The code works, but nobody can change it.

The gap is comprehension, not correctness

When a developer builds a feature over three days, they carry a map in their head. What it touches, what depends on it, what will break if it moves. That map is the actual deliverable, and the code is really just its residue.

When an A.I. model produces the same feature in ninety seconds, the code arrives but the map never gets made.

Nobody notices the difference on day one, because the tests pass and the code runs. The difference shows up the first time somebody needs to change it, and by then the person who prompted it has moved on to other work and remembers roughly as much about the internals as any one else does.

That is the gap. Not “is this code correct” but “does anybody alive know what it connects to.”

A codebase nobody could take on

That gap is not new. AI did not invent it. AI just mass-produces it. The worst case of it I have worked on was written entirely by humans, over two decades, with no A.i. to be found.

The system is a member relationship management platform (an MRM) run by a large membership association. It worked. It also could not be changed. Every small fix had knock-on effects nobody could predict, because the people who carried the map had left, or had only ever been there for one project’s scope. The backlog ran to hundreds of tickets across the codebases, while paying customers waited on bug fixes, security issues and their own scopes of work. Jobs that should have taken six weeks took two or three times that, and everything not on fire was neglected, ultimately ending up with a giant jinga tower of brownfield code.

The textbook answer is a strangler fig migration: incrementally grow the new system around the old one until the old one can be switched off. The textbook answer assumes you can slow down. This team could not stop delivery to refactor, could not afford a big-bang rewrite, and could not keep going as they were. A genuine catch-22.

What worked was making the map machine-built instead of person-carried. We ran OBY, our code intelligence engine, under every piece of code already being done. Each bug fix and each scope item became an opportunity. As a file was touched, its callers and impact surface were mapped, completion gates checked backward compatibility against the migration goal, and the code that had just been understood got nudged toward where the system was heading. Security triage worked the same way, because most vulnerability reports do not apply to how the code is deployed, and our CodeScan tool focused the developers on the ones that did.

The first measurable ground was gained in the first week. Within a couple of months the codebase was in a visibly different state. Over a bit more than three months, on just one of the codebases involved, the work came to 165 authored commits, 536,392 changed lines across 1,591 files, roughly 380,000 added and 156,000 deleted, while normal customer delivery kept running. Then the in-house team stabilised and took the work over, which was always the origianl plan.

The asymmetry underneath: for a human, writing code is expensive and reusing it is cheap, so you find the existing one and adapt it. For an AI it is exactly backwards. Writing is nearly free and finding what exists is the expensive part, so left alone it writes a fresh copy every time and the gap compounds. Hand the model the existing structure before it writes, and the economics flip back.

Why searching for it does not work

You might be thinking the answer is just better search. It is not.

Editor search finds the function name. It also finds it in comments, in strings, in documentation, and in the three other functions that happen to share a prefix. You get a list, you work through the list, and the one caller that mattered was constructed dynamically so it never appeared in the list at all.

Asking a model to trace the callers is worse, not better. You get an answer that reads as authoritative and is wrong often enough that you cannot rely on any single instance of it. An unreliable answer delivered confidently is more expensive than no answer, because you act on it.

What we built instead

OBY parses the code itself and returns every symbol and its callers deterministically, meaning the same question returns the same answer every time with no model guessing involved. Ask what calls a function and you get the list. All of it, every time, including the dynamic ones.

The OBY tool exists because of three years of watching the alternative fail. The AI would not search the codebase. It invented database names. It called APIs that did not exist. It skipped the backend validation that protects a system and wrote confident front-end checks instead. I spent my time typing the same questions: why did you do that?, why didn’t you look?, why aren’t you using what already exists?. People call the failure hallucination, but it is really optimisation. A model is tuned to complete the task in front of it, and if the tool that would inform it is slow or awkward, it will skip the tool and complete anyway, understanding half the problem if you are lucky.

That is how these systems evolve: From a prompt, to a scope document, to hooks and hard gates, and finally to an engine that answers in milliseconds The goal is to make the informed path faster than guessing, not by instructing the model to use it but by making it the quickest route to an answer. The system stops being a gate and becomes the first path the work takes.

Today OBY runs at the start of every engagement, before anybody touches anything.

What it does not do. OBY does not index string literals or configuration values. If you need to find every place a version number appears across JSON and TOML files, use a plain text search, because OBY will return nothing, however, that nothing is not proof of absence. It maps structure: the symbols and their callers, and the impact surface.

How to tell whether you have this problem

You do not need to hire anybody to work this out. Ask your own team three questions.

  1. Pick a function somewhere in the middle of the system, something that has been there a while and gets touched occasionally. Ask who can list what calls it, without running a search first. If the honest answer is nobody, the comprehension gap is already there, whether or not AI wrote a line of it.
  2. Compare how long the last small change took against how long it was estimated to take. A consistent multiple of three or four is not a team that estimates badly. It is a team working without a map.
  3. Ask what happens when the person who understands the trickiest part of the system takes a month off. If that question makes the room uncomfortable, then the code is not your asset. That person is, and they are a single point of failure nobody has budgeted for.

The third one matters most.

Those three questions work on any provider you are weighing up, including us. If somebody selling you a code audit cannot explain how they answer “what calls this” without guessing, they will read your codebase the same way your team already does. Slower, and on your money. We keep a longer list of questions worth asking any integration partner before you sign, built the same way.

This is not an argument against AI-assisted development

We use these tools every day, across Codex and Claude, plus locally hosted models depending on what the work needs and what the client’s data policy allows. They are genuinely fast, and getting faster.

One honest caveat from the MRM engagement: the tooling was the easy half. Not every developer wants to work this way, and an existing workflow will not survive contact with AI unchanged. You redesign the workflow around what the tools are good at, or you get the costs of AI with none of the compounding.

The point is narrower than “be careful with AI”. Speed of writing has moved enormously. Speed of understanding has hardly moved at all. The distance between those two is where the money now goes, and closing it is a tooling problem rather than a discipline problem.

What is not solvable is discovering the gap during an outage.


If you have inherited a codebase nobody fully understands, which is nothing out of the ordinary now, please know, it is fixable. We map systems like this as a fixed-scope piece of work: you get the caller map, the risk surface, and a plain-language account of what is genuinely fragile within an existing code base.

Have a look at how we take AI-assisted code to production, or what a code review and security audit covers. If you would rather talk it through, get in touch or call 1300 840 340.

Similar Posts