Why the AI Gets Worse as Your Codebase Gets Bigger
Which module did that change touch? What else calls it? What falls out if you move it?
Ask those three questions about a two hundred line file and a developer answers on the spot. Ask them about a six thousand line file you have just pasted whole into a model's context, and the honest answer to all three is that nobody knows, the model included.
I learned this the slow way. For years I let files grow, from five hundred lines to a thousand and beyond, because each one still worked and there was always something more urgent than splitting it up. Then the day would arrive when it had to change. The refactor that should have been a dozen small moves made along the way had become one large frightening one, done under deadline on the file the whole system leaned on.
The fix is not a bigger context window. It is a habit, and the habit has a name.
The bottle that overflows
Context is not free space you fill. Pour a big enough codebase into a model's context and it overflows, and the trouble with an overflowing bottle is that you cannot see which drop was lost. The image is not mine originally, its home is a longer piece of writing about how these models think, but it is exactly right for this problem.
The model does not warn you that it dropped the caller three files over. It writes as though that caller never existed. It writes that confidently, and you find out later.
That is the part that takes people a while to see. There is no error message and no degraded tone. A partial answer and a complete one read identically, which means the failure is invisible precisely where it costs most: on your largest, oldest, most valuable system.
PIM: Prompt, Iterate, Modularise
Three moves in order, run as a loop rather than a checklist.
Prompt. Ask for the smallest useful change, described precisely and set against a piece of code small enough to hand over whole. A prompt is a specification, and every part you leave out gets filled in with whatever completes soonest.
Iterate. Read what came back, correct it, and go again. The loop is where the quality is, not the first response. Most reports of poor results from these tools are reports of a first draft, judged as though it were the last.
Modularise. Keep the units small enough that the model can pick up a whole one at once, hold all of it, reason about it, and put it down. Small bits, and it gets faster and faster, because each step fits inside what the model can hold.
That last move is the one people skip, and it is the one that compounds. Do it and every future request is cheaper, because the model reads a fraction of what it read last time and gets more of it right. Skip it and every future request costs more. The two paths look identical for about a month, and then they separate permanently.
Why small is faster for a machine
For a human, writing code is expensive and finding the code that already exists is cheap, so a large file is just more to read. For a model it is the other way around. Writing is nearly free and finding what already exists is the expensive part, so when the existing thing is buried in a huge file, the model does the cheap thing instead. It writes a fresh version, and now you have two, and the duplication compounds every time. Hand it a small module and the expensive part gets cheap, which is the whole game, and it is why we built OBY to give the model the structure before it writes.
AI-native is something you build, not buy
An AI-native codebase is one where the units are small, the names say what they do, and the boundaries are clean, so any one piece fits in the model's working memory with room to spare.
That is not a new idea wearing new clothes. It is the same separation of concerns good engineers always argued for, now with a second, mechanical reason to want it. Clean boundaries used to be about the next human who reads the code, and that argument always lost to the deadline. Now they are also the difference between an assistant that speeds up as it learns your system and one that bogs down as your system grows, which is a payoff the person making the deadline collects today.
Split on the way up, not at the top
The cheap moment to divide a file is while it is still small enough that dividing it is boring. Pick a line the file is not allowed to cross, five hundred is a sensible one, and treat crossing it as the signal to modularise then, not at the end of the quarter. The blast radius of a split made early is a few imports. The blast radius of the same split made late, on a file everything now depends on, is the frightening refactor I described at the top.
One prompt in the moment does it: this file is getting long, split it along its natural seams and keep the behaviour identical. It costs a few minutes while the context is still in your head. The alternative is not splitting it later, it is splitting it under pressure, on a file that has since collected six months of other people's changes.
What PIM does not do
It does not make a legacy monolith AI-native overnight, and it will not rescue a codebase nobody understands. If nobody can say what connects to what, small modules alone do not hand you the map, and you need the map first, which is a different job we have written about in the code nobody can change.
Where the seams go is still a judgement call, and it is still yours. A model will split a file wherever you point it, including in the wrong place, with the same confidence it brings to everything else. Knowing that these two functions belong together and that one belongs elsewhere is domain knowledge, and nobody has automated it, ourselves included.
Over-modularising is real, and I have done it. Forty four-line files with a call graph nobody can hold in their head is worse than one clear file of six hundred lines. The goal is comprehension per unit of reading, not a small number.
PIM is how you keep a codebase from getting there, and how you stay out once you have dug free. It is a habit, not a switch. The honest part is that we are still learning where the right module boundaries sit, because the model's working memory keeps growing and the boundary moves with it. The rule of thumb is a moving target rather than a settled number.
If your AI assistant seems to get slower and less reliable the more of your codebase it sees, that is the problem this describes, and it is fixable without a rewrite. Have a look at how we take AI-assisted code to production or what a code review covers, and if you want a hand making an existing codebase easier for these tools to work in, get in touch or call 1300 840 340.





