Build the AI a Tool. Do Not Feed It Your Codebase.

"Right, I am going to make you a tool."

I said that out loud to a machine, a few years into working this way. It sounds ridiculous written down. It is also the single most useful sentence I have found for getting real work out of these systems, and almost nobody does it.

What everybody does instead

You have a question about your system, so you give the model your system.

The whole repository, or the directory you think is relevant, or a year of exported data, pasted in and handed over. It is the obvious move. It is also expensive twice over. You pay for every token going in, and you get a worse answer than a narrower question would have produced, because what you needed is competing for room with forty files that were not relevant.

The instinct behind it is treating context as storage. It is not storage. It is closer to working memory, and the useful question is not how much can I fit but what does this need in front of it right now.

The alternative is smaller than you expect

Build the model a tool that answers the specific question, and let it call the tool.

Not a platform. A small program that does one job, returns a machine-readable answer, and finishes fast. The model fetches what it needs at the moment it needs it, uses it and lets it go. Nothing is held that is not being used.

We have a lot of these. A sitemap checker that walks a site and reports dead links and orphan pages, which has been running for about two and a half years and still gets used most weeks. An importer built to move a client from Weebly to WordPress that took eight thousand pages across, structure intact. There is OBY, which answers what calls this function deterministically instead of guessing. There is also CodeScan, which answers whether a vulnerability is reachable from code that runs in production.

Each one replaced a task that used to mean handing over an enormous amount of material and hoping. None of them is clever. That is the point: a tool does not need to be clever if it is correct, because the cleverness is supposed to be in whatever calls it.

The charts example, because it is the clearest one

Somebody wants a chart from a spreadsheet, so they hand the spreadsheet to a model that generates images and ask for a chart.

What happens is that your data becomes pixels. An image model burns a serious amount of compute rendering a picture of a bar chart, and what comes back cannot be edited, cannot be checked against the numbers, and often has the labels subtly wrong. Do it fifty times and you have paid for fifty renderings of the same chart.

Or you write a chart tool once. It takes the CSV, it produces the chart, it costs almost nothing per run, and the output is right because it was drawn from the data rather than imagined from a description of the data. That took an afternoon. It has run thousands of times since.

The difference between those two approaches is not really about charts. It is the same choice showing up wherever people ask a general-purpose system to do a specific mechanical job.

Where this puts us on retrieval databases

The popular version of the same problem is to vectorise everything, put it in a database, and retrieve chunks into context on demand.

Our position, and it is a position rather than a settled fact: that is the expensive route to somewhere you can usually reach more directly. You take content you already have in a structured and queryable form, convert it into something lossy and approximate, then pay to search the approximation. If the underlying source can be queried precisely, query it precisely. Code has a parser. Databases have indexes. Filesystems have paths. A tool that uses those gives you an exact answer for a fraction of the cost, and it gives you the same answer twice, which matters more than people expect.

There are genuine cases for retrieval over embeddings, mostly involving large volumes of unstructured prose with no other index. Your codebase is not that. Your database is definitely not that.

What makes a good tool for an AI to use

Five properties, and the third is the one people miss.

One job. A tool that does four things needs a decision before it can be called, and decisions are where the guessing comes back in.

Fast. If the informed path is slower than guessing, a model will guess, not out of laziness but because it is tuned to complete. Speed is not a nice-to-have here, it is the mechanism.

Deterministic. Same question, same answer, every time. A tool that is right most of the time is worse than no tool, because you will act on the wrong one.

Machine-readable output. Structured, parseable, small. Prose is for people.

Honest failure. When it cannot answer, it says so plainly rather than returning something empty that reads like an answer. Silence and zero results are different, and a tool that cannot tell you which one it means will be trusted at exactly the wrong moment.

The part nobody mentions

I have built something like a hundred of these over the years. Some of them are excellent. A number of them have quietly rotted, because a tool is a thing you maintain, and a broken tool that still returns output is worse than nothing at all.

That is the honest cost of this approach. You are not avoiding work, you are moving it: from paying repeatedly for enormous context, to building and keeping a small library of tools that work. For a one-off question that trade is not worth it. For anything you will ask more than about five times, it stops being close.

The test is simple enough. If you find yourself pasting the same category of thing into a model for the third time, stop pasting, and build it a tool instead.


If your team is spending real money handing whole systems to a model and getting confident and unreliable answers back, that is a tooling problem with an ordinary fix. We build this way on every engagement, and both OBY and taking AI-assisted code to production come out of exactly that. Get in touch or call 1300 840 340.

Similar Posts