Internal tools that use a model to do one narrow job.
- AI is worth using where a specific, repeated task is being done by somebody who would rather not be doing it, and where being wrong occasionally is survivable. Most problems in a business are neither. This page is about telling the two apart.
The evidence is not what the market says it is.
There is more confident claiming about AI than about anything else we are asked for, and very little of it is measured. What has been measured is worth knowing before you spend anything.
Measured slower, felt faster
In a randomised trial, 16 experienced developers took 19% longer on 246 real tasks using early-2025 AI tools, while estimating afterwards that they had been 20% faster. METR, 2025. The perception gap is the finding.
Almost right is the problem
The top frustration reported by developers using AI, cited by 66%, is solutions that are almost right but not quite. Second, at 45.2%, is that debugging AI-written code takes longer. Stack Overflow Developer Survey 2025.
The experienced are the most sceptical
84% of developers use or plan to use AI tools, yet more distrust its accuracy (46%) than trust it (33%), and the most experienced are the most distrustful. Stack Overflow, 2025, more than 49,000 respondents.
Believed faster. Measured slower.
This is the most useful piece of evidence about AI at work, because it is randomised, it measures outcomes rather than opinions, and it shows the gap between the two directly.
It is not an argument against using AI. It is an argument for measuring whether a given use actually helped, which almost nobody does.
When an AI tool is the answer, and when it is not.
The useful question is never whether to use AI. It is whether this particular task has the shape that suits one.
Worth doing when
- A specific task is repeated many times a week by somebody senior.
- The input is messy text and the output is a draft a human checks.
- Being wrong occasionally costs minutes, not money or safety.
- You can state what working means before anything is built.
Worth postponing when
- The task happens twice a month.
- Nobody can say what the tool would do, only that it should use AI.
- The output goes straight to a client with nobody reading it.
- The underlying process is broken and the model is being asked to hide it.
What actually moves the number.
Four things move the cost, and none of them is the model.
What people ask before they commit.
The ones that come up most, answered plainly.
- A written answer, not a sales proposal
- What is wrong, and what to fix first
- Yours whether or not you build with us