Picture a leave request form. Four boxes and a submit button.
Eleven years, not one complaint.
Then someone (there is always someone) decides the form needs a chat box stapled to its face. You type need friday off, dentist, a language model squints at the sentence, guesses which Friday you meant and then fills in, I promise I am not inventing this, the same four boxes.
The demo kills.
The VP grins like a man who has just been handed next year’s roadmap slide.
Half a year later the provider retires the model.
Its successor holds private, firmly held opinions about what next Friday means and HR burns a week unpicking leave that landed on the wrong dates, while the original form, which still works flawlessly, sulks behind a grey link at the bottom of the page because somebody promoted the chat box to primary experience in a meeting nobody from HR attended.
I’d wager you’ve seen your own version.
An FAQ page that sprouted a chatbot, maybe, or an invoice parser nobody requested. Hammer, meet screw. Swing anyway.
Why everyone is swinging?
Money, ofcourse!
Just kidding.. Not just money.
Budgets this year arrived with an AI line pre-printed on them like a watermark and anyone who has survived a planning cycle knows the arithmetic of an unspent line: it evaporates and next year you get less. So teams go rummaging for somewhere, anywhere, to pour it.
While it’s slightly dated, the details are still quite valuable. Click below or here to read this document from Deloitte, the surface-level work numbers are quite revealing.
The demo finishes the job. There’s something faintly carnival about watching a model turn a rambling sentence into crisp JSON inside a fifteen-minute slot and card tricks work on exactly the same principle, by keeping your eyes off the other hand. Who in that room is thinking about the 2% of inputs where the date quietly goes sideways? Who’s pricing a million calls a month? Not the person whose promotion packet now reads “added AI to onboarding.” Can’t say I blame them. I’d be tempted too.
Screws, mostly.
Most of what follows looks embarrassingly obvious on paper. At 4pm on a Thursday, in sprint planning, with the AI line item glaring at you from the budget sheet, it doesn’t.
Take any input that can only ever be one of a dozen values. That’s a dropdown. Maybe a keyword rule, with a human queue for the stragglers. An LLM classifier will sort most of them nicely and then, on the ones it fumbles, offer you nothing but a confident shrug.
Or the invoice trick, which baffles me every time. Teams pipe their own PDF invoices through a vision model to read the total and you want to tap someone on the shoulder and ask who printed those invoices in the first place. Their billing system did. The number has been loafing in a database table two joins away since the day it was born.
Arithmetic belongs in code.
Full stop.
Tax, interest, FX conversion, amortization schedules: a model gets them right often enough to lull you and “often enough” is a ghastly bar for figures that end up in front of a regulator.
Natural-language search over structured data is subtler and I’ll concede some ground. “Orders over $500 from Ohio last month” is a SQL query wearing a sentence costume and a plain-English box on top of it can be a real courtesy to colleagues who’d sooner eat glass than write a JOIN. Letting the model compose and execute whatever SQL takes its fancy against production, though? That’s how you earn a postmortem with your name in the first paragraph.
Email validation.
Regex.
Decades of service, zero hallucinations.
And then the prose nobody reads, which irritates me more than it probably should: meeting recaps auto-posted into a channel the whole team muted back in March, commit messages announcing “Updated files to improve functionality.” Tokens paid for in full, read by no one.
Every case above had a predictable answer before the model turned up. Now each has a probable one. Somebody, somewhere, put that on a slide and called it progress.
Debt with no ledger
Ordinary technical debt is at least honest. It lives in a repo, waiting for the rewrite everyone swears is coming and a quick squint at the commit history usually tells you whom to blame.
AI debt is shiftier.
It lurks in the prompt that exactly one engineer understands, wedged into a config file under a comment that reads, in full, “don’t touch, it works.” She’ll leave eventually; good people do. The prompt becomes folklore, recited and never understood.
It lurks in the pinned model version. Your provider retires it on their calendar, not yours and the migration nobody budgeted for lands like a parking ticket on the windscreen. Did anyone write an eval set back when there was time? Rarely. So the retest is two tired people eyeballing outputs at 11pm, hoping.
The bill creeps.
Per-token pricing feels like loose change during a pilot, but production traffic has a way of turning loose change into a line item finance circles in red every quarter, all for a feature a lookup table would have handled for the price of the electricity.
Then trust goes, last and quietest. A tool only needs to be confidently wrong a handful of times before everyone starts double-checking it and from that moment you’re bankrolling the automation and doing the manual work anyway, which is the worst of both worlds and, in my experience, depressingly common.
Nor is this a software-company affliction. My hunch (and it’s only a hunch) is that a fat slice of the internal chatbots perched atop policy PDFs at banks and insurers would have served their users better as a half-decent search index. A few will earn their keep. The rest will be expensive to unbolt.
Before you pick up the hammer
Could a dull rule or a lookup table do it? Ship that. You can always bolt a model on later, whereas prying one off after launch has swallowed entire quarters.
What happens when it’s wrong? If the honest answer involves a customer receiving the wrong refund, you need a deterministic path or a human in the loop. Probably both.
Who owns the prompt a year from now? A name, please. Not a team.
How will you know it’s degrading? No evals, no launch. I won’t budge on this one.
What does it cost at a hundred times today’s traffic? Run that sum before the pilot, while “no” is still a cheap word.
Plenty of features will clear all five and they should be built, gladly. Language models are astonishingly good with messy, open-ended input, the kind where the alternative is some poor soul reading every message by hand until their eyes glaze over. Those are real nails. There are heaps of them and I’d much rather the budget went there.
The boring version
Lately, whenever someone pitches something clever, I try to ask one slightly annoying question: what does the boring version of what you are presenting & pitching look like?
Usually it ships sooner.
Usually it breaks less.
And every so often (rarely, but it happens) the boring version simply can’t cope, at which point you reach for the clever tool carrying a reason you’d happily defend in writing.
AI is a superb hammer.
But, still a hammer.



