The reporting on AI returns has gotten less flattering this year, and I think that is healthy. BCG, Grant Thornton, and others have published survey work with a similar shape: enormous investment, near universal executive conviction, and a small minority of finance leaders who can point to measurable return. One figure stuck with me. Only about 14 percent of CFOs report measurable ROI so far, while two thirds expect significant impact within two years. Conviction is running well ahead of evidence.The usual explanation is that companies are stuck in pilots. That is true, but it is shallow. I have been implementing this stuff inside real operating businesses, recruiting, home services, security, investigations, robotics service, and the pattern I keep running into is more specific than pilot purgatory.AI made the first draft free. Nobody budgeted for the second one.
Take any workflow you have automated. A job description, a service quote, a client summary, a board report, a piece of code. Before, one person spent two hours producing it and their name was on it. Now the draft appears in four minutes and it is roughly 80 percent right.That sounds like a clean win, and often it is. But 80 percent right is a different kind of artifact than a human first draft. A person who does not know a subject writes a draft that visibly signals uncertainty. A model writes the weak 20 percent in exactly the same confident register as the strong 80 percent. The errors are not flagged. They are camouflaged.So the work shifts from production to verification. Verification is harder to do well, harder to supervise, less satisfying, and almost never assigned to anyone specifically. It just lands on whoever is downstream, which is usually a senior person whose time you were trying to protect in the first place.That is how you end up with a company where everyone reports being faster and nothing measurable improved.
I see this go wrong in two directions.The first is rubber stamping. Nobody owns the check, the output looks polished, and it goes out. The mistakes it contains are not random typos. They are plausible, specific, and wrong, which is the worst possible combination. You find out from a customer.The second is shadow rework. A senior person quietly rewrites the output every time because they do not trust it, and they never say so, because the tool was announced as a win. Your reported savings are fiction. Your best person is now doing editing instead of judgment.Both failures share a root cause. The tool got deployed. The standard did not.
A few things, none of them clever.
AI is the most useful set of tools I have had access to in thirty years of running operations. I am not skeptical about the technology. I am skeptical about the assumption that buying it changes how work gets done.Automation exposes your operating discipline. If your standards are written down, your ownership is clear, and your review process matches your risk, AI compounds all of it quickly. If those things live in the heads of a few good people, you will get a lot of confident output that those same few people quietly fix at night.The companies that show real return over the next two years will not be the ones with the best models. Everyone will have roughly the same models. They will be the ones that did the boring work of deciding what good looks like and who is responsible for it.
Connect with Steve: linkedin.com/in/stevepurban