A pattern we see in almost every firm that has "tried AI and it disappointed." The pilot was built. The demo went well. Everyone in the room agreed it was impressive. Six weeks later nobody is using it, and the lesson the firm took away is that AI is not ready.
The model was rarely the problem. Two decisions were missing, and both should have been made before anyone wrote a line of code.
Decision one: who owns it
A system without an owner gets switched off the first time it does something odd. That is not a criticism of the people; it is how organisations work. When an agent chases a client for a document and the client complains, someone has to decide whether the agent was right. If nobody is that someone, the agent gets turned off and the complaint ends there.
The owner is not the sponsor who approved the budget and not the vendor who built it. It is a person inside the firm, with a name, with a set number of hours a month, who is responsible for the system working. They approve the uncertain cases in the first weeks. They decide when the agent is trusted enough to act alone. They are the one the complaint goes to.
Naming that person is the first thing we do in any engagement, and it is the thing most pilots skipped. We will not start a build without it. If there is nobody who can take it, that is useful to know before the money is spent.
Decision two: what done looks like
The second missing decision is a definition of done. Not "it works," which is a feeling, but a measurement that was taken before the build and will be taken again after.
For most workflows the measurement is hours. How many hours a month does the reporting cycle take today, counted, not guessed. How many hours per matter on first-pass drafting. How long from a lead arriving to a human responding. The number before is the baseline. The number after is the result. The gap is the definition of done, written down and agreed by the owner before work starts.
Pilots without this have no way to succeed. There is no number to hit, so the system is judged on the demo, and demos are judged on the most recent impressive or embarrassing thing it did.
Why demos are the wrong test
A demo shows the system on a case chosen to show it well. Real work is the cases nobody chose. The document laid out differently. The client who wrote back in capital letters. The matter that is two matters.
The right test is an evaluation set: real past cases, as many as the firm can spare, with the correct answer recorded, and a pass bar the owner agreed to. The system runs against that set before anyone sees it, and again on every change. When it passes, it runs alongside the manual process for a period, and the owner compares the two. Only then does it act on its own.
That is slower than a demo. It is also the difference between a system that is still running a year later and one that is a story about how AI is not ready.
What this means for a first project
Scope one workflow
The one with the clearest before and after, not the most interesting one.
Name the owner
A person inside the firm, with hours a month, before anything is built.
Count, then agree the bar
Hours today, counted. The pass bar on real past cases, agreed in writing.
Parallel run, then measure
Alongside the manual process until the owner says it has earned trust. Then the number again.
Scope it to one workflow. Name the owner first. Count the hours before you start. Agree the pass bar. Run it alongside the manual process until the owner says it has earned trust. Then measure again.
None of that is about the model. All of it is about the firm. Which is why the second attempt at AI in a firm that did it this way almost always goes differently from the first.
If you have a pilot that died and want a view on whether it is worth reviving, we are happy to look. Most of them are, once the two decisions get made.