When a firm says "we put AI on it," the thing they are picturing is a model. The thing that actually did the work was a harness. The model on its own cannot read a file, open an inbox, run a test or send anything. It takes text in and gives text out. Everything else, which is everything that matters in a business, is the harness.
This is why Claude Code and Codex are not models. They are harnesses around models. Put the same model in two different harnesses and it performs differently, sometimes by a lot. Most of the capability jumps of the last year came from better harnesses around models that had not changed much.
- Takes text in, gives text out
- Starts every conversation blank
- Cannot see your files, inbox or screen
- Can tell you what to do, not do it
- Says its own work looks fine
- Has no way to stop itself
- Reads and acts in your systems through tools
- Carries memory and a compact history of the work
- Starts with context: the project, the rules, the recent changes
- Runs the action and reads the result
- Checks the output with tests, linters, calculators
- Asks before anything risky, inside a sandbox
The five parts
Context is what the model knows when the session starts: what project this is, how the files are laid out, what changed recently, what tools it has, and the house rules. In a coding harness that lives in a file at the root of the project that says things like "use our code style, do not touch this folder, run these tests." You write it once and it stops you from re-explaining yourself in every prompt. Most of the difference between a frustrating agent and a useful one is whether anyone wrote this down.
Memory solves a limit the model was born with. It has a fixed window, and when a long piece of work fills it, two things happen: the model gets a little less sharp, and old instructions start to contradict new ones. A good harness compacts the history into a dense summary, keeps a change log of edited files rather than the files themselves, and lets the work continue past the window in the same direction.
Tools are how text becomes action. The model emits a specific shape of text, the harness recognizes it as a request, matches it to a tool, runs the tool, and feeds the result back. Dozens of these happen a minute without the person watching. The standard shape for plugging in new tools is called MCP, the Model Context Protocol, and it is why Claude Code can reach a CRM, a practice management system or an accounting ledger the same way it reaches a file.
Verification is the part people skip. Ask a model whether its own output is right and it will usually say yes. Hand the output to something that cannot flatter it, a test suite, a linter, a schema check, a calculator, and feed the result back, and the model becomes able to do things it was unreliable at. The early story of models failing at arithmetic ended the moment someone gave them a calculator and told them to use it.
Permissions and the sandbox are the stop conditions. Before a risky action, send data to an unfamiliar site, delete a folder, change a setting, the harness pauses and asks. The mechanism is usually a hook: a small piece of code attached to a point in the loop, such as "after every edit, run the tests" or "before every outbound request, check it against the list of places we have sent data before." The sandbox is where the model acts so that if it does something wrong, the damage stays inside a line you drew.
The loop
Think
The model reads context, memory and the last result and decides one next step. Some harnesses make it argue with itself first, which is what reasoning is.
Act
It asks for a tool: read this, search that, edit this file, run this command.
Check the permission
The harness decides whether that action is allowed, or asks you.
Run it, contained
The action runs, usually in a sandbox, so the blast radius is known.
Trim and read
The harness keeps the useful part of the result and hands it back. The model reads what happened and why.
Repeat until done
Then exit. This is the loop most agents have run since the first one, under the name ReAct, reason plus act.
Two consequences follow from this. First, a harness cannot make a model smarter, but it can make it far more useful, and that is where almost all of the recent progress has come from. Second, harnesses are software. A model with billions of parameters is something only a handful of labs can change. The harness around it is something a small team can change on a Tuesday afternoon.
What this means for a firm
Every system we build for a client is a harness around a model, fitted to one workflow. The model is the least of it. The work is the five parts, done for that firm.
Walk through the demos on this site with that lens and you will see the same five parts every time, in the firm's own vocabulary:
- Context is the playbook in the contract review demo: the firm's positions and fallbacks, written down once, and every incoming contract is read against them.
- Memory is the live record in the property operations center: every message, work order and vendor confirmation in one place, so the morning briefing is drafted from what actually happened.
- Tools are the connections in the AP inbox: invoices read on arrival, matched to purchase orders and receipts, posted to the ledger only after approval.
- Verification is the source chip on every sentence. The knowledge assistant answers only from the firm's documents, shows the passage behind each sentence, and says when the documents do not cover a question. The VC deal screener traces every claim to the page it came from and catches a deck built to mislead.
- Permissions are the review queue in the tax prep autopilot and the approval step in every other demo. The system drafts, flags and chases. A person approves, signs, sends.
This is also why the AI readiness assessment scores "AI ownership" as one of its four areas. The context file, the permission rules and the verification checks need an owner inside the firm. Without one, the harness is whatever the last person who touched it left behind.
Where harnesses go next
The direction is already visible. MCP is showing up in most major platforms, so the tools a harness can reach keep widening without custom work. Teams of agents are being built into harnesses natively, with a shared layer of what every agent learned on similar problems. Harnesses are starting to run a task several times in parallel, each with its own verification loop, and keep the attempt that passed. And the newest idea is a harness that watches its own mistakes, works out which step wasted tokens or went wrong, and tightens itself for next time.
None of that requires a smarter model. All of it requires someone who understands the five parts and can fit them to a real workflow. That is the work.