← Blog

What an AI harness is, and why it decides what the model can do

The model is a brain in a jar. Everything it does in your business, reading files, taking actions, checking its own work, stopping when it should, comes from the harness around it. Here is what a harness is made of and why the same model behaves so differently in two of them.

When a firm says "we put AI on it," the thing they are picturing is a model. The thing that actually did the work was a harness. The model on its own cannot read a file, open an inbox, run a test or send anything. It takes text in and gives text out. Everything else, which is everything that matters in a business, is the harness.

This is why Claude Code and Codex are not models. They are harnesses around models. Put the same model in two different harnesses and it performs differently, sometimes by a lot. Most of the capability jumps of the last year came from better harnesses around models that had not changed much.

The model on its own
  • Takes text in, gives text out
  • Starts every conversation blank
  • Cannot see your files, inbox or screen
  • Can tell you what to do, not do it
  • Says its own work looks fine
  • Has no way to stop itself
The model in a harness
  • Reads and acts in your systems through tools
  • Carries memory and a compact history of the work
  • Starts with context: the project, the rules, the recent changes
  • Runs the action and reads the result
  • Checks the output with tests, linters, calculators
  • Asks before anything risky, inside a sandbox

The five parts

WHAT THE MODEL SEESTHE MODELWHAT IT CAN DOWHAT COMES BACKContextthe task, the rules, the file mapMemorycompacted history, a change logThe last resultwhat the previous step returnedPermissionswhat it may and may not doThe modelreasonsdecides one steprepeatsToolsread, search, edit, run, sendVerificationtests, checks, a calculatorPermission gatehooks that ask before risky stepsSandboxa place to act where damage stays containedA changed file or recordin the system you already runA test resultfed straight back to the modelA question for youwhen it should not decide aloneA trimmed resultonly the useful part kept
One turn of the loop. The model sees context, memory and the last result, picks one action, the harness runs it under permissions, and the trimmed outcome goes back in.

Context is what the model knows when the session starts: what project this is, how the files are laid out, what changed recently, what tools it has, and the house rules. In a coding harness that lives in a file at the root of the project that says things like "use our code style, do not touch this folder, run these tests." You write it once and it stops you from re-explaining yourself in every prompt. Most of the difference between a frustrating agent and a useful one is whether anyone wrote this down.

Memory solves a limit the model was born with. It has a fixed window, and when a long piece of work fills it, two things happen: the model gets a little less sharp, and old instructions start to contradict new ones. A good harness compacts the history into a dense summary, keeps a change log of edited files rather than the files themselves, and lets the work continue past the window in the same direction.

Tools are how text becomes action. The model emits a specific shape of text, the harness recognizes it as a request, matches it to a tool, runs the tool, and feeds the result back. Dozens of these happen a minute without the person watching. The standard shape for plugging in new tools is called MCP, the Model Context Protocol, and it is why Claude Code can reach a CRM, a practice management system or an accounting ledger the same way it reaches a file.

Verification is the part people skip. Ask a model whether its own output is right and it will usually say yes. Hand the output to something that cannot flatter it, a test suite, a linter, a schema check, a calculator, and feed the result back, and the model becomes able to do things it was unreliable at. The early story of models failing at arithmetic ended the moment someone gave them a calculator and told them to use it.

Permissions and the sandbox are the stop conditions. Before a risky action, send data to an unfamiliar site, delete a folder, change a setting, the harness pauses and asks. The mechanism is usually a hook: a small piece of code attached to a point in the loop, such as "after every edit, run the tests" or "before every outbound request, check it against the list of places we have sent data before." The sandbox is where the model acts so that if it does something wrong, the damage stays inside a line you drew.

The loop

01

Think

The model reads context, memory and the last result and decides one next step. Some harnesses make it argue with itself first, which is what reasoning is.

02

Act

It asks for a tool: read this, search that, edit this file, run this command.

03

Check the permission

The harness decides whether that action is allowed, or asks you.

04

Run it, contained

The action runs, usually in a sandbox, so the blast radius is known.

05

Trim and read

The harness keeps the useful part of the result and hands it back. The model reads what happened and why.

06

Repeat until done

Then exit. This is the loop most agents have run since the first one, under the name ReAct, reason plus act.

Two consequences follow from this. First, a harness cannot make a model smarter, but it can make it far more useful, and that is where almost all of the recent progress has come from. Second, harnesses are software. A model with billions of parameters is something only a handful of labs can change. The harness around it is something a small team can change on a Tuesday afternoon.

What this means for a firm

Every system we build for a client is a harness around a model, fitted to one workflow. The model is the least of it. The work is the five parts, done for that firm.

5
parts in every harness
1
workflow per harness, to start
0
actions taken without a person or a check

Walk through the demos on this site with that lens and you will see the same five parts every time, in the firm's own vocabulary:

  • Context is the playbook in the contract review demo: the firm's positions and fallbacks, written down once, and every incoming contract is read against them.
  • Memory is the live record in the property operations center: every message, work order and vendor confirmation in one place, so the morning briefing is drafted from what actually happened.
  • Tools are the connections in the AP inbox: invoices read on arrival, matched to purchase orders and receipts, posted to the ledger only after approval.
  • Verification is the source chip on every sentence. The knowledge assistant answers only from the firm's documents, shows the passage behind each sentence, and says when the documents do not cover a question. The VC deal screener traces every claim to the page it came from and catches a deck built to mislead.
  • Permissions are the review queue in the tax prep autopilot and the approval step in every other demo. The system drafts, flags and chases. A person approves, signs, sends.

This is also why the AI readiness assessment scores "AI ownership" as one of its four areas. The context file, the permission rules and the verification checks need an owner inside the firm. Without one, the harness is whatever the last person who touched it left behind.

Where harnesses go next

The direction is already visible. MCP is showing up in most major platforms, so the tools a harness can reach keep widening without custom work. Teams of agents are being built into harnesses natively, with a shared layer of what every agent learned on similar problems. Harnesses are starting to run a task several times in parallel, each with its own verification loop, and keep the attempt that passed. And the newest idea is a harness that watches its own mistakes, works out which step wasted tokens or went wrong, and tightens itself for next time.

None of that requires a smarter model. All of it requires someone who understands the five parts and can fit them to a real workflow. That is the work.

See which workflow your firm would harness first

Want this applied to your firm?

Tell us what you have in mind. We'll come back with the highest-ROI path to get there.

Book a free callor write to us at [email protected]