About a year ago I wrote a post about AI wrapper applications. My argument was that the software layer between users and AI models would dominate for the next three to five years, and then probably fade away as models got smart enough to do the work natively.
The first half of that prediction held up. The second half didn't, and the reason why is worth a whole post. Because I like to dwell on my mistakes... Just kidding!
The biggest change since then is that the industry has settled on a different word for that layer: the "harness." It might sound like a rebrand, and in some ways it is. But the change in vocabulary carries a real change in the overall argument. A "wrapper" implies a shell around something that's already complete. A "harness" implies structure that the thing can't perform without. After a year of building agentic systems for enterprise clients, I think the second framing is now the accurate one.
(And no, there's no rapper image this time. I've matured... somewhat.)
What Is an AI Harness?
An AI harness is everything around the model call that turns a language model into a working system: the loop that decides what happens next, the tools the model can use, the context it's given, the checks on its output, the permissions that constrain it, and the machinery that recovers when something fails.
One quick disambiguation, because it confuses almost every client conversation I have. You'll hear "harness" used in two ways. An agent harness is the runtime scaffolding I just described. An eval harness is a test rig used to measure how well a model or system performs. They're related, since a good agent harness includes evals, but when I say harness in this post, I mean the first one.
Why the Term Caught On
Here's the observation that changed how the industry talks: put the same model into two different harnesses and you can get dramatically different results. Not slightly different. Different enough that the harness sometimes matters as much as the model choice.
AI coding tools made this impossible to ignore. Developers could watch the same underlying model succeed in one tool and flounder in another, and the difference came down to how each tool managed context, which tools it exposed, and how it checked its own work. It's now standard for benchmark results to name the harness alongside the model, because a score without the harness is close to meaningless.
For decision-makers, the takeaway is simple: you aren't buying a model. You're buying the thing built around it. That's where the real differences between vendors and approaches hide.
Wrapper vs. Harness: Which Do You Actually Need?
Before going further, I want to be clear that plenty of projects still only need a wrapper, and anyone who tells you otherwise is probably selling you something.
You likely need a wrapper if:
- A user asks a question or submits content, and the system returns one formatted answer
- A human decides every next step
- The AI reads from your systems but doesn't change anything in them
You likely need a harness if:
- The model decides what to do next, across multiple steps
- The AI takes actions in systems of record — updating a CRM, sending messages, modifying files
- The system runs unattended, or runs long enough that failures mid-process are a real possibility
The common mistake I see is buying harness-level complexity for a wrapper-level problem. The more expensive mistake is the reverse: shipping a wrapper and expecting it to behave like an agent. That's usually where the "our AI pilot didn't work" stories come from. Sorry, but a custom GPT does not a pilot make.
The Anatomy of a Harness
This is where the work actually lives. Each of these components is something a well-built harness handles deliberately, and each one has a recognizable failure mode when it's missing.
The Loop and Its Stopping Conditions
At its core, an agentic system is a loop: the model looks at the situation, decides on an action, takes it, observes the result, and decides again. The harness runs that loop, and critically, it decides when the loop ends.
When it's missing: agents that never finish, or finish too early and declare victory. Both are common, and both are expensive.
Tool Design
Tools are how the model touches the world — searching a database, calling an API, reading a document. How those tools are designed has an outsized effect on how well the model uses them. Fewer, clearer tools almost always beat many overlapping ones. Tool names and descriptions are effectively prompts. And error messages matter more than people expect, because a good error message tells the model how to fix its mistake and try again.
When it's missing: the model calls the wrong tool, calls the right tool with bad inputs, or gives up after a single failure that a clearer error would have let it recover from.
Context Engineering
In my wrapper post, I described retrieval-augmented generation as a way to give the model access to your data. That's still true, but the pattern has shifted. Instead of retrieving documents up front and stuffing them into the prompt, the stronger approach now is to give the agent retrieval as a tool and let it search for what it needs, when it needs it.
Context engineering is the broader discipline: deciding what the model sees at each step, what gets summarized or dropped as a task runs long, and how to keep the model focused on what matters. Models have limited working memory, and a harness that fills it with noise gets worse answers.
When it's missing: long-running tasks that slowly degrade, agents that "forget" instructions partway through, and inflated costs from sending the same irrelevant material over and over.
Verification
How does the system know it succeeded? This is the question I push clients on hardest, because it's the one they've usually thought about least. Verification can be automated tests, structured output checks, comparison against business rules, a second model reviewing the first, or a human approval step. What matters is that success is defined in terms of your standards, not the model's confidence.
When it's missing: output that looks right, reads confidently, and is wrong. This is the failure that erodes trust fastest.
Durability and Recovery
Agentic workflows can run for minutes or hours and touch multiple systems. Things will fail along the way: an API times out, a service goes down, a model returns something malformed. A durable harness checkpoints progress, retries intelligently, and can resume from where it left off rather than starting over.
When it's missing: a failure at step nine of ten means rerunning all ten, paying for all ten, and hoping the model makes the same decisions the second time.
Permissions and Guardrails
What is the agent allowed to do without asking? What requires human approval? What is it never allowed to do? In my wrapper post I talked about guardrails as a nice benefit. In a harness, they're foundational. An agent with write access to your systems needs clearly scoped permissions, sandboxed execution where appropriate, and approval gates on consequential actions.
When it's missing: you find out what the agent was capable of after it did it.
Observability and Evals
You can't improve what you can't see. A production harness logs what the agent did and why, traces each step, and runs evaluations — that's the other sense of "harness" — so you can measure whether a prompt change, tool change, or model upgrade made things better or worse.
When it's missing: every change is a guess, and every model upgrade is a leap of faith.
What This Means for Cost and Timeline
Here's where I need to correct myself. In the wrapper post, I wrote that these applications "are not as complex to develop as many assume." For single-step wrapper apps, that's still true. For agentic systems, it isn't.
The user interface is still ordinary web development, and AI coding tools have made that part faster than ever. The harness is different. The reliability layer — verification, recovery, permissions, context management, evals — is where the hours go, and it's what clients most consistently underestimate. A demo that works in a meeting might represent a small fraction of the work required to make it work unattended on a Tuesday at 3 a.m.
That's not a reason to avoid agentic systems. It's a reason to budget for them honestly.
Model Portability, Revisited
In the wrapper post, I said that on a platform like Amazon Bedrock, changing models could be "as easy as changing a single line of code." I also wrote a whole post arguing for treating LLMs as interchangeable components. Well, I still stand by the architecture, but I need to walk back the ease of effort a bit.
Models are now trained heavily on how to use tools, how to reason through multi-step tasks, and how to respond to particular prompting patterns, and each provider does this a little differently. That means your harness gets tuned, whether you intend it or not, to the model it runs on. Swapping the model is still a line of code. Getting the system back to the same level of reliability means revisiting prompts, tool descriptions, and evals.
So, you have to build the abstraction layer: It still gives you leverage in negotiations, flexibility on compliance, and a path to better models. Just plan for a model switch to be a tuning and testing exercise rather than a config change. And this is one more reason evals matter: they're what turn a model switch from a gamble into a measurement.
The part of my old argument that aged best is the fine-tuning warning. If harness tuning creates switching costs, fine-tuning multiplies them.
Will the Harness Go Away?
In the wrapper post I said it seemed inevitable that wrappers would become redundant once models could do everything themselves, including generate their own interfaces. Some of that is happening. Models now handle a lot of what thin wrapper apps used to provide, and generated interfaces are real.
But the harness isn't going away, because most of what it does isn't about the model's intelligence. It's about your organization:
Permissions reflect your policies, not the model's judgment
Integration connects to your systems, which the model doesn't own
Verification checks against your definition of done
Audit trails satisfy your regulators
Cost control protects your budget
As models improve, the harness thins in some places: less hand-holding, fewer workarounds for model weaknesses. It thickens in others, because more capable agents get trusted with more consequential work, and consequential work needs more structure around it, not less.
Wrapping Up (One More Time)
A year ago, I called wrappers "in the right place at the right time." Harnesses are something different: not a stopgap until the models catch up, but a durable layer of engineering that determines whether an AI system works reliably in the real world.
If you're evaluating an agentic AI project, the most useful question you can ask isn't "which model are you using?" It's "show me the harness." How does the system decide what to do next? How does it know when it's done? What happens when something fails? What is it allowed to do without asking?
If your team or vendor has good answers to those questions, you're in good shape. If they don't, that's where the work is.