Building better AI agents with harness-driven loops
Originally published on Medium
In this article

Harness is becoming the defining idea in agentic engineering this year. A useful shorthand is: Agent = Model + Harness.
TL;DR:
A fixed pipeline worked until users started changing direction mid-task.
A loop lets the model choose its next action and verify the result.
A harness keeps that loop safe with prompts, tools, memory, skills, sub-agents and policy.
The trade-off is token cost, so memory management matters.
When I started building https://www.nikistudio.cloud - my AI design product, I gave it two agents: one designs ad creatives and campaigns, the other builds slide decks.
I built an orchestrator around them. Each request moved through fixed stages: understand the request, collect missing details, make a plan, get the user’s approval, execute, and review. Each stage decided, in code, what the agent was allowed to do next.
To be fair to that design, it had real strengths. It was predictable, easy to reason about, and cheap to run. It got Niki Studio through its first users. But as the product grew, two limits kept showing up, and no amount of patching made them go away.
Limit one: real conversations don’t move in straight lines
Users change direction mid-task. Someone tweaking a single design suddenly asks for a full campaign. Someone reviewing a plan wants to jump back and change the brief. Each of those jumps is a transition between stages, and in a pipeline, every transition has to be wired by hand.
I responded the obvious way. Every new capability meant more stages, more transitions between them, and more state to keep consistent across two agents. The file managing all of it grew to 1,351 lines. It owned routing, prompt assembly, tool calls, state, and approval rules at once. I knew every line, and I was still careful about touching it, because a change for the design agent could ripple into the slides agent’s states.
Limit two: a pipeline never looks back
The second limit is quieter but deeper. A pipeline moves forward. Each stage finishes its job, hands off, and stops at its boundary. Nothing in that structure lets the agent look at its own output, decide it isn’t good enough, and go back to improve it. Quality was whatever the execute stage happened to produce on its first pass.
The gap between my agents and the agents I use
Here’s what made the problem impossible to ignore. The tools I build with every day, Claude Code and Codex, don’t behave anything like my pipeline.
You hand them a vague, messy problem. They look around the codebase, form a plan if the task deserves one, ask you a question only when they’re genuinely stuck, and then carry the work end to end, checking their own results as they go. Nobody scripted their path. Users of Niki Studio were starting to expect that same intelligence, because those tools set the bar.
So I studied how these systems are actually put together. The pattern I found rebuilt my entire backend: a loop inside a harness.
1. The loop: the model picks its own next move
The old design encoded the next move in stage transitions. The loop gives that choice to the model.
Each cycle, the model sees the goal, the current state, and its available tools. It chooses one next action: ask the user something, run a tool, or verify its output against the goal.
That verify step is the heart of the idea. When the check passes, the work is done. When it doesn’t, the finding feeds back in and the agent iterates.
That is where a sequence becomes a loop. My pipeline stopped at every stage boundary. A loop stops when the work is good enough or a budget runs out.
It also addresses my first limit. When a user changes direction mid-task, a reasoning loop weighs the new instruction and adjusts its next step. The path emerges from reasoning instead of living in my if-else branches.
2. The harness: everything wrapped around the loop
A loop alone is a liability. An unconstrained model with tools can wander, overspend, or confidently do the wrong thing. The harness is everything you build around the loop to make it safe and genuinely capable.
In plain terms, the harness answers six questions about the agent:
Prompts: who is the agent in this mode of work? The instructions need to agree with each other. I learned the hard way that a prompt telling the model to ignore its own earlier instructions is an architecture bug wearing a costume.
Tools: what can it do? Generate an image, extract editable layers, modify the canvas, search stock media. Each one is a clean contract, so the model reasons about what to do while code handles how.
Memory: what survives past one reply? My agent tracks its state, plan, progress, and asset identities in structured form, separate from chat history. Conversation text is a terrible database.
Skills: what expertise does it pull in when the task needs it? A wedding-invitation request can load a design formula the everyday loop doesn’t carry. The base agent stays lean.
Sub-agents: who can take a focused piece of work? Each worker runs with its own narrow context and returns just the result, so the main agent stays clear-headed. My canvas composer already behaves like one, and more are coming. This is the same pattern Claude Code uses when it spins up subagents for a search or a review.
Policy: what is it not allowed to do? This is the set of rules the model can’t talk its way around: which tools exist in this mode, when approval is required, how many attempts a repair gets, and how much a turn may spend. The model proposes; the policy layer has the veto.
That last split became my one-line summary of the whole rebuild: the model chooses the path, the harness holds the boundaries.
What it costs
I want to be honest about the drawback, because nobody selling you “agentic” mentions it.
Tokens. A reasoning loop re-reads its context on every step: the goal, the state, the last tool result. A ten-step turn costs far more than my old pipeline’s single scripted pass. Most of my ongoing work goes into memory: summarizing older context, keeping recent detail raw, and keeping each step’s context small. The harness makes the agent smart. Memory management keeps it affordable.
Why I think this pattern is the one to learn
This isn’t just my product’s story. Look at the tools defining this year: Claude Code, Codex, OpenClaw, OpenCode, Antigravity, Manus. Different companies, different products, and underneath, the same anatomy: a reasoning loop wrapped in a harness of tools, memory, skills, and policy. The frontier models inside them have largely converged in capability.
The harness is where these products now compete, which is why people have started calling 2026 the year of the harness the way 2025 was the year of the agent.
And the loops keep getting longer. Independent researchers who track this measured that the length of a task an AI agent can finish on its own has been doubling every few months: from tasks of a few minutes in 2024 to tasks the length of a full working day now. Every extra hour an agent can hold onto a task makes the structure around it matter more, because nobody lets an agent work unsupervised for a day on vibes.
I think the software people pay for next will be agents that own a problem end to end: a system that investigates, works, checks itself, and comes back with the job done. Companies will need people who can build the structure around the model that makes that trustworthy.
That’s what rebuilding Niki Studio taught me. The model was never the hard part. The harness was.
If you’re building an agent right now, look at two places in your design. Where does your code decide the path for the model? Every one of those spots is where a real user will eventually change direction and your system will have no route. Where does your agent check its own work? If the answer is nowhere, you don’t have a loop yet.
What part of your agent stack needs the biggest rethink: prompts, tools, memory, or the harness?