Loops explained: graph engineering does not replace loops

Originally published on Medium

In this article
  1. 1. THE PROMPT LOOP
  2. DO YOU ACTUALLY NEED ONE?
  3. 1. THE AGENT LOOP
  4. A PUBLIC EXAMPLE
  5. THE VERSION I ACTUALLY RUN
  6. THE SMALLEST LOOP YOU CAN RUN
  7. 3. THE GRAPH AROUND THE LOOPS
  8. THE HONEST PART: HUMAN IN THE LOOP
Diagram comparing a prompt loop with a production agent loop.

Graph engineering does not replace loop engineering. The graph is the outer map. The loop is the repeated work inside one node or task. Production systems often need both.

There are three places a loop can live: inside a prompt, inside an agent or around a group of agents.

I learned the difference while building loops inside Niki Studio . One agent creates an ad design, make it editable, rebuilds the canvas, repairs missing pieces and verify the result.

The first layer is a prompt loop: the model drafts, checks and improves inside one conversation. The second is an agent loop: the architecture inside tools like Codex, Claude Code, or Niki Studio keeps track of progress, acts on real work, verifies results, and decides whether to continue. The third is a graph of loops: agents, tools, validators and human approvals connected through explicit paths.

1. THE PROMPT LOOP

A prompt is one instruction used by individuals. You can use it inside Claude, Codex, etc. You ask, you get an answer, you decide what’s next. A loop is a goal the AI keeps working toward. It acts, checks the result against the goal, and feeds the finding into the next attempt.

Five steps, but only three of them decide whether you’ve built a loop or an expensive way to generate noise:

The check: this is the heart. Without a real check that can reject the work, you don’t have a loop, you have an AI agreeing with itself on repeat. The check can be a test that passes or fails, a number that has to go up, or a strict rubric the model scores against. No check means the model grades its own homework, and the model that did the work is a very generous grader.

The memory. Each pass, the AI has to know what it already tried. Without that, it repeats the same mistake every round. Real loops keep a small record on the side: what’s done, what failed, what’s next.

The stop rule. A loop with no exit runs until it succeeds, breaks, or drains your account. Every serious loop has two ways to stop: the goal is met, or a hard limit says “after N tries, stop and report.”

DO YOU ACTUALLY NEED ONE?

Here’s the part most loop content skips, because it’s less exciting: most tasks don’t need a loop.

A loop earns its cost only when all four of these are true:

  1. The task repeats. A one-off is still better served by one good prompt.

  2. Something can automatically reject bad output. If nothing can fail the work without you in the room, you’re back to reviewing everything yourself, which is the job the loop was supposed to remove.

  3. Your budget can absorb the waste. Loops re-read context and retry. That costs tokens whether the run ships anything or not.

  4. “Done” is objective. If quality is a matter of taste, a human still wins.

Miss one, keep it as a manual prompt. That’s not a failure, that’s the right call. This test is for the smallest loop: a repeatable instruction inside a chat or coding session.

1. THE AGENT LOOP

The prompt loop is useful, but it ends when the conversation ends. A production agent needs architecture around the model that can continue without routing every step through you.

A PUBLIC EXAMPLE

Codex and Claude Code are familiar examples of this pattern: you give them a messy goal, they investigate, plan, use tools, check their work, and continue until the work is done or they need you. Karpathy’s AutoResearch repo makes the same principle explicit: the agent can change the training program, but not the evaluator.

That separation is the idea at every scale: the thing doing the work should not be allowed to redefine the bar.

THE VERSION I ACTUALLY RUN

I built and run this pattern in Niki Studio, an AI design product. The harness asks an agent to generate a design image, make it editable, rebuild the canvas with the real text, repair anything missing, and check quality. Six steps.

Here’s what surprised me: the model could explain this sequence perfectly. Then it would generate the image and confidently call the wrong tool next.

My first fix was everyone’s first fix: a stronger prompt. More emphasis. It helped for a week, then broke under a different model. That’s when the lesson landed: an instruction is not a durable fact. Language is not state.

The fix was structure around the model. Niki Studio, not the model’s memory, tracks which step each item is on. If the agent tries to jump ahead, it gets a structured rejection: “not this tool, this one is expected next.” It gets told no, and why, and it course-corrects itself.

The hardest bug proved why the stop rule matters as much as the check. A user asked for a two-page design. Page one passed its quality check, and one branch of my code declared the whole job finished. Page two was still sitting in the repair queue. The user got one page. Every line of that code looked correct on its own. The system still broke its promise.

A loop isn’t just “try again until good.” It’s a set of promises: about what counts as done, what happens to unfinished work, and who’s allowed to declare victory.

THE SMALLEST LOOP YOU CAN RUN

You don’t need any of the infrastructure to feel how this works. Paste this into Claude or ChatGPT:

Work in a loop until this task is truly done.

TASK:

[describe exactly what you want produced]

SUCCESS CRITERIA (be strict, this is what "done" means):

- [criterion 1, e.g. "every claim has a source"]

- [criterion 2, e.g. "under 500 words"]

- [criterion 3, e.g. "a beginner could follow it"]

EVERY ROUND:

1. Do the work, or improve the last version.

2. Honestly score the result 1-10 against each criterion.

3. Name the single weakest part.

4. If every criterion scores 8+, say DONE and stop.

   Otherwise fix the weakest part first and go again.

RULES:

- Never say DONE until every criterion genuinely scores 8+.

- Don't ask me questions. Make a sensible assumption, note it, keep going.

Watch what happens. The model drafts, grades its own work, finds the weak spot, and rewrites until it clears your bar. You get a later attempt instead of the first thing that looked close.

Notice what’s still missing: you’re still the trigger. Close the tab and it’s gone. Moving from this prompt to a production loop means adding persistent progress, a schedule, and a check the model cannot soften. That’s where the engineering starts, and where the four-part test decides if it’s worth it.

3. THE GRAPH AROUND THE LOOPS

Once a product has more than one agent or tool, a single loop is not enough. The graph is the outer structure: nodes for agents, tools, validators, or human approvals; edges for handoffs, branches, retries, and exits.

One node may contain the Design agent loop described above. Another may be Research Loop. Another may be Review. The graph decides which one runs next. The loop decides whether its own work is good enough to hand off.

That is the useful relationship: prompt loops help individuals, agent loops run systems, and graphs coordinate systems of loops.

THE HONEST PART: HUMAN IN THE LOOP

Loops change the work. They don’t delete you from it.

The faster a loop ships work you didn’t do yourself, the bigger the gap between what exists and what you understand. That gap charges compound interest. And when the loop runs itself, it’s tempting to stop forming opinions and just accept what comes back.

Two people can run the same loop and get opposite outcomes. One moves faster on work they understand deeply. The other avoids understanding the work at all. The loop can’t tell the difference. You can.

Stop being the engine. Keep being the judge.

Back to writing