Build a multi-model relay: Claude architects, Codex verifies, Antigravity implements
Originally published on Medium
In this article

My $20 Claude plan died in 3 tasks. The subscription that looked generous on the first of the month is gone by the second week. Codex has its own caps, unless you get $100 plan (or get resets). And the models keep getting better in a way that makes you want to use them more, not less.
At the same time, something else kept nagging at me. The cheap and open sourced models, OpenCode running MiMo, Antigravity running Gemini 3.1, were getting good at a specific kind of work: well-defined, bounded tasks with clear success criteria. Not architecture. Not debugging a mystery. But “build exactly this, test it exactly this way”? Absolutely.
Those two facts pointed at one setup: organize them like a team where each member’s cost matches their job. Here’s the workflow I now use to build Niki Studio every day, and everything I’ve learned running it.
The team
Claude is the architect. When a feature needs designing or a bug crosses multiple layers, Claude investigates. It reads the code and the logs, separates what it verified from what it suspects, and writes a brief: the root cause, the exact change, the files that must not be touched, and how we’ll know it worked.
I give this role to Claude because Opus, and now Fable, are the strongest models I’ve used at architecture and system design: they hold a whole system in their head and reason about where a change belongs, not just how to write it. It’s the most expensive intelligence I use, so I point it at the work where a wrong decision costs the most.
Codex is the senior engineer. Codex’s current models are at that same frontier standard, and what makes them right for this seat is how well they hold long context and stick to a plan across a long session without drifting. Before anyone implements the brief, Codex checks its claims against the actual code, because a brilliant plan written against stale assumptions produces confident nonsense. Codex handles the delicate integration work itself and reviews every diff the juniors produce.
Its most valuable habit is skepticism: it reads the code, not the summary.
OpenCode and Antigravity are the junior engineers. They run low-cost models and take the contained work: test scaffolding, mechanical implementation from a precise spec, running regression suites and reporting results. The volume of this work is huge, which is exactly why it shouldn’t run on frontier pricing.
I’m above the loop. I set the goals, arbitrate trade-offs, test the result in the real product, and approve every commit. Not because I don’t trust the setup, but because someone has to own what ships, and that someone can’t be one of the models.
The two rules that hold it together
The first rule: expensive models think, cheap models type. Early on I’d hand a junior model a vague task and watch it fail. The problem was never the model’s capability. It was brief. When Claude writes a precise handoff, the junior’s output quality jumps dramatically. Most of the value in this workflow lives in the quality of the handoff, which is why the architect role earns its cost.
The second rule: no model grades its own homework. I learned this from a string of incidents. A junior once reported a task complete while it had quietly added a mock instead of the real integration. Another time a report claimed both pages of a design worked, and my product showed one. The model that made something will always be a generous grader of it. So the diff gets reviewed by a different model, the tests get run by someone who didn’t write them, and the browser gets checked by me.
The report tells me where to look. It never tells me what’s true.
The handoff, done cheaply
Here’s a practical detail: Claude Code and Codex can call another CLI as a tool: the senior model literally invokes the junior and waits for it to finish. It feels elegant. It’s also expensive, because the frontier model sits in an active session, burning context, while the cheap model works. Codex can call OpenCode directly, inside its own session.
My cheaper version: the handoff is a piece of text. Claude writes the brief, I paste it into OpenCode or Antigravity in a separate window, and when it’s done I bring the report back for review. One minute of copy-paste saves a meaningful amount of idle burn.
Here’s the thing: Models write the prompt for the junior model themselves. The whole point of having an architect is that Claude writes those briefs for you. Your job is the prompt at the top of the chain. This is the one I give Claude, and you’re welcome to steal it:
CONTEXT: You are the architect in a multi-model engineering pipeline.
- You (Claude): investigate the problem, diagnose root cause, and produce
an Architecture Brief with a phased implementation plan.
- Codex: senior engineer. Owns the implementation. Verifies your brief
against the actual code before acting on it. Implements the delicate
phases directly, and delegates contained, well-defined phases to
Antigravity/OpenCode. Gets reports back from them but does NOT trust
those reports at face value; runs the tests itself. If a report
reveals a small gap, Codex patches it directly when that's cheaper
than a round trip; only kicks it back to Antigravity/OpenCode as a
new contained task when the fix is non-trivial.
- Antigravity / OpenCode: junior engineers. Take one contained task at a
time: no ambiguity, no judgment calls, just "build exactly this,
verify with this exact check."
- A human reviews and approves before anything is committed.
Your job right now is the first step only: investigate and produce the
Architecture Brief. Do not implement anything yourself.
Act as the architect for this task. Investigate before you design: read
the relevant code and logs, and clearly separate what you verified from
what you still suspect - cite the file, line, or log line behind each
verified claim. A brief built on an unverified guess sends the wrong path
down the whole chain.
If you cannot establish the root cause from the code, say so and state
exactly what you'd need to confirm it - do not paper over it with a
confident brief.
Produce:
1. DIAGNOSIS: root cause (with evidence), the seam where the fix belongs,
and anything about the current code that surprised you.
2. ARCHITECTURE BRIEF - a phased plan, one phase at a time, written so
Codex can hand pieces of it downstream without re-deriving your
thinking. For each phase, state:
- what's true now, and what should be true after
- files it may touch, and files it must not touch
- how to verify it worked - the exact test, command + expected output,
or visual check if there's no test
- whether this phase is SENIOR (Codex should own it directly) or
JUNIOR-safe (Codex can hand it to Antigravity/OpenCode as a
contained task, if written specifically enough that it can't guess)
- "report back, don't commit - the model that wrote it doesn't grade it"
Order the phases so dependencies are respected - nothing downstream should
need a phase that hasn't landed yet.
Claude returns the diagnosis plus ready-to-paste briefs, sorted by who should get them. I paste the junior ones into OpenCode or Antigravity, and the senior ones into Codex with one standing instruction added: verify the brief’s claims against the actual code first, review the juniors’ diffs rather than believing their reports, and when a junior’s work has small gaps, patch them yourself instead of paying for another prompt round-trip.
A follow-up brief costs more than a two-line fix by the model that’s already in the code.
The “report back, don’t commit” line matters more than it looks. Reporting mismatches catches stale assumptions. Not committing keeps every experiment reversible until a human looks.
What I’d tell you to try first
You don’t need this exact stack, and you don’t need all four roles on day one. The pattern is portable to any pair of models: one strong enough to plan and review, one cheap enough to run freely, and you making the final call.
Start with a single handoff this week. Next time your strong model diagnoses a problem, don’t let it write all the code too. Ask it for the brief instead, hand that to a cheaper model, and review what comes back. You’ll learn two things immediately: how good your cheap model really is when the task is precise, and how much of your token spend was going to work that never needed frontier intelligence.
The models will keep changing. The pattern of matching cost to judgment, and never letting the maker be the only witness, will outlast all of them.