Greg Ross

The agent factory: crews of Claude and Codex

Created

This is how I plan and build with AI agents. Claude Code and OpenAI Codex sessions run side by side in herdr, a terminal multiplexer for agents, organized as a small factory. A planner agent agrees each piece of work with me and writes a one-page plan, a crew builds it, and every piece of work is a tracked item on a Beads board.

Why

One long agent session doing everything didn’t work. It stalled, it hid what it was doing, and it lost context. I wanted four things: work I can watch, work done in parallel where it is independent, a second model’s perspective, and as little ceremony as possible. That last one came from experience: an earlier version had review points everywhere, and the ceremony stalled more work than it protected.

How a piece of work flows

1. Plan. Every plan is an epic. It has a one-page PRD (the idea, the goals, numbered stories with acceptance criteria, and a “Later” list) and one tracked item per story. Before any crew starts, a second model (Codex) reviews the plan once. After that it is consulted only now and then: for non-trivial planning, or when something is blocking.

2. Crew. Each epic gets its own crew: a lead (Claude Opus) plus one to three workers chosen for the stories.

Seat Model Gets
Lead Claude Opus Routes stories, verifies each against its criteria, closes them
Worker Claude Opus, low effort Judgment work: research, docs, design notes
Worker Claude Sonnet Hands-on operations: hosts, configuration, scripts
Worker Codex Code: features, refactors, tests

Stories are routed by strength, and independent stories run in parallel. At most three crews run at once. A live system with a single change owner, like my assistant, gets one crew at a time, so two crews never deploy over each other.

3. Review and evidence. Every story is reviewed once by a different agent (the lead reviews when the crew has only one worker), mixing model families where possible: Codex checks Claude’s work and Claude checks Codex’s. A review is “ok” or up to three findings, and there are no reviews of reviews. Each finished story carries an evidence line saying who built it with which model and effort, who reviewed it and the outcome, and how many context resets it took. An occasional audit reads those lines and adjusts the routing.

4. Visible and safe. There are no hidden agents: every agent runs in a pane I can watch. I approve each piece of work once, as a plan, and the lead makes every change undoable before making it. A refused command is escalated, never worked around. When an agent’s context fills up, it writes a handoff note and starts fresh.

5. Me in the loop. Only my decisions and steps that need me in person come to me; everything else is proven from logs. Work is grouped under a few outcomes I own, ranked by me, so a new epic starts from an outcome rather than from whatever came up.

The shape of it

In words, the diagram is:

me ⇄ planner → plan and stories → second-model review → crew (lead and workers) → build, review, evidence → board → lead verifies → wiki updated → done

What I learned

  1. Ceremony is a cost. Each check has to earn its place; the version with checks everywhere stalled.
  2. A second model family catches what the same model misses. Claude and Codex disagree in useful ways.
  3. Evidence lines make routing auditable. Which model is good at what becomes something I can read, not guess.
  4. Visibility beats autonomy theatre. Agents in panes I can watch are easier to trust and to correct.
  5. Approve the plan, not every command. I agree what we’re building; the crew does the work.
  6. The wiki is part of “done”. Agents file what changed, so the next agent starts informed.
  7. This page goes through the same publication checks as every other page. It was written by a crew and independently reviewed by a second model.