Case study · 2026
Overseer
An orchestration system for AI coding agents. It takes a feature request to reviewed, merged code, one small task at a time.

- Agent sessions
- 3,322
- Tasks completed
- 741
- Time span
- 17 days
The problem
Coding agents do well on small, well-defined changes and less well on whole features. Running several at once in one repository causes conflicts, unreviewed code and costs that are hard to see.
I wanted to talk to one orchestrator and have it hand small tasks to agents, instead of juggling CLI sessions and checkouts myself. The constraints were that it runs locally for one user, drives the CLIs I already use instead of its own model loop, and gives every task its own git worktree. In the first version only I could merge, so nothing landed without my review.
How it works
Every task goes through the same six steps. Nothing reaches the feature branch without a review and a passing check.
- 01
Plan
Split the feature request into small tasks, each with its own definition of done.
- 02
Run
Give every task its own agent (Claude, Codex or OpenCode) in an isolated git worktree.
- 03
Review
A second model reviews every change before it moves on.
- 04
Verify
Check the change against the task’s definition of done.
- 05
Merge
Merge passing tasks into a feature branch for human review.
- 06
Learn
Write down what went wrong and feed the lessons back into the prompts.
The office
A live pixel-art office shows the agents at work, so you can see what is running at a glance.
It's the default view, so it's the first thing I see. Each running session is a character at a desk with its harness, model and task above it, and a session that goes quiet for too long dims and wears a clock. The counts live in the room, with a sticky note for open questions, folders on the meeting table for work waiting on my review and a whiteboard with the task board's columns. I mostly use it to spot a stalled agent or a waiting review, since clicking the character or the object takes me straight there.

Cost and learning
Overseer records the cost of every task. When a task fails review or verification, the lesson is written down and fed back into the prompts for later tasks.
One lesson came from agents that finished their work but never committed it. Verification runs inside the task's worktree, so the uncommitted change passed its checks and was then dropped at the merge. Because of this, the prompt now tells every agent to commit each finished piece as soon as it stands on its own and to report a clean git status before it stops.
Results
In 17 days Overseer ran 3,322 agent sessions, and 741 tasks landed on merged branches.
Most of that work built Overseer itself (560 tasks), with 175 tasks for a work project and 6 for this site. 47% of reviewed tasks passed their first review, meaning 299 of 641 landed tasks came back with no findings the first time, and I don't think that's a bad number since the review caught the rest before they merged. The agents and reviews cost about $5.60 per task on average, from the cost each CLI reports or, where it reports none, an estimate from token counts at API prices.