← All work

Case study · 2026

Overseer

An orchestration system for AI coding agents. It takes a feature request to reviewed, merged code, one small task at a time.

Role
Design and development
Stack
TypeScript, Node, React, SQLite
Agents
Claude, Codex, OpenCode
Agent sessions
3,322
Tasks completed
741
Time span
17 days

The problem

Coding agents do well on small, well-defined changes and less well on whole features. Running several at once in one repository causes conflicts, unreviewed code and costs that are hard to see.

I wanted to talk to one orchestrator and have it hand small tasks to agents, instead of juggling CLI sessions and checkouts myself. The constraints were that it runs locally for one user, drives the CLIs I already use instead of its own model loop, and gives every task its own git worktree. In the first version only I could merge, so nothing landed without my review.

How it works

Every task goes through the same six steps. Nothing reaches the feature branch without a review and a passing check.

  1. 01

    Plan

    Split the feature request into small tasks, each with its own definition of done.

  2. 02

    Run

    Give every task its own agent (Claude, Codex or OpenCode) in an isolated git worktree.

  3. 03

    Review

    A second model reviews every change before it moves on.

  4. 04

    Verify

    Check the change against the task’s definition of done.

  5. 05

    Merge

    Merge passing tasks into a feature branch for human review.

  6. 06

    Learn

    Write down what went wrong and feed the lessons back into the prompts.

The office

A live pixel-art office shows the agents at work, so you can see what is running at a glance.

It's the default view, so it's the first thing I see. Each running session is a character at a desk with its harness, model and task above it, and a session that goes quiet for too long dims and wears a clock. The counts live in the room, with a sticky note for open questions, folders on the meeting table for work waiting on my review and a whiteboard with the task board's columns. I mostly use it to spot a stalled agent or a waiting review, since clicking the character or the object takes me straight there.

Cost and learning

Overseer records the cost of every task. When a task fails review or verification, the lesson is written down and fed back into the prompts for later tasks.

One lesson came from agents that finished their work but never committed it. Verification runs inside the task's worktree, so the uncommitted change passed its checks and was then dropped at the merge. Because of this, the prompt now tells every agent to commit each finished piece as soon as it stands on its own and to report a clean git status before it stops.

Results

In 17 days Overseer ran 3,322 agent sessions, and 741 tasks landed on merged branches.

Most of that work built Overseer itself (560 tasks), with 175 tasks for a work project and 6 for this site. 47% of reviewed tasks passed their first review, meaning 299 of 641 landed tasks came back with no findings the first time, and I don't think that's a bad number since the review caught the rest before they merged. The agents and reviews cost about $5.60 per task on average, from the cost each CLI reports or, where it reports none, an estimate from token counts at API prices.

TypeScriptNodeReactSQLite
More work← Back to all projects