One task at a time, verified.
A free Claude Code plugin that builds an approved plan one task at a time, and checks every task with a reviewer that did not write it.
Free and MIT licensed. It runs inside your own Claude Code session as a set of skills: there is nothing to sign up for and nothing to run beyond the two lines below.
claude plugin marketplace add agilepeter/dispatch
claude plugin install dispatch@agilepeter
The loop, drawn.
One approval, up front. Then dispatch runs the tasks itself, one at a time, and checks each one before moving on.
Why a verifier.
An agent grading its own work is the failure mode this plugin exists to avoid. dispatch never asks the implementer whether a task is done: a fresh reviewer subagent, with none of the implementer's context, reads the actual code and the actual diff instead of the implementer's report of it. Then a gate runs the plan's own verification commands and checks their exit codes, never the implementer's claim of a green run. It is the same principle behind harness engineering, aimed at one narrow problem: getting a multi-step coding plan built without taking any agent's word for it.
The receipts.
Every dispatch run appends one row to a local ledger, win or lose. These are that ledger's own numbers, not a sample or a demo.
As of 2026-09-29 (first task 2026-05-16)
Of those 142 tasks, a spec reviewer sent one back to the implementer 36 times, and a quality reviewer rated something Important or Critical 41 times. Both counts are from tasks the implementer had already reported done.
The ledger itself is one row per task, 13 columns, in a local TSV file. dispatch-stats reads it back for a rolling summary; nothing on this page is a plan name, a client, or a task's text. The ledger does record each plan's path and a task's text, which is why it stays on your machine.
The evals.
The plugin ships its own eval suite, run with the plugin loaded and against a baseline without it, so a score says what the plugin changed. These results are the plugin README's own, word for word.
design-doc-before-implementer: a multi-file plan that nobody has approved yet gets a design doc and one approval question before any implementer runs.spec-before-quality: once a plan is approved with its design doc in place, the spec reviewer runs before the quality reviewer, and the task gets ticked.one-ledger-row-per-task: a single task run logs exactly one well-formed, 13-column ledger row.dirty-tree-blocks: an unrelated uncommitted change stops the run cold, with no edits and no git surgery to make it disappear.reviewer-reads-code: an implementer's commit and its "all tests pass" report do not survive spec review, which reads the code and finds what is missing.
| Case | With plugin | Without | Delta |
|---|---|---|---|
design-doc-before-implementer | 1.00 | 0.67 | +0.33 |
spec-before-quality | 1.00 | 0.56 | +0.44 |
one-ledger-row-per-task | 1.00 | 0.33 | +0.67 |
dirty-tree-blocks | 1.00 | 0.63 | +0.37 |
reviewer-reads-code | 1.00 | 0.56 | +0.44 |
Mean delta +0.45, from one complete run of 30 sessions: 17m53s wall clock at concurrency 4, $8.95 at list price. All five cases scored 1.00 with the plugin loaded in all three repeats. Per-grader pass counts for both arms are in evals/reports/0.1.0/result.json.
How much this proves: three repeats per arm is a small sample, and the plugin and the suite were both revised between runs until this one, so these are the scores of the final suite against the final plugin, not of a first attempt. The baseline moves from run to run. Four complete runs were made on 2026-09-27, each after the suite had been corrected, so they do not measure quite the same thing: their mean deltas were +0.42, +0.46, +0.41 and +0.45, and every case scored 1.00 with the plugin loaded in all four. The five cases were written by the plugin's author to show what the plugin is for. They say that it does those five things reliably, not how it will do on a plan of yours.
What you type.
Three entry points. A plugin's skills are namespaced under its own name, so the /dispatch: forms always work; the bare forms below work too, unless another installed plugin already claims that name.
/dispatch <plan-path>also/dispatch:dispatch- Start a plan, or continue one already in progress.
/dispatch-resumealso/dispatch:dispatch-resume- Continue an in-progress plan without retyping its path.
/dispatch-statsalso/dispatch:dispatch-stats- Show the ledger's one-line rolling summary.
A plan is a Markdown file with a Gates: block (the commands dispatch runs after every task, stopping at the first failure) and a list of checkbox tasks. A minimal one, from the plugin's own templates/plan.md:
# Plan: add a health endpoint
Approved: 2026-09-26: yes, ship it
Gates:
pytest -q
python3 -m py_compile app.py
## Tasks
- [ ] 1. Add a GET /health route that returns {"status": "ok"} with a 200 status code.
- [ ] 2. Add a test that requests /health and asserts the status code and body.
- [ ] 3. Wire /health into the existing router and document it in README.md.
What it never does.
The short list, checked against the shipped skill text.
- Never pushes on its own. The loop itself never runs
git push; only a plan's own task text can ask an implementer to, and the coordinator never does it for them. - Never marks a task done on the implementer's word. A fresh reviewer and an exit-code gate both have to agree first.
- Never skips the ledger row, even on escalation. Every task appends one row, success or not.
- Asks for approval once, never again. One design-pass approval per plan, then the run proceeds unattended.
- Keeps plans out of your git history. A plan saved during a session is excluded through git's own local info/exclude, not committed.
- One implementer at a time. Never two in parallel; they would conflict on the same files.
Questions, answered.
The ones worth asking before you point this at a real plan.
Is it free?
Yes. It is MIT licensed, and the whole plugin (the skills, the templates, and the two ledger scripts) is on GitHub. There is no account, no key of its own, and no paid tier.
Does it need an API key?
No. dispatch runs inside your own Claude Code session as a set of skills. Every implementer and reviewer it dispatches is a subagent of that same session, so it counts against your existing Claude Code plan the way any subagent does. The two ledger scripts are plain Python 3 and call no API at all.
Which models does it use?
Sonnet for the implementer and both reviewers by default. Opus only for a BLOCKED retry or an architecture judgment call. Haiku is there as an option for a single-file, mechanical task. All three are tier aliases, never a pinned model id, set in one editable Defaults block in the dispatch skill that you can change.
How is it different from asking Claude to "build the plan" in one go?
A single subagent that writes the code, tests it, and reports back is grading its own homework: the same context that wrote a bug is the context judging whether the bug is fixed. dispatch splits each task across subagents that never share context: one implements, one checks the code against the spec, one checks the diff for quality, and a gate runs your own verification commands instead of trusting any of their reports.
Where do plans and the ledger live?
A plan you point dispatch at by its path stays exactly where you put it. A plan that only exists in the conversation is saved to .dispatch/plans/ inside the repo, which dispatch excludes from git through the repository's own local info/exclude rather than editing your .gitignore. The ledger is a TSV file at $DISPATCH_LEDGER if you set it, otherwise ~/.claude/dispatch/runs.tsv, shared by every plan on the machine.
Does it work outside Claude Code?
The loop itself does not: it is a set of Claude Code skills that dispatch Claude Code subagents, so without Claude Code there is nothing to run it. The exception is the ledger: bin/dispatch-ledger and bin/dispatch-stats are plain, stdlib-only Python 3 files, so a shell script or a CI job can append to or read the same ledger with no Claude Code involved.