Status
Open source · v0.11
Period
June 2026
Stack
Shell · Python · Claude Code · Beads
The agent that did the work should not be the one that grades it.
Every task closes through the same path: deterministic gates first, then a read-only evaluator that starts every criterion at FAIL and flips it only on evidence.
1Problem
Left to run on their own, coding agents fail in a small number of recognisable ways. They report a task as finished without running anything. They quietly cut a feature down to a stub and call it done. They disable the test that was failing. They ship code that is green and wrong. They close a task while the change is still uncommitted. And over a long run the planner, the loop and the task tracker drift apart until none of them describes what actually happened.
Most harnesses answer this by asking the same agent that did the work to check the work. That is the part this project refuses.
…they ask the same agent that did the work to grade the work.docs/DONE-IS-EARNED.md
2Question
Can a check the agent cannot skip make an unattended loop trustworthy?
Trustworthy enough to leave running overnight on a real codebase and read the result in the morning.
What does it cost?
In tokens, compared with a single session doing the same work. The method for measuring this is written; the measurement is not done yet.
3Method
The evaluator is the only part that may say done.
One source of truth
Plans go straight into the Beads tracker as an epic, tasks and dependency waves. There is no markdown plan to drift away from the queue.
A fresh worker per task
Each task runs in a new session with a clean context, so nothing from the previous task leaks into the next one.
Gates that block
Hooks exit with code 2 and stop the agent: lint and types after every edit, dangerous commands, commit-before-close, and a fence around the files the task may touch.
A default-FAIL evaluator
A read-only agent (Read, Grep, Glob, Bash) checks that the changed path was actually exercised with a negative control, that the user-visible behaviour changed, that the code was read, and that nothing was stubbed or scoped down.
Parallel waves
Tasks with disjoint write zones run at the same time; tasks touching the same file run one after another. An integration gate checks the batch works together.
Built for unattended runs
Stalls are detected by activity, orphaned tasks are reaped by heartbeat, and the run report refuses to say DONE with an orphan or a dirty tree.
The harness is checked with itself: completely.toml runs ruff and the contract suite on its own repository. The README reports 77 contract tests passing at the time of writing.
4Findings
The evidence is four incidents from real runs, written up in docs/DONE-IS-EARNED.md. They share one shape: the check passed, and the thing it checked was broken.
The value of a default-FAIL gate is not that it catches bad work — it’s that it catches plausible work.docs/DONE-IS-EARNED.md
The rule that came out of it: every layer that validates something must itself be validated against the production path, with a negative control. A test that only proves the check runs proves nothing about whether it works.
5Limitations
- Token cost is a multiple of a single session and has not been measured. The benchmark method is written, the benchmark has not been run.
- Unattended mode still needs an active session to keep the workers alive.
- The evaluator is only as good as the evidence it is given.
- Open issues: an error from the tracker is read as an empty queue — a silent, false “done”; heavy concurrency on the tracker is unsolved.