Spec · 2026-08-28 · for red-pen

Closing the Disk Loop

Four located seams, ~53GB reclaimable on your word today, and the wiring that means you never think about disk space again.

Draft for Robert’s annotation · three decisions inside · nothing armed or deleted yet

Jump by type: Your three calls · Reap table · Honest risks · The math · How every delete gets courted

Your three calls — everything else is already moving

1 · Reap-now signature. ~43GB of dead verification clones + 10GB of expired-lease worktrees. Every path: zero open files, owner provably finished, receipts on execution. (Also: a 1.3GB personal video in temp space — keep, or move to the archive drive?)
2 · Policy amendment v2. Five rows: a new janitor zone for root-level temp clones · leases stamped at birth (and the worktree zone gets real delete hands after its court) · scratch dirs register their owner at creation · caps and floors become loud instead of silent · the fleet boxes' historical run backlog joins your signed scope by its literal root (found mid-arc: the boxes are refusing dispatches at 2–4GB free, and the 08-26 signature names a folder that turns out to be empty while the debris sits one directory over — new box runs already clean up after themselves as of today).
3 · Name the quiet moment — ANSWERED: "you pick." Two switch-flips remain (the self-clean fix turned out to be already live): pointing new runs at the SSD, and swapping the janitor's config. I watch for the lull and flip them with receipts.

The journey — follow one run's exhaust through the machine

  1. The leak, located. Your janitor works — it reclaimed ~10GB today alone. But four seams sit outside its reach, and they explain the whole drift: workers mint 3GB verification clones one folder outside every zone the janitor may touch; worktrees are born without the ownership tag the rules require before anything can be reclaimed; the overflow SSD has signposts but no road (nothing ever reads them); and the only alarm that fired did so as a silent banner at 3:30am, four nights running.
    ↘ go deeper — the receipts
    Four independent read-only investigations, 2026-08-28. (a) The /private/tmp root clones (~43GB) are minted by sealed-run workers' own ad-hoc bash, violating their brief law ("never touch anything outside RUN_DIR"); one clone was forensically matched to a specific run's execution window via the ledger. (b) The 08-27 "bootstrap the leases" ruling ran as a one-time backfill (75 leases) and was never wired into creation — 50 of 71 worktree dirs today are lease-less; the worktree pool is mode: report in the janitor's runtime config ("no delete path exists for it in this build"). (c) NEURON_SPILL_ROOT/FLEET_RUN_ROOT: ~30 readers, zero writers — every reader falls back to a hardcoded internal path; Fleet-Scratch holds 804KB of filesystem bookkeeping. (d) The GC alarm's three trigger conditions are all "janitor broke" — a healthy janitor losing the race writes status=ok forever; it has never fired. Canonical detail: DISK-LOOP-CLOSURE-2026-08-28.md §Seams 1–4.
  2. Immediate relief — EXECUTED 08-28: 9GB free → 55GB free, no restart. You signed with one condition — "damned sure they're safe in the past" — and the condition earned its keep: an independent adversary broke my first dead-list, killing five rows (a hand-written "do not reap" note a blanket pass had steamrolled, the checkout your live Cloudflare Worker deployed from, three trees at open pull-request tips). The 37 survivors were archived to the SSD (48GB of fingerprinted tars, receipt per path, working undo) and then reaped, each re-probed for life at the instant of deletion. The five kills now carry protective markers and live leases on disk.
    ↘ go deeper — the evidence table
    Clones (each: 0 open file handles in a fresh lsof snapshot · untouched 8h–3d · owner run terminal in the bay ledger · no reference in any launchd/cron/hook registration surface): neuron-fast-verify 3.1G, b7c-fresh 3.1G, b7c-base 3.1G, nf106-repair 3.0G, nf-88Mahl 3.0G, neuron-fast-issue110 3.0G, court2136 3.0G, court-2136 3.0G, nf-pr106-repair 2.5G, m2-boundary-red 2.5G, governance-debt-fresh 2.5G, neuron-verify-64553 2.0G, court106r5 1.6G + smaller siblings. Expired-lease worktrees (19, all holder: lease-bootstrap-20260827, expired 2026-08-28T05:05): courtA-131, courtC-pageid ×2, forms-bay-courtC, p0-worker-deploy, plane-tools-anchor, w1-contracts-cure (2.5G), w2-p0-court, w3-canon, w3-fast, w5-guard-court + 8 small. Delete-time recheck per path: open-handles probe + registration grep + lease re-read. Receipts append to a reap log in the disk-plan folder.
  3. Runs clean up after themselves — CORRECTED: this was already live. Mid-arc the warp desk proved the "run reaps its own scratch when its evidence banks" fix shipped on 08-27 and fired on this very Mac today — my investigator had graded a stale local copy of the machinery while the truth lived on trunk. The branch I'd flagged as "built but never shipped" is a superseded duplicate marked do-not-adopt. What actually remains here, chartered with the warp desk: the guard that makes rogue clones outside a run's own folder structurally impossible (that's the ~43GB source), and the schedule backstop.
    ↘ go deeper — what ships
    Branch bay/reap-at-completion-r3 in ~/genome-fleet/warp, 2 commits, built 08-26 22:22, ledger outcome SUCCESS, pushed_to_origin: false (a branch-shape rule refused the direct push; it ships via PR). Wiring lands in the sealed-run harvest path so the run dir is swept when the terminal ledger row + banked report + harvested branch all exist — the same predicate the ratified sweeper computes today. The worker-side guard blocks git clone/git init targets outside the run dir at the sandbox level, so the brief law gets teeth instead of trust.
  4. Every byte born with an owner and an expiry. Three moves, all wiring of things that exist: the lease-stamper runs on the janitor's half-hour tick so no worktree is ever ownerless for more than 30 minutes; the scratch registry (built two weeks ago, zero callers) gets called wherever scratch is minted; and after a grace week, unowned scratch becomes reapable — the janitor's own log says that flip alone releases 5.1GB today. Cap breaches stop being silent: over-cap with nothing legally reapable becomes an alarm, not a diary entry.
    ↘ go deeper — the mechanism
    (1) Re-run the ratified lease bootstrap now (covers today's 50 lease-less dirs), then schedule it — new worktrees get LEASE.json with expires_at at birth or within one tick. The worktree pool flips report→reap ONLY after the lease layer passes its own planted-victim court — the 08-26 court explicitly warned this flip is one config word, which is why the court comes first. (2) Registry call sites land in the shared scratch-minting helpers, not in 2,816 call sites individually. (3) unregisteredPolicy: keep → tmpfiles after the grace period, per the janitor's own counterfactual (88,681 dirs / 5.1GB). (4) The new temp-root zone's done-proof: owner provably dead + ≥6h age floor + zero open handles + registration grep clean — ships report-only, arms after court. (5) The 5GB floor becomes a delete-time circuit breaker inside the reap loop (today it's only a config-shape check). (6) Box backlog by literal root: the signed 08-26 box row names box:/tmp/engine-bay-runs, which dry-runs EMPTY (warp desk receipts, 08-28) while the real debris lives at /home/exedev/warp-box-runs — the literal path binds, so the warp desk correctly reported instead of sweeping; the new row extends the same instrument + predicate + any-doubt=keep to that root. New box runs self-reap at completion since bay 0.6.178, so this clears history, not a growing pile.
  5. Heavy exhaust moves to the SSD — and the wall becomes physical. One environment variable points the bay's runs at the 140GB quota'd scratch volume (the per-run temp redirect follows it automatically); one more export carries ~2,800 scratch-minting call sites and the compiler cache with it. From then on, a runaway job hits the wall of its own 140GB box and fails alone — it cannot take your machine down. The internal disk keeps only what's durable.
    ↘ go deeper — the two traps we're building around
    Trap 1, the phantom mount: if the SSD is unplugged, every producer would silently write to a plain folder on the internal disk at the same path. No script today checks the volume is real. The mount guard ships in the same change as the repoint, never separately: no mount, no run, loudly. Trap 2, the instant-copy downgrade: sealed runs hand each worker a ~570-package dependency tree in ~5s via copy-on-write cloning, which doesn't cross volumes — naively moved, every run pays a full multi-hundred-MB copy. Cure: the shared dependency store co-locates on the SSD. Also decided here: worktrees stay internal for now (bounded by the lease layer); durable evidence (reports, ledgers) stays internal by design — only ephemeral exhaust spills.
  6. When the loop is under pressure, you hear about it — in daylight. LANDED 08-28. Two hours below your hard line now files a GitHub issue (the one channel with a proven round-trip to you), with an 18-hour cooldown so a bad week pages daily, not half-hourly — tested in both directions before arming. Every new session now opens with a one-line disk status (the line built on 08-26, finally plugged in). Every worktree now carries a declared owner-and-expiry tag, refreshed by a half-hourly backfill, so nothing is ever ownerless again. Still queued: retiring the 3:30am banner and consolidating the redundant watchers.
    ↘ go deeper — alarm mechanics
    Fourth trigger in the existing alarm leg: free below hard on 4 consecutive heartbeats (~2h), hard-line value read from the ratified policy file so alarm and janitor cannot drift; per-trigger cooldown 12–24h. Heartbeat lines gain pressure-tier and cap-overflow fields; cap breach sustained over N runs alarms too. The 52MB shared alarm log splits: the fleet-obligation logger (164k of its 187k lines) gets its own rotated file. Watcher consolidation: janitor + alarm + nightly mirror keeper stay; the hourly logger, the daily 80% banner, and the retired daily janitor fold in.
  7. Nothing gets delete hands without surviving its court. Every change that can destroy ships the same lawful lane the 08-26 work established: dry-run receipts first, planted victims it must refuse AND planted victims it must take (a check that cannot fail proves nothing), then a fresh non-author court tries to break it, then your signature, then it arms. The courts already caught a real remote-code hole in this janitor once — the lane works.
    ↘ go deeper — per-build court scope
    New temp-root zone: planted should-keep (live pid; registered path; never-touch overlap) + should-reap (dead owner, aged, clean) victims, both directions. Worktree report→reap flip: lease-layer court first (expiry honored, live lease refused, missing lease refused). Self-clean merge: re-run the subprocess/verdict-parsing census from the 08-26 recourt on any new predicate code, since shell-injection via crafted dir names is the exact class the first court caught. Config↔policy parity re-diff after amendment (runtime reads the config file; the policy is provenance). Dry-run is the default everywhere: only the armed daemon's own launch arguments carry live mode.
  8. What "done once and for all" means. Steady state: internal disk holds only durable things (~150–200GB free); all transient exhaust lives behind the SSD's 140GB physical wall; scratch vanishes at run completion with the janitor as backstop; a sustained breach pages you same-day; and a runaway fleet costs one failed job, not your machine. The restart you're deferring stays optional — it buys back ~30–60GB of system-side space (staged macOS update, system caches) whenever it's convenient, but the loop closes without it.
    ↘ go deeper — the before/after table
    Today: ~9–14GB free, below the 25GB hard line; reclaim latency 30 min where allowed, infinite where not; pressure discovered by feel. After relief (decision 1): ~60GB free same-day. After all workstreams: ~150–200GB free; reclaim at completion; ENOSPC contained to the job's own box; page on sustained breach + session-start line. The three laws this obeys: ownership is declared, never inferred · instruments get teeth only after a non-author red-proof · silence is a bug, sustained breach pages.

Honest risks

Mid-flight arming. Your time-sensitive work is running now. Mitigation: everything builds and courts in isolation; the three switch-flips batch to the quiet moment you name (decision 3). Nothing touches a live run.
The worktree delete-path flip. One config word separates "report-only" from "deletes worktrees." Mitigation: the flip is last, behind the lease-layer court, and the reap-at-completion scar (a pruner once ate a live launch pad) is exactly why ownership is declared, never inferred.
Unmounted-SSD phantom writes. Named, guarded, and shipped in the same change as the repoint — never separately.
Config/policy drift. The janitor runs from its machine config; your signature lives in the policy file. Every amendment ends with a parity re-diff so the two can never quietly diverge.
The one rule under all of it: no agent deletes outside a scope you signed — and every scope you sign gets machinery with real hands, real courts, and a pager. The 08-26 report said it best: sign nothing, and the next reclaimer also ships report-only, and the disk fills again on Thursday.