14 — Validation plan: riskiest assumptions → cheapest disproving tests
Brief §17. This is not the pre-mortem (
13-pre-mortem.md, sibling — 18-month failure modes with owners). This is the near-term test plan that decides whether to grow past the walking skeleton at all. It operationalizes12-prd.md§20 open questions (Q1–Q7) and §19 risks (R1–R5) into runnable experiments, sequences them into 2 weeks, and states exactly what result flips which decision — feeding this plan’s own00-SYNTHESIS.md(pending; not to be confused withplans/001-factory-apps-validation/00-SYNTHESIS.md, already decided).Method note (Torres-adjacent, per the advisory board’s own repeated demand): feasibility of the claim primitive is not re-tested here as a risky unknown — four research passes (competitive, oss, priorArt, temporalCase digests) plus
10-architecture.mdD1–D3 already over-proveSKIP LOCKEDat 10→64 concurrent. Budget goes to what’s actually uncertain: whether the problem is felt yet, whether agents can use the API unassisted, whether Tomas will read the output, and whether the thing is worth Robert’s maintenance hours. Agents are participants, not just infrastructure — an agent-based usability test is cheap, repeatable, and the fleet’s own primary-user population (04-user-stories.md: “agents are personas too”).Cost figures below (Robert-hours, tokens) are planning estimates, not measured data — marked as such throughout, per house rule against fabricating numbers. Every test re-reports actuals at its checkpoint.
1. Assumption register
12 assumptions across the 7 categories named in the brief. Scored on importance (Critical/High/Medium/Low), evidence today (house confidence tags: confirmed/supported/hypothesis/assumption/unknown), and uncertainty (High/Medium/Low). Sequencing rule: test first = high importance × high uncertainty × weak evidence tier (unknown/hypothesis); cheap-confirm-only = high importance but already supported/confirmed with low uncertainty; park = out of window entirely.
| ID | Category | Assumption | Importance | Evidence today | Uncertainty | Treatment |
|---|---|---|---|---|---|---|
| VA-01 | Desirability | The collision/rework problem is real at current fleet scale (~10 agents), not only at a modeled 64 | Critical | unknown — zero named Groupon incident; arXiv:2606.19616 is hypothesis-tier and about a different system [priorArt §2] | High | Test first |
| VA-02 | Desirability | work-ledger measurably reduces rework/duplicate work vs. worktree discipline alone (today’s status quo) | Critical | unknown — never compared | High | Test first |
| VA-03 | Usability (agent) | A fresh Claude session, given only the CLAUDE.md claim-protocol paragraph, correctly claims/transitions on first attempt, zero-shot | High | assumption — 01-initial-recommendation.md asserts “the error message is prompt engineering,” never tested (Ive advisory demand) |
Medium-High | Test first |
| VA-04 | Usability (agent) | Error strings alone (no docs) let an agent recover from an illegal-transition or lost-claim-race without escalating to a human | Medium-High | assumption — same family as VA-03 | Medium | Test first |
| VA-05 | Feasibility | Zero double-claims holds at simulated 64-concurrent | Critical (NF1, hard invariant) | supported — SKIP LOCKED battle-tested (oss/priorArt digests), T1 spec’d in 10-architecture.md §12 |
Low | Cheap confirm only |
| VA-06 | Feasibility | grpn_ external auth works end-to-end for a non-Encore Claude session |
Medium | confirmed — commit eada68c, live mechanism, repoRecon §6 |
Low | Cheap confirm only |
| VA-07 | Feasibility | Exactly-once projection from Temporal (T2) | Low for v1 | n/a — no workflow-modelled units exist (repoRecon §3) | n/a | Parked — not scheduled this window |
| VA-08 | Viability | Robert-hours to build + run < Robert-hours saved from avoided rework | High | unknown — never measured | High | Test first |
| VA-09 | Viability | The concierge alternative (hand-curated markdown/CSV) does not already beat the service at current scale/duration | High (drives Q1) | supported (analytical only) — 11-build-vs-buy.md scored concierge 3.19 vs. 4.53, flip trigger named, never run head-to-head |
Medium | Test first |
| VA-10 | Trust | Tomas will accept the disposition table (SQL/Grafana) as kill-gate evidence, not just the ADR log he already reads | Critical (blocks half the product’s reason to exist) | unknown — literally unasked, per every advisory-board member’s demanded question | High | Test first, day 0 |
| VA-11 | Politics | A1 (CODEOWNERS-partitioned lane) is granted, or A-lite is triggered cleanly | High (blocks default placement) | assumption — carried from plan 001’s A1, likelihood Medium, research/14 unsent |
Medium | Test first, day 0 |
| VA-12 | Data | Ledger’s summed cost_cents reconciles with actual Anthropic billing within tolerance |
Medium-High (North Star numerator integrity) | unknown — never checked | Medium | Test week 2 |
What this list deliberately excludes: feasibility past 64-concurrent (no published prior art exists to test against, 10-architecture.md NF5 — this is a hard ceiling by evidence, not a 2-week experiment) and VA-07 (no substrate to project from yet). Spending validation budget there would be re-deriving what four research passes already settled — exactly the over-investment Torres and Cagan’s advisory positions warn against.
2. Sequencing logic
Two things move independently of the priority score: cost and latency. VA-10 and VA-11 are near-zero-cost (one message each) but slow (waiting on a human), so they fire day 0 regardless of rank — the schedule’s critical path is Tomas’s reply time, not compute. Everything else sequences by priority: test-first assumptions run before cheap-confirms; cheap-confirms run inline with the walking-skeleton build since they’re nearly free once it exists; VA-07 never runs in this window.
flowchart LR D0["Day 0<br/>fire slow+cheap asks"] --> D1["Days 0-2<br/>walking skeleton + cheap-confirms"] D1 --> D2["Days 2-3<br/>zero-shot usability (VA-03/04)"] D2 --> D3["Days 3-7<br/>3-agent live trial vs concierge week (VA-01/02/08/09)"] D3 --> CP1["Checkpoint 1 — day 7"] CP1 -->|GO| D4["Days 8-11<br/>full v1 surface"] D4 --> D5["Days 11-12<br/>one real Orders unit to terminal disposition<br/>Tomas reads it live (VA-10 close)"] D5 --> D6["Days 12-14<br/>cost reconciliation (VA-12) + ramp micro-check"] D6 --> CP2["Checkpoint 2 — day 14 → feeds 00-SYNTHESIS.md"]
3. Test specs
Compressed view of every assumption’s experiment. Full design for the three headline experiments (walking skeleton + live trial, concierge week, fake-door) follows in §4–§6.
| ID | Method | Sample / participants | Success threshold | Failure threshold | Cost (est.) | Gates |
|---|---|---|---|---|---|---|
| VA-10 | One direct written message to Tomas: “will you read a disposition table at M4, or the ADR log you already read?” — not a doc | Tomas (n=1, the only relevant reader) | Explicit written yes, or he opens v_disposition_summary at the day-12 live read and engages with it |
Explicit no, or no reply by day 14 | ~5 min Robert time | PRD kill-trigger row 2 (fold disposition half into ADRs) |
| VA-11 | Send research/14’s A1 ask (already drafted elsewhere — not re-authored here) |
Tomas (n=1) | OWNERSHIP.md agreement, or explicit A-lite go-ahead, within 14 days | No response/no agreement | ~15 min Robert time | PRD R1 (trigger A-lite immediately — named fallback, not a program-ender) |
| VA-05 | Vitest race test against Docker PG: N claimers vs M<N ready rows, N up to 64 simulated | Synthetic (no fleet agents needed) | AC2.1/AC2.2 exact: M distinct units claimed, zero doubles; claim p99 < 200ms (12-prd.md NF2) recorded from the same race test — closes 04-user-stories.md US-01’s latency pointer |
Any double-claim observed | ~2 Robert-hrs (part of walking-skeleton build) | Stop-the-line bug if it fails — not a concept pivot, feasibility is already over-proven |
| VA-06 | One real grpn_ token, one real external Claude session, one real /units/claim call end-to-end |
1 live agent session | 200 with full unit state returned | Any auth failure | ~15 min Robert time | Same — stop-the-line bug, not a pivot |
| VA-03 | Zero-shot usability test: N=8 fresh Claude sessions, given only the CLAUDE.md claim-protocol paragraph (08-made-to-stick.md’s drafted artifact), pointed at the real walking-skeleton endpoint, asked to claim a seeded unit and do one transition |
8 fresh Claude sessions (small-N, directional — not a powered study; appropriate for a 2-human-reader internal tool) | ≥6/8 (75%) execute the correct claim call on first attempt | <4/8 (50%) | ~1 Robert-hr setup+review; ~8×15k tokens (~120k) | <50% → rewrite the CLAUDE.md paragraph / error strings before wider rollout (PRD Q5 spirit: fix the primitive, don’t add docs) |
| VA-04 | Same 8 sessions, 4 of them get a deliberately induced illegal-transition or lost-claim-race; observe recovery using only the returned error string | 4 of the 8 above | ≥3/4 recover correctly without escalating to a human | <2/4 | included in VA-03’s cost | Same as VA-03 |
| VA-01 / VA-02 | 3-agent live trial (§4) vs. concierge week (§5), same-domain comparable batches, run in parallel | 3 live fleet agents + Robert (concierge) | See §4/§5 combined thresholds | See §4/§5 | See §4/§5 | PRD Q1 (service vs. fold-into-monitor) and Q7 (has a collision actually happened) |
| VA-08 / VA-09 | Robert-hours ledger: tally build+supervision hours (walking skeleton + full v1) vs. concierge week’s hours, extrapolated to program duration | Robert’s own time log | Build+run hours < concierge’s extrapolated ongoing cost at current unit volume | Concierge stayed near-zero-cost and showed no degradation signal | ~30 min/day logging, no extra cost | PRD Q1 fold trigger; 11-build-vs-buy.md runner-up flip condition |
| VA-12 | Pull Anthropic console usage for the sessions behind one real terminal unit + the live-trial units; diff against ledger’s summed cost_cents |
1 real Orders unit + 3 live-trial units | Variance ≤ tolerance (initial target ±10% — a calibration knob, not a fixed constant, same treatment as claim TTL in 10-architecture.md §4) |
Variance > tolerance | ~1 Robert-hr | Blocks calling v1 “done” (cost integrity feeds the North Star numerator, 09-north-star.md §7); not a kill trigger for the concept |
4. Headline experiment 1 — walking skeleton + 3-agent live trial on real Orders prep units
This is the cheapest test that could disprove the concept, because it is also literally the first production increment (10-architecture.md §12 T1; PRD Phase 1 — same artifact, not a throwaway). Reusing PRD’s own release-strategy build here; what this document adds is the falsification framing and the concierge comparison arm, which the PRD’s phase gates don’t carry.
Build (days 0-2, ~1 Robert-day): work_units migration + claim() atomic guard (pull-next + named) + v_queue_depth + one Grafana panel + the T1 concurrency test. Nothing else — no transition/block/dispose yet.
Live trial (days 3-7): 3 real fleet agents claim, work, and (informally, via the walking skeleton’s minimal surface) report on 3 real Orders characterization/audit prep units — not synthetic fixtures. Comparable in size/type to the concierge arm’s batch (§5), so the two arms are apples-to-apples.
| Signal tracked | Success (supports building) | Failure (undermines urgency) |
|---|---|---|
| Double-claims | Zero, across the whole trial | Any occurrence — genuine surprise, stop-the-line (VA-05 already over-proven) |
Rework (attempt_count > 1, stale_reclaimed events) |
≥1 real instance observed, or agents self-report a near-miss | Zero instances and zero near-misses across 3 real units and 5 days |
| Agent friction (unprompted deviation from claim protocol, confusion in transcripts) | Low, consistent with VA-03 usability result | High — compounds a VA-03 failure, same fix |
| Robert supervision hours | Lower than concierge arm’s, or comparable | Higher than concierge arm’s with no offsetting collision-prevention benefit |
Why 3 agents and 3 units, not more: cheap enough to run this week without disrupting real Orders prep throughput; small enough that a null result (zero collisions, zero rework) is itself informative — it’s the number a skeptic would ask for before committing more (Cagan/Christensen advisory demand: “has an actual collision happened yet, at real scale, not the projected ceiling”).
5. Headline experiment 2 — the concierge alternative (one week, hand-curated)
Runs in parallel with §4 (days 0-7), on a comparable batch of 3-8 real Orders prep units, tracked by Robert in a flat markdown or CSV table — unit id, claimant, state, disposition, cost, evidence — updated by hand or by agents appending a line, with no atomicity (first-come-first-served by convention, exactly the “dumb” baseline 11-build-vs-buy.md scored analytically but never ran).
Purpose: calibrate whether the collision problem is felt yet at all (VA-01), and what the honest ongoing Robert-hour cost of the dumbest possible alternative is (VA-08/09) — the number the fold-into-monitor kill trigger (PRD Q1) needs to be decided on evidence, not the analytical score alone.
| Signal tracked | Threshold that argues FOR building (service beats concierge) | Threshold that argues AGAINST (concierge already sufficient) |
|---|---|---|
| Observed collision or lost-update in the markdown table | ≥1 instance (two agents claim/edit the same row inconsistently) | Zero, all 7 days |
| Robert curation time | >30 min/day sustained, or visibly rising with unit count | <10 min/day, flat |
| Data integrity at week end | Any silently-dropped or contradictory field (matches the arXiv:2606.19616 file-tracker failure mode — priorArt §2) | Table internally consistent throughout |
Decision rule combining §4 and §5: if the concierge week runs clean (all three “against” thresholds hold) and the live trial also shows zero collisions/rework, that is the single strongest evidence available that VA-01 is not yet true at current scale — it does not kill the schema (the durable table is needed regardless per 11-build-vs-buy.md and the temporalCase digest, since the fleet monitor forces it day one) but it does argue against urgency on the standalone-service framing, feeding PRD Q1 toward “fold” rather than “keep separate,” on schedule rather than waiting for the PRD’s own 4-week gate.
6. Fake-door equivalent (internal tools have no landing page — this is the adapted form)
Commercial fake-door = a button that doesn’t work yet, measuring click-through before building. There is no purchase funnel here (N/A — internal, 2 human readers), so the internal analog is: ship the instruction before the mechanism, and watch whether agents even attempt to use something that isn’t fully real yet — organic compliance, not a staged demo.
Design: Day 0, before the walking skeleton exists, add the CLAUDE.md claim-protocol paragraph (the same artifact 08-made-to-stick.md names as “the highest-leverage, highest-risk artifact in the plan”) to the fleet’s system prompt. Point it at a stub POST /units/claim that logs the call shape and returns 501 not_implemented with a body asking the agent to describe what it attempted. Runs passively for 48h of normal fleet operation — zero dedicated agent spin-up, zero incremental token cost beyond what the fleet already spends.
| Signal | Success (protocol is legible, worth building the real thing) | Failure (protocol itself needs rework first) |
|---|---|---|
| Attempt rate | Agents that hit a claimable-work decision point call the stub at a rate consistent with “they read and tried to follow the instruction” | Near-zero attempts — agents ignore or never reach the instruction; a prompt-placement problem, not a schema problem |
| Call shape correctness | Logged calls are well-formed (right method, right field names) even though they 501 | Malformed/garbled calls — same fix path as a VA-03 failure, but caught before any backend exists |
This runs first, is nearly free, and can independently justify pausing before the walking-skeleton build if the attempt rate is near-zero — cheaper than discovering the same thing after 2 Robert-days of build.
7. Two-week timeline
gantt dateFormat YYYY-MM-DD axisFormat Day %d section Day 0 (parallel, cheap+slow) Fake-door prompt injection :d0a, 2026-07-19, 2d Tomas trust question (VA-10) :d0b, 2026-07-19, 1d A1 lane ask sent (VA-11) :d0c, 2026-07-19, 1d Concierge week begins (VA-09) :d0d, 2026-07-19, 7d section Build + confirm Walking skeleton build :d1, 2026-07-19, 3d Cheap-confirms VA-05/VA-06 :d2, 2026-07-21, 1d section Usability + live trial Zero-shot usability (VA-03/04) :d3, 2026-07-21, 2d 3-agent live trial (VA-01/02) :d4, 2026-07-22, 5d Checkpoint 1 :milestone, cp1, 2026-07-26, 0d section Week 2 (if GO) Full v1 surface build :d5, 2026-07-27, 4d One real unit to terminal + Tomas live read (VA-10 close) :d6, 2026-07-30, 2d Cost reconciliation (VA-12) :d7, 2026-07-31, 1d Ramp micro-check :d8, 2026-08-01, 1d Checkpoint 2 :milestone, cp2, 2026-08-02, 0d
8. Checkpoint decisions
Checkpoint 1 — day 7
| Condition | Decision |
|---|---|
| Fake-door attempt rate low, or usability <50% correct | PIVOT — fix the CLAUDE.md paragraph / error strings before building further; hold walking-skeleton scope where it is |
| Concierge clean AND live trial shows zero collisions/rework | PIVOT (scope, not concept) — proceed, but push toward Q1’s “fold into monitor” framing rather than a standalone-service timeline; do not treat this as reason to stop the schema, which the monitor needs regardless |
| ≥1 real collision/rework/near-miss observed in either arm, usability ≥75% | GO — proceed to full v1 surface (week 2) as scoped in 12-prd.md §21 |
| A1 rejected and A-lite not confirmed | PIVOT (placement) — do not build homeless; escalate per PRD R1 before continuing week 2 build, even if the technical signal says GO |
| Tomas answers “ADR log, not a table” | PIVOT (scope) — continue building claim/transition/block; do not build out the disposition/cost audit half until this is revisited (PRD kill-trigger row 2) |
Checkpoint 2 — day 14, feeds 00-SYNTHESIS.md
| Condition | Decision |
|---|---|
| VA-10 closed positive (written yes or engaged live read), VA-12 within tolerance, no double-claims across the whole 2 weeks, A1 resolved (either path) | GO — full build as scoped in 12-prd.md, ramp gating per PRD §24 |
| VA-10 closed negative | FOLD the disposition/cost audit half into ADRs; keep claim+transition+block as a thin primitive, possibly inside the monitor’s schema (Q1) |
| VA-12 outside tolerance | Ship v1 but flag North Star “verified throughput” as provisional until cost instrumentation is fixed — not a kill trigger, a data-quality gate |
| Live-trial-vs-concierge gap never materialized across the full 2 weeks | FOLD — schema ships regardless (monitor needs it), standalone-service framing does not survive; revisit PRD Q1 as closed toward “fold” |
| A1 rejected, A-lite also blocked | ESCALATE, do not build homeless — same as Checkpoint 1’s placement branch, now with a hard deadline reached |
This checkpoint’s output — GO / FOLD / ESCALATE, with the specific evidence that produced it — is the input this validation plan hands to 00-SYNTHESIS.md. It does not replace 12-prd.md §25’s own day-28 kill gate (ledger coverage <50% of fleet work); it is the earlier, cheaper leading indicator that should make that gate’s outcome unsurprising either way.
9. Total cost roll-up (estimate, not measured)
| Item | Robert-hours | Tokens |
|---|---|---|
| Fake-door + prompt edit | 0.5 | ~0 (marginal to existing fleet ops) |
| Trust + politics messages (VA-10, VA-11) | 0.3 | 0 |
| Walking skeleton build (days 0-2) | ~6-8 | ~200k-500k (one build session) |
| Cheap confirms (VA-05, VA-06) | ~1 | ~10k |
| Zero-shot usability test (VA-03, VA-04, N=8) | ~1 | ~120k |
| Concierge week (VA-09) | ~3-5 (across 7 days) | 0 |
| 3-agent live trial supervision | ~2-3 | (real fleet work, not incremental) |
| Full v1 surface build (week 2) | ~8-12 | ~500k-1M |
| One real unit + Tomas live read | ~1 (+Tomas’s ~15 min) | (real fleet work) |
| Cost reconciliation (VA-12) | ~1 | 0 |
| Total, 2 weeks | ~24-33 Robert-hours | ~1-2M tokens, mostly real productive build work, not validation overhead |
Marked as a planning estimate throughout (per house rule against fabricating numbers) — re-report actuals at each checkpoint; if actual build hours run materially above this, that is itself VA-08 signal.
Cross-references
| Doc | What it owns that this plan doesn’t |
|---|---|
12-prd.md |
Canonical open questions (Q1-Q7), risks (R1-R5), MVP scope, ramp gating (§24), kill criteria (§25) this plan operationalizes |
10-architecture.md |
T1/T2 test definitions, claim algorithm, the exact schema the walking skeleton implements |
07-ux-ui.md |
The 5 verbatim error strings used in VA-04’s induced-failure test |
08-made-to-stick.md |
The CLAUDE.md claim-protocol paragraph used verbatim in the fake-door test and VA-03 |
09-north-star.md |
Verified-throughput metric definition VA-12 protects the integrity of |
11-build-vs-buy.md |
The concierge alternative’s analytical score and flip trigger this plan tests empirically |
13-pre-mortem.md |
18-month failure modes (sibling; distinct time horizon from this 2-week plan) |
00-SYNTHESIS.md (pending) |
Consumes this plan’s Checkpoint 2 output |