Skip to content
GRPNR.

Validation plan — dual-run parity harness

2026-07-18 · Author: Fable · Template §17. Binding inputs read in full: prd.md, research.md, build-vs-buy.md, plus 00-context-brief.md / initial-recommendation.md. Role of this document: not a delivery plan — the fastest, cheapest tests that could disprove the concept before the Orders week-6 clock makes disproof expensive. Where a test succeeds it also, incidentally, retires risk for delivery; that is a side effect, not the goal.

Internal-product reinterpretation (binding, context brief). No A/B test, no landing page, no signup funnel, no pricing test — none apply here, said once and dropped. “Desirability” = does the program (Tomas’s gates + the fleet) actually change behavior on our verdict, not whether anyone “likes” the product. “Trust” = would an auditor and a bypass-prone agent fleet both keep using it under pressure. Evidence tags: MEASURED / MODELED / DESIGN; epistemic split confirmed fact / supported interpretation / hypothesis / assumption / unknown, per Zaruba’s audit law.


1. The nine highest-risk assumptions

Everything else in the PRD is a design decision executed after these hold. If any FAILs at its threshold, the PRD section named in “if false” must be rewritten before further build spend — that is the whole point of testing cheaply first.

# Assumption Category Importance Current evidence + confidence Retired by
A1 Tomas will bind PR-lane promotion to a harness verdict row — not just build it, actually enforce it under schedule pressure Desirability CRITICAL — without this the harness is optional theater, exactly the TSB failure mode it exists to prevent (research.md D2 §9) Assumption. FR-C2 (PRD §9.6) designs the structural gate but no one outside Robert’s lane has agreed to honor it; Zaruba’s decision log has zero dual-run entries (build-vs-buy.md §5.3) E1
A2 A pinned-EOL Docker stack boots deterministically and records real traffic at acceptable cost, on at least the cheapest pilot Feasibility CRITICAL — gates every FR-L requirement (PRD §9.2) Unknown/hypothesis. Manifest-v1 pull rot is live and unresolved (docker-library/ruby#452); MODELED 2-5 eng-days first-of-stack (D3 §8, inference by analogy, not sourced) E2
A3 Recorded-fixture replay reproduces legacy behavior with fidelity high enough that a real divergence shows up as DIVERGED, not as noise or a false PARITY Feasibility CRITICAL — this is the kernel’s entire claim to correctness (PRD principle 7) Unsupported until tested — no existing self-replay precision number exists anywhere to borrow (D6 Q2) E3
A4 MBUS shadow-taps can diff JMS/STOMP/Kafka events between two systems at all, with no OSS prior art to lean on Feasibility HIGH — 16/33 legacy repos are MBUS-coupled; Orders alone has 25+ topics + a Kafka signal + an outbox durability contract (PRD §9.3 FR-A3) Confirmed market gap, zero prior art (D1 §11) — genuinely unknown whether it’s buildable at this cost at all E7
A5 Per-domain adapter cost stays roughly flat (the 2nd adapter does not cost >2× the 1st) — the generalization the harness is betting its post-MVP roadmap on Viability HIGH — 001’s own risk #1; detection threshold already named (PRD §20 Stage 2) Assumption, untested — only 1 of 3 Orders oracles is even built today (D4 §b) Baselined, not retired, by E6 — the 2×-detection ratio needs an adapter #2, which is post-MVP (PRD §20 Stage 2). Retired only when Stage 2 runs.
A6 Verify cost/unit (tokens + lab minutes) stays cheap enough that the fleet’s cost of running the harness stays below the cost of the regressions it prevents Viability MEDIUM — determines whether JTBD-1’s “cheaper than re-reading the diff” bar is met (PRD §6) No data yet; MODELED lab-setup cost exists (D3 §8) but no per-verify runtime cost has been measured E2/E4 (partial), full measurement is post-MVP
A7 Self-replay false-DIVERGED rate is at or near zero on a real recorded corpus, not just on toy fixtures Trust CRITICAL — nonzero and unbounded false-DIVERGED is the harness’s own stated top usability risk (initial-rec §10, PRD principle 3) Unknown — Diffy’s real-world false-positive rate was never published by Twitter or anyone (D6 Q2, D1 §13); no external number to borrow E5
A8 Whoever owns QSA/SOX evidence format at Groupon will accept the evidence-bundle shape (verdict → smallest counterexample → normalization ledger → replay recipe) as sufficient for DONE row 9 Trust HIGH — the persona itself is marked assumption, not measured, in the PRD (§5 note) Assumption — zero auditor engagement to date (journey.md §1.5) E4
A9 Zaruba’s plane will actually grant the live ≥99.999% money-gate lane per Option B, and accept the harness’s oracle-generalization pattern as the shared verdict contract Boundary CRITICAL — PRD §13 states this is “proposed, not ratified”; it is the single biggest scope fork in the whole document Unknown — zero decision-log entries either way (D4 §c/d) E1

2. The seven experiments, ordered by information-per-day

Cheapest and most scope-determining first. A conversation that forecloses half the design space in an afternoon outranks a five-day spike that only confirms one leaf assumption — do the conversation first, spend the days only on what survives it.

E1 — The Zaruba boundary conversation (retires A1, A9)

  • Method. Not a build. A structured, documented decision conversation — Robert presents PRD §13 (Option B) and FR-C2 (the structural promotion gate) to Tomas as two forced binary decisions, each with a proposed default and a decision date.
  • Design. Two questions, asked together, answered in writing:
    1. “Does the live ≥99.999% ledger-reconciliation gate live in your plane (Option B), or does the harness own it?” Default proposed: Option B.
    2. “Is ‘no verdict row ⇒ no PR-lane promotion’ (FR-C2) a binding structural rule you will hold under schedule pressure, or an advisory signal builders can override?” Default proposed: binding.
  • Evidence required. A dated written answer (doc, ticket, or email) to both questions, not a verbal impression.
  • Success threshold. Both questions answered within 5 business days; Option B ratified or a named alternative owner assigned; the gate confirmed binding.
  • Failure threshold. No written answer within 10 business days, or the gate is confirmed advisory-only.
  • Cost. ~0.5 agent/human-day (prep note + meeting + writeup). MODELED.
  • Sequence. Day 0-1. Before any other experiment spends a token — every other experiment’s design assumes an answer here.
  • If false / decision triggered. If the gate is advisory: PRD principle 2 (“structural gates beat human judgment”) is unfalsifiable in practice and the whole anti-TSB argument for building collapses to “nice-to-have tooling” — rewrite §3 Goals and re-score against build-vs-buy.md Option 8 (no new software) before spending further. If Option B is rejected (harness must own the live gate): PRD §13, §19, §20 all need to be rewritten around S1 (live shadow), which is a materially larger, differently-shaped build — do not proceed to E6/E7 until this is resolved, since their scope depends on the answer.

E2 — EOL boot PoC (retires A2)

  • Method. Throwaway spike, not production code. Pin forex-ng (2 endpoints, no DB, no MBUS — the ranked-cheapest pilot, D5 §e / architecture.md L261).
  • Design. (1) Pin the image by exact digest, no :latest. (2) Boot via docker-compose, health-check pass. (3) Record ~100 real requests through mitmproxy as an external proxy, no legacy code changes. (4) Replay the recorded 100 requests against the same still-running legacy instance (self-replay). (5) Diff must yield 100% PARITY — this is the falsification recipe named in the initial recommendation (§11) and PRD FR-K8.
  • Evidence required. Boot log with image digest; recorded fixture set (masked); self-replay verdict output, byte-identical across ≥3 runs.
  • Success threshold. Image pulls and boots deterministically; ≥95 of 100 recorded requests replay to PARITY on first attempt; any failures are diagnosable (not silent).
  • Failure threshold. Image will not pull at all (manifest-v1 dead end with no rescue-pull workaround), or self-replay produces >5% spurious DIVERGED with no traceable cause.
  • Cost. 1-2 agent-days. MODELED, consistent with D3 §8’s low end (forex-ng is a modern, non-EOL-tier stack, so this should land near the cheap end, not the JRuby/Java-7 tier).
  • Sequence. Day 1-3, immediately after E1 (does not depend on E1’s answer, can run in parallel with it).
  • If false / decision triggered. If forex-ng — the easiest stack in the estate — cannot boot deterministically or self-replay cleanly, the harder EOL tiers (Ruby 1.9.3/JRuby 1.7, PRD FR-L1) are not worth attempting yet. Do not proceed to giftcard_service (Rails 3.2) or any Orders-chain lab work; escalate straight to the kill-criterion in PRD §24 (“self-replay cannot reach zero false DIVERGED”) and re-scope toward S3 (impact-selected tests) before committing further lab budget.

E3 — Deliberately-wrong-stub test (retires A3)

  • Method. Piggybacks on E2’s harness — near-zero marginal cost. Stand up a second “candidate” target that is a known-wrong stub of forex-ng (e.g., swap one field’s transformation, drop a header, reorder a list) — a manufactured, minimal, known-in-advance divergence.
  • Design. Replay the same ~100 recorded fixtures from E2 against the wrong stub instead of the real legacy instance. The differ must (a) flag DIVERGED, (b) minimize to the smallest counterexample which must actually contain the manufactured defect, (c) never silently pass it as PARITY, (d) never crash on it.
  • Evidence required. The verdict output for the stub run, with the minimized counterexample diffed against the known defect by hand.
  • Success threshold. 100% of manufactured defects are caught as DIVERGED; the minimized counterexample is legible and traceable to the actual defect (no noise burying the signal).
  • Failure threshold. Any manufactured defect is missed (false PARITY — the more dangerous failure mode than false DIVERGED) or the counterexample is unreadable/too large to act on.
  • Cost. 0.5-1 agent-day beyond E2 (same fixtures, same harness, a second target). MODELED.
  • Sequence. Day 2-3, immediately following E2, same spike.
  • If false / decision triggered. A missed defect (false PARITY) is a correctness bug in the differ itself, not a tuning problem — it means the kernel cannot be trusted as a gate at all (PRD principle 1: “any nondeterminism in the harness itself is a defect of the same severity as a wrong verdict”). Block all further experiments until the differ is fixed and this test is rerun green; do not proceed to E4-E7 on an unproven differ.

E4 — Auditor-evidence dry run (retires A8)

  • Method. Not a build — a review meeting. Take the evidence bundle produced by E2/E3 (verdict + smallest counterexample + normalization ledger + replay recipe, PRD FR-V2) and walk it, unmodified, past whoever currently owns SOX/QSA evidence format at Groupon (identify this person as part of the experiment if unknown today — that itself is diagnostic).
  • Design. Present the bundle cold, without narrating around its gaps. Ask three questions: (1) “Is this the right shape of evidence for a DONE-row-9 sign-off?” (2) “What’s missing that you would need before signing?” (3) “Would you accept a replay recipe you can run yourself as a substitute for a live walkthrough?”
  • Evidence required. Written notes from the session; a named list of gaps, not a vague “looks fine.”
  • Success threshold. The reviewer confirms the bundle’s shape (verdict-first, falsifiable, sourced) is directionally right, even if specific fields are missing; no structural objection (e.g., “we need live production evidence, a replay recipe is not enough” would be structural).
  • Failure threshold. A structural objection that recorded-replay evidence cannot ever satisfy SOX/QSA regardless of polish — this would invalidate the bundle format for the money-gate portion specifically (functional-parity bundles for non-money capabilities are unaffected).
  • Cost. ~1 agent-day (assembling a clean bundle + the meeting + writeup). MODELED. Can run as soon as any real bundle exists — does not require Orders-scale data, a toy bundle from E2/E3 suffices for the shape question.
  • Sequence. Day 3-5, once E2/E3 produce a real bundle; independent of E5-E7.
  • If false / decision triggered. If the objection is structural, PRD §13’s claim that recorded replay can carry any part of DONE row 9 is falsified for that scope — narrow the bundle’s claimed authority to functional/long-tail parity only, and escalate the money-gate evidence question back into the E1 conversation (which lane’s evidence format actually satisfies the auditor is now a live open question, not a design detail).

E5 — Noise-rate measurement on a real endpoint (retires A7)

  • Method. Extends E2’s self-replay test from “does it work once” to “how often is it wrong” — the actual precision number nobody has ever published for this class of tool (D6 Q2).
  • Design. Pick one real, moderately busy forex-ng or orders-ext endpoint. Record a larger corpus (target: ≥1,000 requests across a realistic time window, not a synthetic burst). Self-replay the full corpus against the same legacy instance. Count any non-PARITY verdict as a false DIVERGED (there is no legitimate reason a self-replay against an unchanged instance should ever diverge).
  • Evidence required. Full verdict log for all ≥1,000 replays; a triaged list of every non-PARITY result with root cause (timestamp leak, ordering, float formatting, genuine bug).
  • Success threshold. Zero false DIVERGED across the full ≥1,000-replay corpus, matching the absolute bar FR-K8 (PRD §9.1) and the PRD §24 kill criterion both state for self-replay against an unchanged legacy instance (“must yield 100% PARITY” / “self-replay cannot reach zero false DIVERGED… must not ship”). A ≤1/1,000 rate is not a passing result here — this document does not get to loosen a ship-gate the PRD states as zero. Any non-PARITY result found during the run must be triaged in real time: if it traces to a named normalization gap, fix the rule and rerun until that case reads PARITY; only a residual with no traceable normalization fix counts against the threshold.
  • Failure threshold. Any false DIVERGED that survives triage without a traceable, fixable normalization cause — i.e., root causes that are structural (not fixable by adding a normalization rule — e.g., genuine legacy nondeterminism the differ cannot model). This is deliberately the same bar as PRD §24’s kill criterion, not a looser one: E5 is the experiment that tests whether that zero bar is achievable, not a place to redefine it downward.
  • Cost. 2-3 agent-days (corpus recording + triage of every non-PARITY result). MODELED.
  • Sequence. Day 5-8, after E2/E3 confirm the kernel works at all.
  • If false / decision triggered. This is the PRD’s own kill criterion (§24, third bullet: “self-replay cannot reach zero false DIVERGED… the noise model is unsound”). If the rate cannot be driven down after a bounded normalization-tuning effort (cap this sub-effort at 2 additional days), do not proceed to E6/E7 — fall back to S3 (impact-selected tests) per the documented rip-cord, and report this as a hard finding to Robert/Tomas before Orders depends on the kernel.

E6 — Orders-adapter spike (retires the wrap-not-rewrite premise behind A5; baselines A5 itself, partial A1/A6)

  • Method. Not a toy — a real spike wrapping the existing 3-oracle reconciliation pattern (crossCheckBillingFeed), read-only, on one real Orders read-path unit (PRD §18 M4).
  • Design. Take one signed HOT/WARM Orders read-path capability. Wire it through the kernel’s verdict interface consuming the existing billing-events feed (no parallel feed, per FR-A2). Declare provider-settlement and contract-check oracles UNVERIFIABLE (they don’t exist yet) rather than faking them. Produce one real PARITY verdict end-to-end with a full replay recipe.
  • Evidence required. The verdict output; a time/token log of the spike (this is also the A6 cost-per-unit data point); a written note on how much of crossCheckBillingFeed had to change vs. wrap verbatim.
  • Success threshold. One real PARITY verdict produced; the existing oracle is wrapped, not rewritten (retires the wrap-not-rewrite premise behind A5, confirming A2 from research.md, “generalize one working instance”); elapsed effort is recorded as the adapter-#1 baseline for the 2×-detection threshold (PRD §20 Stage 2) — this baselines A5 but cannot retire it: a ratio needs a numerator and a denominator, and adapter #2 does not exist until Stage 2.
  • Failure threshold. The existing oracle cannot be wrapped without a substantial rewrite (signals the “generalize the pattern” premise itself, not just the 2nd adapter, is wrong), or the internal-endpoint seam (PRD §25 Q2 — expose:false billing feed) blocks the read entirely without a resolvable workaround.
  • Cost. 3-5 agent-days. MODELED, scoped explicitly as a spike (throwaway-tolerant), not production-grade adapter code.
  • Sequence. Day 6-10, after E1 confirms the money-gate boundary (E6 must not accidentally build money-gate machinery E1 assigned elsewhere) and after E2/E3/E5 confirm the kernel is trustworthy enough to spend real Orders time on.
  • If false / decision triggered. If the seam blocks the read: this becomes the first concrete instance of PRD open question #2 (internal-endpoint seam) — resolve it (Core-side read adapter vs. an in-Encore hop) before Stage 1 hardening starts, since every subsequent Orders unit hits the same wall. If the wrap requires substantial rewrite: revise the adapter-cost model in PRD §17/§20 before committing to the 2×-detection threshold as written — the baseline itself may be wrong.

E7 — MBUS shadow-tap spike (retires A4)

  • Method. The single most expensive, most uncertain experiment — sequenced last on purpose, and deliberately not combined with an EOL boot. Rehearse on orders-ext (modern, non-EOL stack that shares a fraud/payment topic family with Orders, D5 §e / architecture.md L406) so a failure here is isolated to “can we diff a bus event at all,” not confounded with EOL-lab pain already tested in E2.
  • Design. Pick one real MBUS topic on orders-ext. Build a minimal shadow-tap: capture one event on the legacy side and the corresponding event on the candidate side, normalize (strip generated IDs/timestamps), diff payload + envelope. Explicitly test the outbox-durability angle at small scale: does “durably queued exactly once” hold, not just payload equality.
  • Evidence required. One real captured event pair; verdict output; a written note on what the shadow-tap mechanism actually required (library, protocol quirks, JMS/STOMP vs Kafka differences).
  • Success threshold. One real topic diffs cleanly to a PARITY/DIVERGED verdict with a legible counterexample when a manufactured mismatch is injected (same wrong-stub logic as E3, applied to a bus event).
  • Failure threshold. The bus protocol cannot be tapped non-invasively at all (would require legacy code changes, violating the no-modification constraint), or event ordering/at-least-once semantics make “diff” structurally undefined without invasive instrumentation.
  • Cost. 5-8 agent-days. MODELED, highest of the seven — reflects D1 §11’s “confirmed market gap, no prior art” finding; this is genuinely unknown-unknown territory, not a scoped build.
  • Sequence. Day 10-18, last. Depends on E1 (confirms MBUS shadow-taps are in-scope for the harness, not something Zaruba’s plane already owns per flip-condition 2 in build-vs-buy.md §5) and benefits from E2/E3’s kernel-correctness patterns being proven first.
  • If false / decision triggered. If bus tapping proves structurally infeasible non-invasively: this is the PRD’s “highest-uncertainty build” (§9.3 FR-A3) failing outright — MBUS parity would have to degrade to UNVERIFIABLE-by-default for all 16/33 MBUS-coupled repos, a major and immediate scope cut communicated to Tomas before Stage 1 hardening (PRD §20) is planned around it.

3. Sequence and decision calendar

Day Experiment(s) Gate to proceed
0-1 E1 (Zaruba conversation) Written answer to both questions, or escalate and proceed provisionally with Option B as working assumption (flagged)
1-3 E2 (EOL boot PoC) — parallel with E1 ≥95% self-replay PARITY on forex-ng
2-3 E3 (wrong-stub test) — same spike as E2 100% of manufactured defects caught, no false PARITY
3-5 E4 (auditor dry run) — parallel track once a bundle exists No structural objection to the bundle shape
5-8 E5 (noise-rate measurement) Zero false DIVERGED after bounded normalization tuning (PRD FR-K8/§24 bar)
6-10 E6 (Orders-adapter spike) — after E1 lane is clear One real PARITY verdict; adapter-#1 baseline logged
10-18 E7 (MBUS shadow-tap spike) — last, isolated from EOL risk One real topic diffs cleanly with a legible manufactured-defect counterexample

Total elapsed: ~18 working days (~3.5 calendar weeks with E1/E2/E3 and E4 partly parallel) — inside the week-6 Orders deadline with margin, by design (PRD §18’s own MVP gate is explicitly “before Orders is load-bearing”).


4. Maximum justified investment before the next go/no-go

Cap: 20 agent-days total across E1-E7, hard stop. MODELED, derived from: program M4 ≈ $0.9M (MEASURED, context brief); the harness is explicitly “Robert-lane, weeks not months” (context brief, binding); D3’s own lab-setup model tops out around 5 eng-days for the hardest single stack, and E7 (the single costliest experiment here) is deliberately scoped to a non-EOL rehearsal, not the worst-case JRuby-1.7 combination. Twenty days of validation spend against a $0.9M milestone and a TSB-class downside is proportionate; anything beyond it without a PARITY-shaped result starts eating into the delivery budget the PRD’s own kill criterion (§24, “harness became the schedule”) warns against.

This number does not exist elsewhere in the plan set — PRD §25 Q5 and build-vs-buy.md §5 Q5 both flag “the investment cap, numerically” as an open, unresolved DESIGN gap. This document sets it for the validation phase specifically (pre-go/no-go); it does not set the separate, larger per-flagship delivery cap that PRD §17/§24 still need as a named budget line — that remains open (§5.5 below).

At day 20, regardless of which experiments remain: stop, tabulate every experiment’s success/failure verdict against its stated threshold, and make one decision — proceed to PRD §18 MVP scope as designed, proceed with a named scope cut (e.g., MBUS degraded to UNVERIFIABLE-by-default, or money-gate evidence narrowed post-E4), or invoke the PRD §24 kill criteria and fall back to S3. No “extend and see” — the cap is a decision date, not a soft target.


Open questions

  1. Who owns QSA/SOX evidence sign-off at Groupon today? E4 cannot run until this person is identified — not currently named anywhere in the plan set (PRD §5 flags the persona itself as assumption).
  2. Does Zaruba’s plane already have informal MBUS-diffing precedent (script, one-off tool, tribal knowledge) that E7 would duplicate? build-vs-buy.md §5 flip-condition 2 names this as unchecked; a 30-minute conversation before E7 could save 5-8 days.
  3. What is the per-flagship delivery investment cap (distinct from this document’s 20-day validation cap)? Still DESIGN, not a number, per PRD §25 Q5 — should be set immediately after this validation phase concludes, using its real cost data instead of MODELED estimates.
  4. Does forex-ng’s traffic volume actually reach the ≥1,000-replay target for E5 within a reasonable recording window, or is a longer capture period or a busier endpoint needed?** Unverified — D5 characterizes forex-ng as “cheapest” (2 endpoints, no DB, no MBUS) but not as high-volume.
  5. What is E6’s fallback if the expose:false billing-feed seam (PRD §25 Q2) blocks the read entirely — is a Core-side read adapter buildable within E6’s own 3-5 day budget, or does it require a separate, unbudgeted spike first?