Skip to content
GRPNR.

Executive recommendation — Dual-run parity harness

2026-07-19 · Author: Fable (own synthesis over the full plan set + all discovery notes). This document supersedes initial-recommendation.md where they disagree, and states where they disagree.


1. Recommendation

BUILD — a limited MVP on an OSS foundation, in a pivoted shape. Not build-from-scratch, not buy, not abandon, and not the shape the initial recommendation described.

The composite, precisely (from build-vs-buy.md §4):

  • Build on OSS foundation for the kernel: adopt mitmproxy (recorder), apache/activemq-classic + cytopia/mysql-5.6 (lab base images), VCR/WebMock + WireMock (lab outbound-dependency isolation only). Pattern-only, zero code adopted: Diffy’s relative-disagreement noise floor, ApprovalTests’ versioned-baseline audit model, Scientist’s control/candidate vocabulary + write-safety default.
  • Build custom (the genuinely net-new surface, confirmed uncovered by any tool): deterministic differ, verdict engine (PARITY / DIVERGED / UNVERIFIABLE), content-addressed masked fixture store, MBUS shadow-taps (confirmed market gap — no OSS diffs JMS/STOMP/Kafka between two systems), evidence bundle, work-ledger verdict rows + structural PR-lane gate, CI runner.
  • Partner for the money gate: the live ≥99.999% ledger reconciliation stays in Zaruba’s oracle plane (Option B). The harness contributes the two missing oracles (provider-settlement, contract-check) and the generic verdict/evidence wrapper — never a second money gate.
  • Hold S3 (impact-selected tests) as the capped, documented fallback if the EOL lab blows its budget.

The pivot from the initial concept (state it, don’t bury it)

Three of the initial recommendation’s thirteen claims did not survive research:

  1. Recorded replay is NOT the money-gate mechanism. Every primary-sourced financial-grade cutover (Uber ×2, Stripe) ran live/continuous dual-run or dual-write; no counterexample exists (D6 Q4). Recorded replay is re-scoped to what it honestly does: functional parity of read paths and the long-tail/COLD/DARK surface where the legacy system is the only spec.
  2. The mission is reframed from “prove parity” (Frame A) to “make the legacy-off decision safe and defensible” (Frame B, research.md §1.2). Parity is the mechanism; the paid-for product is a defensible off-decision with a replayable counterexample. TSB is the negative control: nine successful rehearsals + board sign-off still failed catastrophically because the decision rested on judgment, not a structural gate.
  3. Two verdict tiers, one surface (Christensen lens): a cheap, fast, low-noise per-unit functional verdict for the fleet agent, and a rigorous live-reconciliation verdict for the program gate. One artifact for both either over-serves the agent or under-serves the gate.

What survived intact: the product identity (evidence machine for turning legacy off), CLI/CI-first with no UI beyond one Grafana panel, the structural gate (no verdict row → no PR lane), deterministic differ as sole verdict authority (LLM triage-only — fully corroborated), kernel + lab outside Encore, noise as the #1 usability risk, and the self-replay falsification recipe.

2. Confidence

  • HIGH on the portfolio shape and the build/partner/fallback split. It survived a full-strength steelman attack (research.md §4), an 8-option comparison, and three adversarial audit passes; the 001 flagship verdict is now evidence-hardened rather than assumed.
  • MEDIUM on feasibility until the validation experiments run. Three feasibility unknowns stack: EOL image bootability (live manifest-v1 deprecation risk), the noise floor (nobody — including Diffy — ever published a false-positive rate; our self-replay test is the only available evidence source), and the MBUS tap (no prior art anywhere).
  • MEDIUM on the boundary until Tomas ratifies. Option B is proposed, not agreed; his decision log has zero dual-run/shadow/parity entries.

3. Evidence balance

Supporting (strongest five). (1) No OSS/commercial tool covers the kernel; the bus-event-diff gap is total — confirmed by live repo-metadata verification of ~30 tools (D1). (2) Every org that did this at scale (Uber, Slack, Zalando, GitHub, Shopify, LinkedIn) built bespoke in-house dual-run tooling — building is the empirically normal answer (D2). (3) Verification is the program’s measured binding constraint, and the greenfield’s own gap list names the missing harness (ASSESSMENT.md top-5). (4) TSB: rehearsal + human sign-off without a machine gate is the documented catastrophic configuration (£48.65m fine, ~£366m total; medium-confidence secondary sourcing). (5) Cost ratio: Robert-lane weeks against a $0.9M M4 milestone that is unclaimable without proof.

Contradictory (strongest four, all absorbed into the shape rather than ignored). (1) Recorded replay has zero evidence of clearing a five-nines financial gate — hence Option B. (2) Diffy, the closest prior art, was abandoned by its own inventor and its precision was never published — hence the self-replay zero-false-DIVERGED ship-gate. (3) No published writeup solves the full EOL-lab difficulty stack end-to-end — hence the pilot ladder + rip-cord instead of a committed lab build. (4) The TDAD ~70% fallback number is a single uncorroborated off-domain study — hence S3 is a fallback, never the primary.

4. Critical unknowns (each mapped to a validation experiment)

Unknown Retired by
Money-gate lane: does Tomas ratify Option B + the binding structural gate? E1 (written, two binary questions, decision date)
Do ruby:1.9.3 / jruby:1.7 / openjdk:7 images still pull; does the easiest stack boot + self-replay clean? E2
Does the differ catch a known-wrong candidate with a minimal counterexample (no false PARITY)? E3
Will the SOX/QSA evidence owner accept the bundle shape? (Owner not yet identified — itself diagnostic) E4
Self-replay false-DIVERGED rate on a ≥1,000-replay real corpus — the number nobody has ever published E5
Does wrapping crossCheckBillingFeed work verbatim (adapter premise) + the expose:false seam E6
Can a bus event be diffed non-invasively at all? E7
DARK Orders batch surface (Resque/Quartz uninventoried); noise budget for the 99.999% row; numeric investment cap Stage-1 inventory pass; E1 conversation; ratification of the pre-mortem’s 40-eng-day proposal

5. Conditions that must be true for BUILD to stand

  1. E1 lands Option B ratified + gate confirmed binding (advisory-only gate = the anti-TSB argument collapses; re-score against no-new-software before further spend).
  2. E2 passes on forex-ng: if the easiest stack cannot boot and self-replay clean, recorded replay is dead in this estate and the harness shrinks to wrapper + oracles + S3.
  3. E5 reaches zero false DIVERGED after bounded tuning — the PRD’s own ship-gate; nonzero-and-unfixable kills the kernel as a gate.
  4. The investment cap becomes a ratified number (validation cap 20 agent-days is set; delivery cap proposed 40 eng-days through Orders M4, DESIGN, needs Robert+Tomas sign-off).

One verdict/evidence-bundle surface, two tiers. Tier 1 (fleet): recorded-replay functional verdicts, cheap and low-noise, gating PR-lane promotion structurally through work-ledger rows. Tier 2 (program gate): live oracle reconciliation in Zaruba’s plane, wrapped in the same verdict/evidence vocabulary, plus the two missing oracles. Underneath: parity kernel + legacy-runtime lab as a workspace package + CLI/CI runner outside Encore; per-domain adapters built strictly per flagship gate; MBUS shadow-taps custom, per-domain, rehearsed on a modern stack before ever touching an EOL boot. UX = the evidence bundle (verdict → smallest counterexample → normalization ledger → replay recipe), UNVERIFIABLE as an honest first-class state, one Grafana panel, nothing else. Architecture, data lifecycle (record→mask→freeze→replay→expire), threat model, and ADRs: architecture.md.

7. Immediate next steps

  1. Today: E1 — the Zaruba boundary conversation, two written binary questions with proposed defaults (Option B; gate = binding), 5-business-day decision window. Costs half a day and forecloses half the design space.
  2. In parallel (does not depend on E1): E2 — forex-ng boot + record 100 + self-replay PoC, 1–2 agent-days.
  3. Then E3 → E4 → E5 per the validation calendar (validation-plan.md §3): wrong-stub differ test, auditor dry run, ≥1,000-replay noise measurement.

Maximum justified investment before go/no-go: 20 agent-days across E1–E7, hard stop (validation-plan §4). At day 20: proceed to MVP as designed, proceed with a named scope cut, or invoke kill criteria and fall back to S3. No “extend and see.”

8. Pivot and kill criteria

Pivot triggers (build-vs-buy.md §5): lab blows its cap → S3 + oracle-only for that domain, keep the wrapper; internal MBUS-diff precedent surfaces → wrap it, don’t build taps; Tomas claims the whole verification lane → harness collapses to wrapper + adapters; an OSI-licensed Diffy successor or a bus-capable integration tool appears → re-open the differ/recorder decisions.

Kill triggers (PRD §24 + pre-mortem hardening): self-replay cannot reach zero false DIVERGED after bounded tuning → the noise model is unsound, do not ship the kernel as a gate; one confirmed false PARITY on a money-chain unit → freeze the rule version, re-verify everything issued under it (single-instance trip, not a threshold); a confirmed PII/PCI leak outside the lab → halt all recording program-wide; the E1 answer makes the gate advisory → the product loses its reason to exist as a gate; “harness became the schedule” with the cap exceeded and no PARITY-shaped result → S3, publicly.

9. Five-advisor verdict (principle-based simulations, per research.md §3 — not real opinions)

  • Jobs: back only the one-widget proof — one command, one artifact, one flagship; kill “kernel/adapter/generic” vocabulary until one real Orders unit goes claim → verdict → off. The MVP as scoped (M1–M6, pilot-first) satisfies this; the platform words stay out of the code until Stage 2.
  • Ive: back once honesty is structural — UNVERIFIABLE surfaced plainly, counterexample minimization a kernel requirement, and the self-replay proof green. A noisy report is dishonest material; the normalization ledger is the honesty mechanism.
  • Cagan: feasibility risk is HIGH and correctly attacked first — discovery (E1–E7) before delivery, the flagship schedule must not depend on an unproven lab. Backs investment only after the spikes.
  • Torres: the opportunity tree shows money (Opp 1) is mostly solved in Zaruba’s plane; discovery energy belongs on the invisible surface (Opp 4, batch/MBUS) and noise trust (Opp 3). Assumption test per branch before widening.
  • Christensen: two customers, two jobs, two service tiers behind one surface — never make the agent pay six-nines cost, never let the gate accept a token verdict. Backs it if the cheap tier beats re-reading the diff and the rigorous tier hits the live number.

Net: all five converge on the same conditional — build small, prove feasibility first, keep the money gate live and partnered, and let the structural gate be the product.

10. The ten hardest founder questions

  1. If Tomas answers that the promotion gate is advisory, not binding — do we still build anything beyond the two missing oracles? (The anti-TSB argument is the product’s spine.)
  2. Whose lane is the live ≥99.999% reconciliation — in writing, with a date — and what exactly is the shared verdict contract across the two planes?
  3. What is the numeric investment cap per flagship, who tracks actuals against it, and who pulls the S3 rip-cord when it’s hit — under week-6 pressure, not in a calm retrospective?
  4. If forex-ng — the easiest stack in the estate — cannot self-replay clean in two days, do we admit recorded replay is dead here and re-scope immediately, or do we “try one more stack”? (Name the answer now, before sunk cost exists.)
  5. Is voucher-inventory-service (Ruby 1.9.3 on JRuby 1.7, hand-patched Rails) bootable at all within a bounded time box — and if not, what verifies that repo when its flagship comes?
  6. Who owns QSA/SOX evidence format at Groupon today, and have they ever accepted a replay recipe in place of a live walkthrough? (Nobody on the plan can currently name this person.)
  7. What is the stated noise budget for the 99.999% row, and what structural source does it cite — what corruption/staleness rate does the legacy estate itself produce? (Uber cites S3 durability; Slack cites replica lag; we currently cite nothing.)
  8. Are we prepared to let one confirmed false PARITY freeze the entire verdict lane and force mass re-verification — accepting the schedule cost that single-instance trip implies?
  9. Who operates the harness at M3 write-ramp when Robert is unavailable, and who co-reviews normalization-rule changes so a solo operator cannot silently bless a real divergence?
  10. The DARK batch surface (uninventoried Resque/Quartz jobs) is structurally invisible to an HTTP recorder — who signs the risk of cutting over Orders with that surface uncharacterized, and at what inventory coverage does that signature become defensible?

Plan status: research complete. Next action is E1/E2, pending Robert’s acceptance of this recommendation.