Skip to content
GRPNR.

PRD — Dual-run parity harness (“the quality-control bench”)

2026-07-18 · Author: Fable · Plan 002, template §15. Binding inputs read in full: 00-context-brief.md, initial-recommendation.md, and this plan’s research.md, journey.md, user-stories.md, north-star.md, architecture.md, build-vs-buy.md, competitive-analysis.md. This is the buildable spec; where a section merely restates a decision already made in those docs, it cites and moves on rather than re-arguing.

Internal product (binding, context brief). No external customers, no pricing pages, no procurement cycles. Reinterpretations used throughout, stated once here: buyer = the program (Tomas’s gates + Robert’s investment call); procurement = lane approval from Tomas + security sign-off, not a purchase order; churn = the fleet routing around the harness (bypass); pricing = engineer-days + agent tokens. Every SaaS-template beat is read that way; where one does not survive translation it is scaled down in one line, not padded.

Evidence discipline (binding, Zaruba audit law). Numbers carry MEASURED / MODELED / DESIGN. Epistemic split: confirmed fact / supported interpretation / hypothesis / assumption / unknown. External claims carry URL + date. Advisory-lens content is labelled principle-based simulation. No invented market data, quotes, or endorsements.


1. Executive summary

The harness is an evidence machine for turning legacy OFF — a machine-checked parity verdict with a replayable counterexample, per signed capability, wired to the Definition-of-DONE rows that pay the program’s bonuses. It exists because verification, not codegen, is the program’s binding constraint (analysis/legacy/SYNTHESIS.md §3, confirmed fact) and because today exactly one domain slice produces a receipt (merchant-accounting’s ~300-LOC reconciliation oracle, MEASURED); ASSESSMENT.md lists “no parity/dual-run harness” as a top-5 gap.

Research reshaped the initial concept in three load-bearing ways:

  1. The mission is the defensible legacy-off decision, not parity itself (research.md Frame B). Nobody is paid for parity; they are paid for legacy provably off. Chasing 100% parity is rejected by every rigorous source found (Zalando: “fixing those last few percentages has a cost higher than the value it brings”, engineering.zalando.com 2021-11).
  2. Recorded replay does not clear the money gate. Every primary-sourced financial-grade cutover ran live/continuous dual-run, never recorded-fixture replay (Uber >99.999%, uber.com 2024-10-03; Stripe 99.9999%, stripe.dev 2024-02-16). The ≥99.999% ledger row is Zaruba’s oracle plane (Option B); the harness owns recorded-replay functional/long-tail/DARK parity plus the verdict/evidence wrapper both lanes share.
  3. The net-new product is the wrapper, not a platform. Build the verdict/evidence surface generic; build the lab and adapters strictly per-flagship-gate (build-vs-buy.md §4). No tool covers the kernel — best achievable stitch of existing tools is ~35-40%, entirely on the HTTP leg, 0% on the message bus (competitive-analysis.md §3.1, confirmed fact).

MVP is a walking skeleton proving the two riskiest assumptions — EOL-lab bootability and a zero-false-DIVERGED noise floor — on one cheap pilot before the Orders week-6 clock runs. Economics: program all-in $6.7M, M4 ≈ $0.9M (MEASURED); harness is Robert-lane weeks-not-months. It converts a $0.9M milestone from unclaimable to claimable and insures against a TSB-class false-off (supported interpretation).


2. Problem statement + evidence

Core problem. The program cannot turn legacy OFF without a machine-checked, replayable, auditable proof that the Encore.ts rebuild behaves like a 2013-2017 EOL estate nobody can safely run — produced fast enough that the proof is not itself the schedule.

Evidence it is real (confirmed fact unless tagged):

  • Verification is the binding constraint (analysis/legacy/SYNTHESIS.md §3).
  • One domain slice has a receipt; ~40 non-ledger R4 use cases, the read paths, and the DARK batch/MBUS surface have none.
  • 16/33 legacy repos are MBUS-coupled — event parity is day-one scope for any cutover.
  • Orders alone: 127 HTTP endpoints, 25+ MBUS topics + a Kafka signal, dual MySQL with a dedicated outbox DB, DARK Resque batch jobs (MEASURED, D5). MODELED 50-60+ fixture classes.
  • Dual-run must be live by week 6 or M4 slips publicly (reasons-and-decisions/03-orders-flagship.md).
  • No OSS/commercial tool covers the kernel; none diffs JMS/STOMP/Kafka events between two systems (D1, confirmed gap).

Cost of the problem. A wrong off-decision on the money chain is a TSB-class event (£48.65m fine, ~£366m total cost — Slaughter and May review via D2 §9, medium-confidence secondary sourcing). TSB is the negative control: nine successful dress rehearsals + a 1,600-person pilot + board sign-off still produced catastrophic full-scale failure — because the decision rested on human judgment, not a structural machine gate.

Why today’s solution is insufficient. Zaruba’s oracle covers one domain and even there only 1 of 3 oracles is built (D4). Process (R4 + panel + human sign-off) is exactly the TSB-failure configuration when relied on alone. Vibes do not survive 64 concurrent agents, and there is no surviving published evidence that shadow traffic catches AI-agent regressions specifically (D6, CONFIRMED gap) — the harness must generate that evidence for itself.


3. Goals / non-goals

Goals.

  • G1. Produce a deterministic, replayable, auditable verdict (PARITY / DIVERGED / UNVERIFIABLE) per signed capability, gating PR-lane promotion.
  • G2. Characterize legacy behavior where no independent oracle exists (read paths, long-tail, DARK batch/MBUS) via recorded replay against an EOL lab.
  • G3. Manufacture the evidence bundle Tomas and QSA/SOX sign (DONE row 9).
  • G4. Serve two customers at two service levels behind one surface: a cheap per-unit functional verdict for the fleet, a rigorous reconciliation verdict for the money gate.
  • G5. Prove EOL-lab bootability and the noise floor on a cheap pilot before the Orders schedule depends on them.

Non-goals (explicit, north-star.md §10):

  • No live production traffic ownership. If the ledger gate needs live traffic, that machinery is Zaruba’s oracle plane (§13).
  • No performance parity in v1. The p99 ≤ +10% DONE row is production telemetry the program owns; the harness proves functional parity only.
  • No UI product. CLI-first, CI-native, JSON verdicts; one Grafana panel for humans.
  • No generic QA platform. Adapters generalize only per flagship gate.
  • No second money gate. The harness wraps Core’s reconciliation; it does not rebuild it.
  • No LLM equivalence gate. The deterministic differ is the sole verdict authority; an LLM may triage, never pass.

4. Product principles

From north-star.md §6 / architecture.md §1, restated as build invariants:

  1. Determinism is the product, not a feature. A verdict is a pure function of (fixture set, adapter, normalization rules). Any nondeterminism in the harness itself is a defect of the same severity as a wrong verdict.
  2. Structural gates beat human judgment under schedule pressure. A unit cannot reach the PR lane without a verdict row; there is no “probably fine” path (anti-TSB, D2 §9).
  3. Noise normalization is the entire game, not polish. Kernel v0 scope. False DIVERGED → trust collapse → bypass is the top usability risk.
  4. Nobody chases 100% parity — and says so with a number. State the noise floor and its source; never treat nonzero divergence as automatic failure.
  5. Side-effecting paths are never executed twice. Recorded replay by default; mutating verbs off unless explicitly enabled (Diffy’s own default).
  6. The report states its own limits. Every verdict ships a normalization ledger (what was ignored + why) and a replay recipe (how any reader falsifies it). “A parity report that hides its exclusions is a lie with good typography.”
  7. The tool proves its own claims or it doesn’t ship them. Self-replay-to-PARITY is the harness’s only falsification recipe for its own noise model — no external false-positive baseline exists to borrow (D6 Q2).
  8. Adapters are the product; the kernel stays minimal. Extract kernel API from the second adapter, never speculatively.

5. Personas

Only stakeholders with an evidenced role (journey.md §1, user-stories.md).

Persona Type Interface Owns
Fleet agent Machine (primary) CLI harness verify, JSON verdict, work-ledger row Claims a unit, runs verify, consumes verdict mechanically
Strike-unit builder (×4) Human (primary) Evidence bundle, CLI Resolves DIVERGED: real regression / acceptable difference / new normalization rule
Robert Human — operator/investor CLI, Grafana, repo Investment decision, lab lifecycle, fixture cadence, adapter-scope calls
Tomas Human — program owner Evidence summaries, DONE-row report Flagship gate approval, adapter-scope expansion, lane/oracle boundary, “week-6 slip” call
QSA/SOX auditor Human — evidence consumer Evidence bundle + normalization ledger, read-only DONE row 9 sign-off; needs falsifiable replayable proof
Security reviewer Human — evidence consumer / gate Masking spec, lab isolation design PII/PCI masking + legacy-credential isolation sign-off before any fixture leaves the lab
Fleet ops Human — operator Lab, recorder, Grafana Boots lab images, keeps fixtures fresh, panel truthful

Note: the QSA/SOX profile is assumption, not measured — no auditor has touched program artifacts yet (journey.md §1.5). Confirm before finalizing the evidence-bundle format.


6. Jobs-to-be-done

Two distinct customers hiring the harness for two different jobs; conflating them is the platform-creep trap (research.md §3.5, Christensen lens).

  • JTBD-1 (fleet agent): “When I claim a work-ledger unit, tell me whether my change is safe to merge, cheaper and faster than re-reading the diff, without crying wolf.” Bar for firing vibes: low-noise, fast, mechanical. Over-serving it with money-gate rigor is over-shoot — keep the per-unit verdict lightweight (S2/S3).
  • JTBD-2 (program gate): “When I decide to turn legacy OFF, give me irreversible-decision-grade proof I can sign and defend to a regulator.” Bar: a structural gate a human cannot override with confidence alone, backed by a live reconciliation number with a stated noise budget. This customer wants more rigor — under-serving it is fatal, and this is where live oracle reconciliation belongs (S4/S1), not recorded replay.

Resolution: two verdict tiers behind one vocabulary — never confused, both auditable.


7. Journey summary

Full analysis in journey.md; not duplicated. The core loop: agent claims a unit → harness verify <unit> → replay recorded fixtures against the booted legacy lab AND the candidate Encore service → normalize → diff → verdict. PARITY promotes the unit to the PR lane with an evidence link; DIVERGED/UNVERIFIABLE blocks it and (on real regression) mints a deduped triage question. The human path is the exception: a builder opens the evidence bundle, reads the smallest counterexample first, and decides real-regression / acceptable-difference / add-normalization-rule — the decision itself an auditable scope record. Stakes escalate at the money chain, where the verdict is fed by Zaruba’s live reconciliation, not the harness’s recorded replay. Ten journey stages (onboarding → recording → verify → verdict → promote-or-block → triage → normalization decision → DONE-row rollup → retirement) with owners and failure states are tabulated in journey.md §2; the capability→FR mapping is in journey.md §4 and realized as the numbered FRs below.


8. User stories — critical path only

Full set (US-1…US-19) with acceptance criteria, edge cases, and instrumentation in user-stories.md; not duplicated. Critical path:

  • US-1 boot a pinned-EOL lab image, credential-isolated. US-3 security signs masking rules before any fixture leaves the lab. US-4 record ~100 requests through an outside-in proxy. US-7 harness verify returns a machine-readable verdict within one CI-step timeout. US-8 first real PARITY on a real Orders-chain unit (activation). US-9 self-replay always yields PARITY (the harness’s falsification recipe). US-10 every normalization-ledger change reviewed + logged with rationale. US-11 evidence bundle shows smallest counterexample first. US-13 work-ledger structurally refuses PR-lane promotion without a PARITY verdict row. US-14 one DONE-row-8 rollup combining harness verdicts + Zaruba’s reconciliation. US-17/US-18 auditor independently replays a verdict and reads the normalization ledger as a plain dated list.

Verdict contract (DESIGN, user-stories.md): JSON on stdout, one object per unit; exit codes 0 PARITY · 1 DIVERGED · 2 UNVERIFIABLE · >2 harness internal error (CI must retry once then page Robert, never treat as pass/fail).


9. Functional requirements

Grouped. Each FR is a buildable capability the journey/stories demand. MVP tag marks walking-skeleton scope (§18).

9.1 Kernel (record → normalize → replay → diff → verdict)

  • FR-K1 (MVP) The kernel records legacy traffic through an outside-in network-layer proxy without modifying legacy code — adopt mitmproxy (MIT, 44k★, pushed 2026-07-18, D1 §3) as an external process. VCR-class in-process capture is rejected (requires legacy Gemfile changes, stuck on a 2015-era line for the oldest stacks, D3 §5).
  • FR-K2 (MVP) Recording covers three traffic classes with distinct mechanisms: proxy for sync HTTP; shadow-tap for MBUS + outbox; DB-transaction-log tailing or job instrumentation for batch/cron. Batch/cron is structurally uncapturable by an HTTP proxy (D5 §d) — the recorder emits UNVERIFIABLE(batch-class) rather than silently omitting it.
  • FR-K3 (MVP) Fixtures are content-addressed (hash of masked content = identity), immutable once frozen, versioned by (domain, class, stack_version, rule_set_version). BUILD — no adoptable tool is content-addressed + masked + cross-language (D1).
  • FR-K4 (MVP) The normalizer strips known noise (timestamps, UUIDs, ordering, float formatting) from both sides before diffing and records every rule it applied into the normalization ledger. Two strategies: (a) explicit hand-curated ledger (the load-bearing feature), (b) A/A self-replay noise floor (Diffy’s relative-disagreement-rate design, not its code — opendiffy is CC BY-NC-ND, legally blocked; twitter-archive is abandoned, D1 §2a).
  • FR-K5 (MVP) The differ is a total, deterministic structural diff over normalized pairs (JSON/raw bodies, allowlisted headers, status codes, bus envelopes/payloads), emitting an ordered minimal counterexample. Never throws on malformed input — emits a structured “unparseable” that becomes UNVERIFIABLE.
  • FR-K6 (MVP) The verdict engine emits PARITY | DIVERGED(smallest counterexample) | UNVERIFIABLE(reason). Classification is three-way at the domain level (posted / received-expected-no-op / unpublished-gap), mirroring the ConsumptionStatus enum (interfaces/merchant-accounting.interfaces.ts:62-66) — a naive “not in ledger ⇒ DIVERGED” misclassifies every lifecycle/void event (REDEEMED, UNREDEEMED, FULFILLED, VIEWED, CANCEL, RESCIND_AUTHORIZATION) that is correctly received but never posted (D4 §b). Ambiguous → UNVERIFIABLE, never a coin-flip.
  • FR-K7 (MVP) Counterexample minimization is a kernel requirement, not a polish pass. When minimization itself fails (diff too large/structural), the bundle states so rather than dumping the raw diff.
  • FR-K8 (MVP) Self-replay gate: record from the legacy lab, replay against the same legacy instance — must yield 100% PARITY. A kernel or normalization-ledger version that fails its own self-replay cannot ship. This is the harness’s only precision evidence (D6 Q2).

9.2 Legacy-runtime lab

  • FR-L1 (MVP for 1 image) The lab boots pinned-EOL stacks via docker-compose profiles. The Orders chain needs 6 distinct pinned runtime/framework images, no reuse (MEASURED, D5 §c): Ruby 2.4.6/Rails 3.2 (orders), Ruby 1.9.3-on-JRuby-1.7/Rails 3.2 hand-patched (voucher-inventory, hardest boot), Ruby 2.6.10/Rails 4.1 (accounting), Ruby 2.7.5/Sinatra (users), Java 8/Spring (billing-record), Java 8-11/Dropwizard (orders_mbus_client).
  • FR-L2 (MVP) Boot is deterministic from a pinned manifest — no :latest, no live bundle install against upstream rubygems at boot. Mitigations baked in from day one: rescue-pull old images to an internal registry (works around Docker manifest-v1 deprecation, docker-library/ruby#452, live/unresolved 2026-07-18, D3 §1); vendored gems / internal gem mirror (works around the RubyGems TLS-1.2 floor since Jan 2018, D3 §2).
  • FR-L3 (MVP) If an image will not pull or the app fails its boot health check, the run is UNVERIFIABLE(lab_boot_failed) — never a false verdict against a half-booted app, never a silent fallback to a stub.
  • FR-L4 The lab is a hostile-adjacent sandbox: no egress (default-deny), internal-only Docker network, non-root USER, read-only root FS + tmpfs scratch (D3 §6). Every image is EOL by design and carries unpatched CVEs; EOL-alerting allowlists lab-tagged images while staying strict fleet-wide.
  • FR-L5 Lab-boot duration, success/failure, and image digest are logged per boot (audit trail for “which lab produced this fixture”).
  • FR-L6 Adopt current OSS base images where they exist: apache/activemq-classic (ActiveMQ 5.x is actively maintained, not truly EOL — lowest lab risk, D3 §1), cytopia/mysql-5.6 (third-party, fills the stale-official-image gap). Adopt WireMock/Mountebank/VCR narrowly to stub what legacy calls out to, so the lab boots deterministically.

9.3 Adapters (per-domain, the product)

  • FR-A1 (MVP: Orders read path) An adapter implements: fixture classes + how to record them; the independent oracle source(s); domain-specific normalization rules; thresholds (from program config); counterexample shaping. Shape (DESIGN): Oracle<TFeedEvent, TCheckSource> = (feedWindow, independentSource) => { posted | gap | unverifiable }[].
  • FR-A2 (MVP) The Orders adapter wraps the existing 3-oracle reconciliation almost verbatim (D4 §b): order-store cross-check (built: crossCheckBillingFeed, exposed as _reconciliationOrderCrossCheck), provider-settlement (typed stub only — interface fixed at interfaces/merchant-accounting.interfaces.ts:171-188, endpoint _reconciliationProviderSettlement is TODO pending the payment adapter), contract-check (not built, not stubbed). Do not assume all three exist — only 1 of 3 is in code. The adapter consumes the existing billing-events feed (_billingEventsFeed.controller.ts:1-38, GET /order/internal/billing-events, keyset-paginated, expose:false auth:false) — it must not invent a parallel feed (D4 §a.1).
  • FR-A3 MBUS shadow-taps capture and diff async events — legacy JMS/STOMP + Kafka — including the outbox durability contract (“durably queued exactly once”), not just payload parity. Orders has 25+ jms.topic subscriptions + Kafka txn.signals.high; orders_mbus_client consumes 8 more; Orders runs its own outbox DB (orders_msg_prod). Highest-uncertainty build, no OSS prior art (D1 §11). Harness-internal only — the standalone MBUS bridge is DROPPED (binding, 001).
  • FR-A4 The kernel API is extracted from adapter #2, not designed speculatively (principle 8). Adapter cost is tracked explicitly against the 2×-first-adapter detection threshold (§17).

9.4 Verdict / evidence

  • FR-V1 (MVP) The verdict type is three-state, not boolean; no existing repo type has this shape (closest is ReconciliationReport.balanced: boolean, D4 §a.3).
  • FR-V2 (MVP) Each verdict ships an evidence bundle: verdict, smallest counterexample, normalization ledger (what was ignored + why), and a replay recipe (documented CLI commands, no harness-team-specific knowledge). Bundle is content-addressed and auditor-grade (DONE row 9). Rendering follows Ive-lens ordering: verdict → smallest counterexample → everything else behind links.
  • FR-V3 (MVP) UNVERIFIABLE is a first-class state, not an error — it blocks promotion exactly like DIVERGED but routes to a different triage queue (record more traffic / fix lab vs. explain a divergence). A retry loop must never silently treat UNVERIFIABLE as PARITY (correctness requirement).
  • FR-V4 (MVP) Normalization-ledger writes are append-only with field/topic, rationale, author, date, and which fixture classes the rule was validated against (self-replay). No silent edits; a “change” is a new version referencing the old (supersedes chain).
  • FR-V5 Any past verdict reproduces exactly from its cited (fixture_set_hash, rule_set_version) — or the harness explicitly reports it can no longer do so (post-PII-destruction). Never silently returns a different answer for the “same” historical query.

9.5 Integrations

  • FR-I1 (MVP) work-ledger client writes append-only verdict rows into an Encore-side service — the only harness write into Encore. New service self-registers via registerInternalService; append-only rows use keyset pagination; writes into any Core-owned table go through a Core proposal endpoint (submit → validate → apply/reject → audit log), never a direct table write; migrations via pnpm drizzle only (D4 §d).
  • FR-I2 (MVP) Triage hook: on DIVERGED = real regression, mint a deduped question referencing the evidence-bundle URI, reusing Tasks INFORMATION_REQUEST + createFromWorkflow idempotent mint path (D4 §a). The triage service itself is out of scope (separate plan). Two frictions to resolve at build: minting requires an owning Temporal workflow run (thin wrapper workflow, accept the coupling); no assigneeGroupSlug registry exists (define the first slug set). Completion is a Postgres write via complete() (Admin-ACL-gated) or completeByToken(), which signal back into the owning workflow (_workflowSignalSend) — the harness reads “triage answered” as an external signal into work-ledger, not as a Task read (D4 §a).
  • FR-I3 Zaruba-oracle read (read-only): generalize crossCheckBillingFeed’s shape (reconciliation.utils.ts:79-104); read the billing-event feed (_billingEventsFeed.controller.ts:1-38, expose:false auth:false) via keyset pagination — today consumed only by merchant-accounting’s reconciliation service and its own subscription (billingEvent.subscription.ts, topic billing-event); the harness becomes its second consumer ever (D4 §a). The harness never writes into merchant-accounting.
  • FR-I4 (MVP) Grafana/OTel panel (reuse existing local-first fleet monitor, not a new product): parity coverage, divergence rate, self-replay pass rate, UNVERIFIABLE rate, fixture staleness, lab-boot success. Aggregate counts only, no PII.

9.6 CI

  • FR-C1 (MVP) A new backend-ci.yml (first backend CI in the repo, D4 §e) runs pnpm validate (root package.json:20: check:ci && docs-sync --check && type-check && test) + harness verify, path-triggered, git diff-scoped, concurrency-cancelled, timeout-bounded — mirroring frontend-ci.yml.
  • FR-C2 (MVP) Structural promotion gate: work-ledger’s promotion transition requires a verdict-row reference with exit code 0 for the unit’s current fixture-set hash, enforced at the data/proposal-endpoint layer — not a convention, not a skippable CI check, not a lint rule. A unit whose implementation changed after verification has its verdict invalidated by content-hash mismatch and must re-verify. Fail closed (work-ledger unavailable → promotion blocked). This is the harness’s only truly load-bearing structural claim; every other property degrades gracefully, this one cannot (journey.md §3).

10. Non-functional requirements

  • NFR-1 Determinism. harness verify is pure given (fixture set, adapter, rule set) — no wall-clock, RNG, map-ordering, or locale leakage into the diff. Same inputs → same verdict byte-for-byte, on any machine. Target: ~100% verdict reproducibility on repeat replay with no upstream change (Reliability metric, §16).
  • NFR-2 Verify latency budget. A verdict must return within one CI-step timeout (DESIGN — a slow or flaky harness is worse than none; agents retry-loop against it, journey.md §1.1). Tracked as time-to-value; exact budget calibrated on pilot volume.
  • NFR-3 Fixture safety. No fixture is written to durable storage outside the lab boundary until masking-rule sign-off exists for its fixture class; every fixture write logs its masking-ruleset hash; an unmasked-fixture scanner runs as a CI gate on the fixture store.
  • NFR-4 Isolation. Legacy creds live only inside the lab’s isolated network namespace, never in Encore secrets or any fixture that leaves the lab (binding). The two secret domains never mix.
  • NFR-5 Self-test as regression gate. The self-replay suite (FR-K8) runs in CI on every kernel/normalization-ledger change, not just locally; green is a merge precondition.
  • NFR-6 Falsifiability. Every stated number in an evidence bundle carries source + window + confidence; any reader can re-derive the verdict from the replay recipe without harness-team assistance.

11. Data requirements

Entity model in architecture.md §4. Load-bearing points:

  • Fixtures are content-addressed, masked, versioned by stack + rule-set; carry retention_class ∈ {PII, PCI, non-sensitive} and a mask_manifest (which rules masked which fields). Source of truth for functional parity = the frozen fixture set (the live legacy estate cannot be kept always-on, D1 §2a). For the money gate, source of truth = the billing-event feed in Zaruba’s plane.
  • Masking happens before a fixture is frozen and leaves the lab, inline during capture, never after landing. No off-the-shelf tool masks recorded HTTP fixtures (GoReplay’s masking is a paid Pro feature, D3 §7, confirmed) → custom scrubber, BUILD. Field scope: PAN, IBAN/SEPA, address, device/fraud IDs, email, OAuth tokens (D5 §b — every Orders-chain fixture class touches PII/PCI somewhere). Ambiguous field → default to mask, escalate, never default to record raw.
  • Retention is retention-class-driven: PCI shortest-lived + access-logged; GDPR-erasure fixtures (orders_mbus_client consumes gdpr.account.v1.erased by construction) must themselves honor erasure — a fixture derived from an erased account is deleted on erasure propagation. Exact windows per class are unknown — needs security sign-off before any fixture is frozen (§19).
  • Retirement archives lab image definition + fixture set + verdict history + normalization ledger as one retained, queryable, falsifiable unit; never deletes PII-bearing fixtures without a separate security retention sign-off; never reuses a domain’s work-ledger unit IDs.

12. Integrations, roles/permissions

Integrations (§9.5): work-ledger (write, proposal endpoint), Tasks/triage (mint), Zaruba oracles (read-only feed), CI (backend-ci.yml), Grafana/OTel (metrics). Seam to resolve (D4 §d, confirmed): the billing-event feed and reconciliation reads are expose:false, reachable only via Core’s in-cluster generated client. An external CLI/CI runner cannot call them directly → a Core-side authenticated read adapter, or that one hop runs inside Encore. Real seam, not assumed away (§19).

Roles / permissions.

Action Who may perform Rule
Change a normalization rule Robert (MVP scale); strike unit co-review post-MVP Append-only + rationale + self-replay-passed; auditable scope decision
Approve a fixture-masking spec Security reviewer Per fixture class, before recording; recording without sign-off = security incident, fixture quarantined + destroyed
Approve adapter scope beyond Orders Tomas Must map 1:1 onto R4 shadow-checkability classes; no invented fourth class
Set verdict thresholds (99.999% etc.) Program-owned config, adapter-declared Harness reads the number, never invents it (ADR-5)
Read raw pre-mask recordings Nobody outside the lab No audit exception; auditors get masked fixtures only
Override the promotion gate No one Structural; there is no audited-override path in MVP

13. Money-gate boundary (the load-bearing lane decision)

Zaruba’s R4 defines dual-run over three shadow-checkability classes — YES / COMPARE-ONLY / NO — and rows 24/25 (order money → ledger) say verbatim “YES — this IS the reconciliation harness” (D4 §c). Recommendation: Option B — the live ≥99.999% ledger-reconciliation gate belongs to Zaruba’s oracle plane (architecture.md §2.3, ADR-6). Rationale: every financial-grade precedent runs live continuous, not recorded replay (D6 Q4); the computation already exists in Core; recorded replay has zero evidence of clearing this row. The harness contributes exactly two things to the money gate: (i) the two missing oracles (provider-settlement, contract-check — cheap independent-source oracles, not lab work), and (ii) the generic verdict/evidence wrapper so the reconciliation output becomes an auditable, PR-lane-gating artifact instead of a boolean logged by a cron. This boundary is proposed, not ratified — Zaruba’s decision log has zero dual-run/shadow/parity entries; needs an explicit Robert↔Tomas lane decision + a shared verdict contract before M3 (§19).


14. Security / privacy

Threat model in architecture.md §7. Non-negotiables:

  • EOL image exploitation → hostile-adjacent sandbox (FR-L4).
  • PII/PCI leakage → mask-before-freeze; mask-verification gate quarantines any fixture tripping a PII scan; retention-class expiry; access-logged PCI class. Masking coverage is proven per fixture class, not asserted once for the whole harness (journey.md §1.6).
  • Accidental replay against prod → replayer enforces a target allowlist (refuses any host not on the explicit candidate/lab list); lab has no network route to prod; mutating verbs OFF by default.
  • Legacy credential exposure → creds stay in the lab, never in Encore secrets or CI logs; lab env is compose-local, never committed.
  • Poisoned fixture → content-addressing (tampering changes the id); immutable once frozen; the self-replay gate fails if a fixture no longer reproduces PARITY against its own stack.

LLM run evidence (global rule): the advisory triage assistant (post-MVP) runs only on DIVERGED, batch-clustered, cache-keyed by counterexample hash, hard per-run token cap; every run logged in the project’s LLM run-evidence ledger with lane + billed cost re-summed from real usage. No silent spend.


15. Operational requirements

  • Local-first. Runs on WSL2 + docker-compose; no cloud dependency until the billing gate clears (binding). Kernel = workspace package + CLI; lab = compose profiles; Encore surface = verdict rows.
  • Fixture freshness. A scheduled (weekly, DESIGN — no measured basis yet) lab-image rebuild + self-replay check; for genuinely frozen legacy (orders ~2017) this degrades to “did our lab recipe rot” (TLS/apt drift), not “did legacy change” — distinguished in the report. A failed rebuild does NOT auto-invalidate existing fixtures (avoids a cascading false-DIVERGED storm from unrelated lab-infra breakage).
  • Harness self-observability. Self-replay pass rate (must be 100%), parity coverage, divergence rate, UNVERIFIABLE rate (rising = lab-boot rot), fixture staleness, lab-boot success — all to Grafana/OTel, local-first.
  • Setup cost. MODELED ~2-5 eng-days first repo per stack combination, <1 day follow-on (D3 §8 — the follow-on figure is inference by analogy, not sourced; tracked against the Cost metric to test the model).

16. Dependencies, constraints

Dependencies.

  • Zaruba-lane boundaries (§13): the money gate is Zaruba’s plane; the harness depends on the billing-event feed and on the R4 shadow-checkability classification, but owns the recorded-replay surface and the shared verdict contract.
  • Billing gate: no cloud spend until it clears — the harness stays local-first until then.
  • Adopted OSS: mitmproxy (recorder), apache/activemq-classic, cytopia/mysql-5.6, WireMock/VCR (lab isolation). Pattern-only (zero code): Diffy (noise floor), ApprovalTests (baseline audit model), Scientist (control/candidate vocabulary + write-safety default), Hoverfly (live-vs-recorded shape).
  • Existing infra: Grafana/OTel fleet monitor, triage-dedup service, Tasks workflow (INFORMATION_REQUEST + createFromWorkflow, tasks.service.ts:43-84) — reused, not built. Corrected per d4 recon: work-ledger’s Encore-side verdict-row surface does not exist in Zaruba’s repo today (D4 §a, confirmed) — the work-ledger service itself is new build (FR-I1), not reused infra; only the fleet-level unit-claiming concept it extends is pre-existing.

Constraints (binding).

  • Kernel + lab live outside Encore — Cloud Run cannot run pinned EOL images or tap MBUS (architect correction #4). Encore surface = verdict rows only.
  • Terminology is binding: work-ledger (never bare “ledger”), parity kernel, legacy-runtime lab, adapter, verdict, evidence bundle, normalization ledger, shadow-tap.
  • Legacy code is never modified to record it (rules out in-process VCR).
  • The deterministic differ is the sole verdict authority; LLMs triage only.

17. Risks

Risk Evidence Mitigation in scope
Diff noise → false DIVERGED → bypass Top usability risk; Diffy’s FP rate never published (D6 Q2) Noise normalization = kernel v0; self-replay gate; UNVERIFIABLE as honest third state; pause a noisy adapter rather than tolerate it
EOL lab won’t boot at acceptable cost Highest-uncertainty tier; manifest-v1 + TLS + native-gem, no end-to-end solve published (D3) Pilot ladder de-risks before Orders depends on it; rescue-pull + vendored gems; S3 rip-cord
MBUS shadow-tap has no prior art Confirmed market gap (D1 §11) Rehearse on a modern non-EOL stack (orders-ext) before combining bus-tap with an EOL boot; budget as top-risk line item
Recorded replay can’t clear 99.999% Every financial-grade precedent runs live (D6 Q4) Option B — money gate is Zaruba’s plane; harness does functional parity
Harness doesn’t generalize past Orders 001 risk #1 Extract kernel from adapter #2; detection = 2nd adapter >2× the first
“Harness became the schedule” — and its opposite, under-resourcing 001 pre-mortem #2; Uber Invoicer 10% in 6mo as side-work (D2 §6d) Cap investment per flagship + named S3 fallback trigger; harness is a named workstream, not ambient side work
Auditor requirements are assumption, not measured No auditor engagement yet (journey.md §1.5) Confirm before finalizing bundle format; design to falsifiability regardless

18. MVP scope — walking skeleton

The smallest credible product that tests the most important assumptions (EOL-lab bootability + zero-false-DIVERGED noise floor), on the cheapest pilot, before the Orders week-6 clock runs. Not a compressed full vision — deliberately one cheap endpoint chain end-to-end plus an Orders-adapter spike.

Pilot: forex-ng (ranked cheapest — 2 endpoints, no DB, no MBUS, D5 §e), then rehearse the EOL boot on giftcard_service (Rails 3.2) and the MBUS tap on orders-ext (modern stack, shares a fraud/payment topic family with Orders) as capped throwaway spikes.

# MVP item FRs Acceptance criteria
M1 Parity kernel (record → normalize → replay → diff → verdict, deterministic, content-addressed) K1-K7, V1-V3 On a fixed (fixture set, adapter, rule set), harness verify returns the same three-state verdict byte-for-byte across ≥3 machines/runs; malformed input yields UNVERIFIABLE, never a crash or a guessed PARITY
M2 One pilot lab image booting one endpoint chain L1-L3, L5-L6 forex-ng boots deterministically from a pinned manifest (no :latest, no live bundle install); one smoke request returns the manually-verified oracle response; boot digest logged. Plus a dated pass-or-fail recipe for the giftcard_service Rails-3.2 boot — a written failure with the exact blocker (manifest/TLS/gem-build) is an acceptable MVP outcome, silent abandonment is not
M3 Self-replay falsification test (record → replay same legacy instance → PARITY) K8, NFR-1, NFR-5 Self-replay yields 100% PARITY on the pilot’s real recorded corpus, run in CI; any kernel/ledger change that breaks it is blocked from merge
M4 Orders-adapter spike wrapping the existing 3-oracle reconciliation on the Orders read path A1-A2, I3 One real Orders read-path unit goes claim → verify → verdict, consuming the existing billing-events feed (no parallel feed); the built oracle (crossCheckBillingFeed) is wrapped verbatim; provider-settlement/contract-check are declared UNVERIFIABLE, not faked. Produces the first real PARITY on a real Orders unit with a full replay recipe (US-8 activation)
M5 Verdict rows + structural gate in work-ledger V4-V5, I1, C1-C2 A unit cannot reach the PR lane without a matching current PARITY verdict row, enforced at the data/proposal-endpoint layer; a stale verdict (post-change hash mismatch) forces re-verify; work-ledger unavailable → promotion blocked (fail closed)
M6 Evidence bundle + masking gate + one Grafana panel V2, K2 (masker), I4, NFR-3 Every verdict ships a bundle (verdict → smallest counterexample → normalization ledger → replay recipe); no fixture leaves the lab without per-class masking sign-off + a passing unmasked-fixture scanner; one panel shows coverage + divergence + self-replay health from verdict rows

MVP gate = the two riskiest assumptions retired on the pilot before Orders is load-bearing: (a) an EOL runtime is bootable/recordable at acceptable cost, (b) self-replay yields zero false DIVERGED on a real corpus.


19. Explicitly excluded (with reasons)

Excluded Reason
Live production traffic shadowing owned by the harness Every rigorous case paid a real doubled-traffic tax + side-effect-suppression machinery that overlaps Zaruba’s plane, not ours (D2); if the ledger gate needs it, it is a Zaruba-lane decision, not a harness feature (§13)
Second money gate The money computation already exists in Core; rebuilding it is waste (build-vs-buy.md §1a)
LLM as verdict authority 2024-26 LLM-judge-for-equivalence literature is benchmark-stage; every survey frames it as “triage and assistance, not definitive gates” (D6 Q6, arxiv 2510.24367). Advisory triage only, post-MVP, non-authoritative
Generic multi-domain platform up front Premature generalization is 001 risk #1; extract kernel from adapter #2 (principle 8)
Standalone MBUS bridge app Dropped in 001; shadow-taps live inside adapters only (binding)
Performance parity (p99 ≤ +10%) Production telemetry the program owns, not a functional-parity harness function
Embedding-based divergence clustering, auto-generated adapters, DB-value diffing, multi-region replay Out of MVP per initial-rec §9; no evidence yet the simple oracle-generalization pattern is insufficient; datafold/data-diff is confirmed dead (2024-05-17 sunset)
Commercial shadow platforms (Signadot/Speedscale) K8s/mesh-native (estate is pinned-EOL Docker booted ad hoc); Enterprise-gated; procurement won’t close before week 6; closed control plane in the audit chain; confidentiality cost of a vendor POC on a confidential program (build-vs-buy.md §2.3)
UI product beyond one Grafana panel The fleet reads structs, not screens

20. Post-MVP roadmap (staged per flagship gate)

Each stage is gated on the prior flagship clearing, never a parallel build-out (north-star.md §9).

  • Stage 1 — Orders adapter hardening (M2→M3): full Orders read-path fixture coverage (MODELED 50-60+ classes); MBUS shadow-tap across both bus technologies (JMS/STOMP + Kafka) + outbox durability; the two missing oracles (provider-settlement, contract-check) for the money-gate wrapper; DONE-row-8 rollup (US-14).
  • Stage 2 — Second adapter = the generalization test: extract the real kernel API here. This is where the 2×-cost detection fires — if adapter #2 costs >2× adapter #1, freeze generalization and fall back to per-domain bespoke + S3 (§21).
  • Stage 3 — advisory LLM triage (cluster + explain DIVERGEDs, suggest candidate rules for human approval — never a verdict); A/A self-replay noise-floor automation; batch/cron capture mechanism (DB-txn-log tailing) once the Resque/Quartz job inventory is measured.
  • Stage 4 — next hub by in-degree (voucher-inventory 8 → deal-catalog-jtier 7 → users-service 6 → billing-record 5) → long tail, only after Orders clears M4.

21. Release strategy

  1. Pilot service (forex-ng self-test, then giftcard_service/orders-ext spikes) — retire lab + noise risk off the critical path. Days not weeks per easy repo (DESIGN).
  2. Orders read paths — M2 reads-authoritative; recorded-replay functional parity on the read surface; first real PARITY (activation) before week 6, or the “M4 slips, publicly” trigger fires.
  3. Orders write-ramp support — M3 1→100%. The harness supplies functional-parity verdicts + the evidence wrapper; the ≥99.999% ledger number is Zaruba’s live reconciliation (§13), reported as one combined DONE-row-8 rollup with its confidence basis (sample count, window, source) stated — write-ramp dwell time sizes to a confidence interval (Uber’s model: ~1 day at ~1,000 comparisons/sec buys six-nines), not a flat N-days rule (US-14).

22. Success metrics

From north-star.md. North Star: % of signed HOT/WARM Orders-chain capabilities with a current parity-coverage verdict (PARITY or DIVERGED-with-open-question; UNVERIFIABLE does not count). Denominator = the panel’s signed capability list (independent of the harness team → un-gameable, unlike raw verdict counts). Reported per flagship.

Metric Definition Target/note
Activation First PARITY verdict end-to-end on the pilot Binary, one-time; gates “the kernel exists”
Time-to-value Median hours, unit-claimed → first verdict Tracked so verify never silently becomes the schedule
Quality — false-DIVERGED rate Self-replay sampling (any non-PARITY on self-replay = false DIVERGED) Near-zero; nonzero triggers a normalization-ledger review. Only precision evidence available (D6 Q2)
Trust Bypass attempts (should be structurally zero) + normalization-rule audit-acceptance rate Any nonzero bypass = process breach, not a metric to optimize
Cost Verify cost/unit (tokens + lab minutes); lab minutes/fixture Tracked against the MODELED 2-5 eng-day / <1-day follow-on model to test it
Reliability % fixtures reproducing the same verdict on repeat replay ~100%; drift = kernel bug

Guardrails: fixture staleness (a stale fixture drops out of the numerator even if its last verdict was PARITY); normalization-rule growth outpacing audit-acceptance (silent widening of what’s ignored); coverage-without-depth (a “covered” capability must include its Given/When/Then use-case scope, not one happy-path fixture).


23. Pivot criteria

Concrete, watchable triggers (build-vs-buy.md §5):

  • EOL lab blows its investment cap on the pilot ladder (manifest-v1 pull unrecoverable, TLS/native-gem intractable) → drop the recorded-replay lab for that domain, flip to S3 impact-selected tests + oracle-only where an independent oracle exists. The harness shrinks to wrapper + money-gate partnership; the lab-build is abandoned, not the harness.
  • Internal MBUS-diffing precedent surfaces in repos/monorepo-development/ or Zaruba’s oracle code → don’t build shadow-taps, wrap the existing mechanism (removes the highest-uncertainty line).
  • Tomas assigns the whole verification lane to Zaruba’s plane → collapse toward Option 5: harness = verdict/evidence wrapper + adapters only, no independent lab.
  • Recorded replay shown sufficient for the money gate (no evidence today) → the harness could own the live gate; the partnership narrows to a read-only feed integration.
  • A single tool clears >60% of the kernel including MBUS (nothing does today) → narrow BUILD to adapters + evidence wrapper.

24. Kill criteria

  • Flagship gate fails on the harness’s account. If no Orders unit reaches PARITY before week 6 because the harness cannot produce a verdict (not because of a real regression), the harness has failed its one job — “if dual-run is not live, M4 slips, publicly” (context brief). Fall back to S3 impact-selected tests for the flagship and re-scope the harness to the money-gate wrapper only.
  • “Harness became the schedule.” If harness build/operation is itself delaying the flagship past its capped per-flagship investment budget — detection: adapter #2 costs >2× adapter #1, or the pilot ladder exceeds its capped time-box by a wide margin — freeze generalization, ship what exists as the Orders bespoke proof, and hold the line at S3 (impact-selected tests, TDAD 6.08%→1.82%, ~70% — single uncorroborated off-domain study, D6 Q7, cited as fallback economics only, not a proven number). The rip-cord requires the cap to be a budget line with a named S3 trigger — currently DESIGN, not a number (§25).
  • Self-replay cannot reach zero false DIVERGED on a real corpus after reasonable normalization effort → the noise model is unsound; the harness cannot be trusted as a gate and must not ship as one. Fall back to S3.

25. Open questions

  1. Money-gate lane, unratified. Option B assigns the live ≥99.999% reconciliation to Zaruba’s plane, but his decision log has zero dual-run/shadow/parity entries — proposed, not agreed. Needs an explicit Robert↔Tomas lane decision + a shared verdict contract before M3. (architecture.md Q1, research.md §5 Q1.)
  2. Internal-endpoint seam. The billing-event feed and reconciliation reads are expose:false, reachable only via Core’s in-cluster client. An external CLI/CI runner can’t call them → a Core-side authenticated read adapter, or that hop runs inside Encore — either partially contradicts “kernel + lab outside Encore.” Affects work-ledger-client + Zaruba-oracle-read.
  3. EOL bootability at full combined difficulty. Is jruby:1.7 / openjdk:7 actually pullable on a current Docker Engine (manifest-v1)? Does apache/activemq-classic serve 2013-17-era JMS/STOMP clients? Both plausible-but-unverified; the pilot ladder is designed to answer them cheaply before the flagship depends on them.
  4. Self-replay false-DIVERGED rate. Nobody published theirs (D6 Q2); the self-replay gate is the only precision evidence. Must be proven zero-false-DIVERGED on a real corpus before the fleet trusts it, or bypass follows.
  5. The investment cap, numerically. “Weeks not months” and “cap per flagship” are DESIGN, not a number. The kill-criterion rip-cord has no brake until the cap is a budget line with a named S3 trigger.
  6. Retention windows per PII/PCI class — unknown, needs security sign-off before any fixture is frozen.
  7. The legacy estate’s own structurally-expected noise floor (analogous to Uber’s ~10 corruptions/trillion from S3’s 11-nines durability) — needed so the ≥99.999% gate is set to a sourced floor, not arbitrarily. Needs input from Zaruba’s plane on the legacy MySQL/MBUS stack’s known failure modes.
  8. DARK batch surface for Orders — Resque/Quartz job classes uninventoried; the 50-60+ fixture-class estimate stays MODELED until an app/workers/*.rb + schedule-file grep lands. HTTP proxies cannot capture this class at all.
  9. assigneeGroupSlug registry + Tasks-workflow-stamp question-minting — neither exists today; a build decision not yet assigned an owner, and whether the harness owns a wrapper Temporal workflow is unresolved even inside Zaruba’s repo (D4 §a.4).
  10. QSA/SOX’s actual evidence-bundle requirement — §5 profile is assumption, not measured engagement. Confirm before the bundle format is finalized.