Southbridge
Engineering study · draft for review

Methods: source-bound entity resolution, costs and clocks

The headline unit is a donor dependency family, not a row or model question. Cases sharing assigned/context atoms, accepted donor identities, donor-side registry identities or required candidate-comparison dependencies remain together. Shared recipients do not establish donor identity. A family agrees with the reference only if every constituent case does; repeated donors cannot earn extra headline votes.

There are two frozen held-out panels. Panel 1 is the first comparison, published on September 17. Panel 2 is the replication drawn on September 18 after the Jev-first question criteria were revised using panel 1's retained answers. Panel 1 therefore also serves as the development set for that revision, and its post-hoc re-run is reported as a development number.

Frozen test denominatorPanel 1Panel 2
Dependency families200200
Cases204205
Unique assigned atoms / case incidences275 / 284279 / 293
Unique original rows / case incidences829 / 844903 / 951
Unique assigned pairs113104
Reference relationships14 household, 1 filer12 household
Single-atom cases146141

An atom groups source rows before resolution; a case assigns atoms and evidence for a decision. These are workload measures, not independent accuracy samples. Neither panel's reference contains an unknown pair or an uncertain atom, so any model abstention scores as disagreement.

Where the reference comes from

The reference is the accepted D28 Ohio resolution set. Each accepted case was proposed by a swarm of cheaper models and accepted when they agreed or when an adjudicator ruled. For panel 1's 204 cases, reference/source-acceptance.jsonl records gpt-5.6-luna as a proposer on every case, deepseek-v4-flash as adjudicator on 193, Luna itself on 9 and deepseek-v4-flash-vision-exp on 2; 200 carry the acceptance label swarm. Panel 2's 205 cases: Luna proposed on every one, deepseek-v4-flash adjudicated 200 and Luna 5, all labelled swarm.

Three consequences follow. The cheap arm's reasoning model is the reference's main proposer, so its agreement is partly self-agreement. The Jev-first arm's Luna and Gemini fallbacks share that dependence. Fable is the only arm whose answers are independent of how the reference was built, and its three panel-1 disagreements were two unknown verdicts on suffix conflicts (DANIEL R. OKAFOR against DANIEL OKAFOR JR; ALVAREZ^ SR. against ALVAREZ) and one committee label for a county party judicial fund the reference kept as organization. The comparison measures agreement with this reference, not human truth, and the scoring never rewards abstention.

What entered the panels

The census covers the seven active D28 Ohio releases recorded in reference/release-index.json (calibration, checked reuse, ordinary-bank, risk-cleared and three nonregistry-risk releases). The external index later drifted to thirteen releases; panel 2 pins the retained seven-release copy so both panels come from one population. Superseded releases and D27 are ineligible. Selection uses seeded SHA256 ranking within semantic and row-count strata, six round-robin opportunities per stratum, then seeded fill — never ascending case size. It is a purposive accepted-reference sample, not a prevalence estimate.

Panel 1 excluded 636 families for prior exposure across earlier study inputs, challenges, label files and replays. Panel 2 adds every panel-1 assigned atom, context atom, family and source case to that exclusion (836 exposed families) and uses a new seed; the two panels share no case, family, atom or source case. The extractor asserts all of this.

Panel 1 includes person, organization, committee, joint and placeholder typing; matches, separations, suffix and name-form conflicts, registry mapping, donor/recipient namespace distinctions, household links and one filer link; four T2 cases. By the time panel 2 was drawn, the fresh population had no committee, joint, placeholder, registry-mapping, filer or joint-member families left. Its remaining fresh higher-tier family was excluded because the retained original capture was truncated, as recorded in the frozen manifest; it was not a measured high-tier case. Panel 2 is all T0 and tests typing (285 person, 8 organization atom incidences), identity (99 same, 19 different pair incidences) and household links (12). Neither panel has a fresh explicit-uncertainty family.

Extraction checks assigned-row counts, recovers enumerated context and candidate endpoints, retains all case-associated references, and verifies capture and body hashes. Models receive complete retained text or candidate records. Accepted resolutions, review narratives and workflow wrappers are excluded from model inputs; native workers have no reference-file mount.

Earlier three-arm inference paths

All arms receive the same original evidence and semantic output opportunities, with concise JSON and legitimate uncertainty allowed. The no-Jev arm is a competent system control, not the same atomic-question pipeline with one API swapped.

ComponentRoute and settings
Lean Fable baselineAnthropic Messages API, claude-fable-5-1; adaptive thinking, medium effort; cacheable static task/schema prefix; no tools or agent runtime
Cheap primary/referred assertionsHankweave/Pi pi/openai/gpt-5.6-luna; high thinking
Cheap-branch adjudicationHankweave/Pi pi/google/gemini-3.7-flash; medium thinking
Typed judgmentsTypeSafe /v1/systemone, jev-1.13.0; Choice questions with offered labels and probabilities

Fable runs at medium effort while Luna runs at high thinking; that asymmetry favours the cheaper model on quality and Fable on cost, and it is the same in both panels. Native codons use the pinned Hankweave 0.10.0 no-retry image, fresh sessions, no compaction, a 300-second deadline and a 16,384-output-token ceiling. Fable shares that deadline and ceiling. Native tools remain installed but are prohibited by the prompt; acceptance rejects tool use, retries and prompt-delivery mismatches.

The cheap arm asks Luna for complete resolutions, then sends missing, incomplete or uncertain cases to Gemini. Jev-first judges typing, unit coherence and identity before registry and relationship assertions. Chosen-label probability thresholds are 0.80, 0.75, 0.70, 0.90 and 0.90 — not the separate confidence statistic — and are identical in policy-v3.json (panel 1) and policy-v4.json (panel 2). Unknown labels, uncertain units and below-threshold assertions are batched for Luna; unresolved cases go to Gemini. Views above 64,000 bytes bypass Jev intact. Packing targets 32 questions and 80,000 pre-encoding bytes per Jev request.

The rubric revision

claims.ts retains v1, the original question criteria, and v2, the selected identity/household revision. The later experimental v3 is retained for research history but is not the selected policy. v2 changes only the identity and same_household questions: it states the evidence that suffices under the reference standard, not only the evidence that fails.

Questionv1 told the modelv2 adds
Identity, same"supports the same donor identity, including corroborated clerical or name-form variants"Same surname and the same or clerically equivalent given name at the same residential street address, or same postal code with the same employer, is sufficient; a middle initial or suffix present on only one record is a filing variant
Identity, different"a genuine generation/name conflict can distinguish people, but missing corroboration or different addresses/jobs alone cannot"Different given names, or explicitly conflicting suffixes on both records (JR against SR, II against III), are sufficient
Household, yes"supported between distinct compatible entities"; "a shared employer/mail destination is not by itself a household"Same house number and street on both records is the accepted evidence; abbreviation and punctuation variants do not break it; employer, business or PO-box addresses do not count

For that earlier gated comparison, every other question, threshold, byte bound, batch policy and routing rule was unchanged. The revision was frozen in policy-v4.json after the panel-1 probes and before its panel-2 outcomes.

The probes (runs/preflight/rubric-probe-v1.json, rubric-probe-v2.json) re-asked panel 1's 87 identity questions and the 13 household questions that fit the Jev byte bound, using a provisional partition composed by the pipeline's own composer. Under v1 the probe reproduced the held-out failure: 56 definite identity answers and 31 unknown; 13 of 13 household answers no at p 0.75–0.99. Under v2 on the same requests: 87 definite identity answers, 86 agreeing, none unknown; 13 of 13 household answers yes, twelve at or above 0.90. Two earlier probe attempts are retained as rubric-probe-v1-contaminated-*.json: they passed the reference resolution itself as the upstream proposal, and Jev answered yes 13/13 under the unchanged v1 wording because the proposal's relations, and then its basis narrative, contained the answer. That is a state-design observation, not a rubric result.

What the gate accepted, and what it should have

audit-decisions.ts re-reads every retained Jev decision against the reference. For panel 1 under v1 (runs/analysis/heldout-completion-1-jev-first-decision-audit.json):

Question kindAskedGate-acceptedAccepted and agreeingDefinite answers agreeing
type268250250268 / 268
unit coherence268239239265 / 265
identity87131353 / 54
registry201 / 2
relation (household)10900 / 10

Of the 132 panel-1 cases that needed no reasoning model, 123 were single-atom cases. Counterfactual threshold sweeps on the same recorded answers show identity at 0.55 would have accepted 33 with 33 agreeing, and type or unit at 0.50 would have accepted every definite answer with none disagreeing; thresholds were nonetheless left unchanged so the revision isolates the wording. The same audit for the revised panel-1 re-run and for panel 2 lives beside it.

Development and accounting

The original architecture and thresholds used 26 previously exposed, labelled development families; its held-out outcomes did not tune selection or routing. The question revision then used panel 1's retained answers and probes. The later Jev-only research again used panel 1 for development and froze its candidate and review ranking before inspecting panel-2 outcomes in that research pass. It reused the earlier control runs.

Both panels use 55 and 52 input batches respectively, rotated arm order and two batch workers. Component timing includes awaited dispatch, native startup where applicable, parsing and accounting; it is neither provider-only latency nor per-arm makespan. The driver stops new dispatch after any unpriced component. Panel 1 paused once on a Luna connection error. Panel 2's dispatch paused on one Jev request of two questions and 89,794 wire bytes that returned HTTP 400 max_tokens_exceeded (both questions fell back to Luna) and on Luna connection errors that Gemini adjudication absorbed; each pause was inspected and resumed with the identical plan, completed batches retained, commitments held. The resume record is runs/editorial-replication-v4/resume-log.txt.

Funding has three original authorizations — $25, $10 and $50 — plus a later $10 research allocation drawn from the extension's unused capacity, not an additional authorization. All four phases are sealed and the total authorization remains $85. A comparison's model estimate includes its priced failures and fallbacks but excludes development, reference construction, local compute and authoring. The whole-study receipt separately includes development and failed/superseded study calls. Catalogue estimates are not invoices; unknown charges retain their commitments.

Jev alone, followed by a budgeted family review

The selected research policy is policy-jev-only-v2.json: the corrected wording, complete literal source records, up to four cases per batch, and no automatic reasoning fallback. Valid selected choices are consumed without applying the legacy probability gates. The accepted_by_gate field is diagnostic in this arm. Explicit unknowns, contradictions and source-capacity exclusions remain visible and stay in the denominator.

After the first frozen Jev pass, whole families are ranked by invalid or missing output, then explicit uncertainty or incomplete assigned work, then the minimum selected-option probability across their recorded decisions. Stable family hashes break ties. This is a heuristic ordering, not a calibrated probability of family error. The ranker receives membership metadata but no reference resolutions or correctness flags.

The predeclared review prefixes contain 0, 1, 2, 5 or 10 families. Each selected family is a fresh, source-only Luna request under Hankweave, with the original evidence and resolution contract but no Jev proposals, probabilities or reference answers. It cannot fetch more evidence. Structurally valid case answers replace the corresponding Jev answers before scoring; a failed review retains the base answer or its unresolved slot. More review can therefore reduce agreement.

One actual ten-family invocation supplied the calls used by the prefix comparisons. Smaller prefixes reuse complete calls, never a fraction of a larger call's cost. They are derived policy evaluations, not separately timed deployments. One malformed case identifier in the ten-family run received one explicit same-source/same-prompt recovery attempt. Both the failed original call and the recovery are included in that final point; no reference outcome selected the retry.

Four timing quantities, not one speed number

ClockWhat was measuredWhat it must not be called
Whole Jev invocationStart to finish of one single-arm run, with two batch workersProvider-only latency
Accumulated component serviceSum of full awaited call durations, including unsuccessful included attempts and wrapper overheadElapsed end-to-end runtime
Review-stage clockActual elapsed time of the full ten-family review invocation; the recovery has its own clockThe runtime of the smaller prefixes
Median batch serviceMedian of the summed component-call times within each original input batchA whole-panel wall clock

The first corrected Jev invocation is the base for every review prefix. The other two corrected repetitions show observed cost, outcome and elapsed-time variation; they are not an ensemble. The full review and explicit recovery stage clocks may be added to that base clock, but the result is labelled a sum of separately recorded stage clocks, not one observed combined deployment. For smaller prefixes, elapsed time is left unmeasured. Dividing summed service by worker count does not turn it into a measured wall clock.

The retained Fable, cheap and earlier Jev-first controls ran interleaved in a different execution wave. They have component-service sums and median batch-service values, not independently observed per-arm makespans. Timing comparisons are not a provider-capacity test, a latency SLA or evidence of petabyte throughput.

Cost, byte and rate derivation

economics-data.json is generated offline by derive-economics.ts. Its historical comparison remains bound to the report selected by results.json; its required research block is separately bound to the accepted Jev-only result, pinned report-input manifest and verification receipt. New costs and clocks are reconciled against original component and invocation records. Only aggregate counts, descriptive run labels and hashes are publication material; raw source records, benchmark identifiers, native prompts and private filesystem paths remain outside this JSON.

For the historical comparison, the common byte denominator is the UTF-8 length of stateText(sourceState(batch)), counted once per original input batch and compared byte-for-byte with its retained Fable-primary request. The Ohio census is retained separately. This is a logical evidence-workload measure, not raw CSV volume, network payload or a token count; kB means 1,000 bytes.

evidence kB / service second = serialized evidence kB / sum(component seconds)
priced dollars / evidence kB = priced model subtotal / serialized evidence kB

Each model bill reconciles its component costs, call counts and times. The new review points show Jev, completed original Luna calls, failed original reviews and the explicit recovery separately. Zero-cost parts mean no recorded charge in that category; an unpriced attempt remains unknown and is never turned into zero. Displayed values are rounded; the JSON retains the full estimates. Gate acceptance is not correctness.

Full Ohio scope and scale scenarios

The held contribution substrate is four raw CSVs: CAN_CON_2022, CAN_CON_2024, PAC_CON_2022 and PAC_CON_2024, with 883,012 data rows and 189,908,438 raw bytes for report years 2022 and 2024. The corpus census re-hashed the raw files and counted their record terminators against the retained quote-aware transform receipts; the independent normalization receipt compares all 883,012 rows and 283,349 strict evidence bundles with zero mismatches. A strict bundle groups equivalent source evidence; it is not a resolved donor identity.

The scale explorer applies linear normalizations to the measured panel-2 workload:

row-equivalent factor     = target original rows / panel unique rows
bundle-equivalent factor  = target strict bundles / panel unique atoms
evidence-byte factor      = target serialized evidence bytes / panel evidence bytes
effective factor          = chosen factor × remaining-work share
priced model scenario     = measured priced subtotal × effective factor
service-effort scenario   = measured component-second sum × effective factor

These are arithmetic on the observed workload, not measured statewide runs or quotes. They hold the panel's case mix, batching, referral frequency and error behaviour constant; panel 2 is more ordinary than the full state, which contains committee, joint and higher-tier families that panel 2 does not. The token-budget calculator applies the retained Jev input tariff of $0.042 per million tokens: a decimal petabyte submitted whole implies $10.5 million; one million 2,000-token questions imply $84. No experiment here establishes that such a reduction preserves a petabyte task.

Reading the interactive report

The opening breakdown describes the later Jev-only and family-review research. Its cost chart, repeated-run table and clock disclosures remain readable without JavaScript. The historical performance controls, model receipt, before/after table, family explorer and scale calculator retain the earlier comparisons; their labels identify that scope. The anonymous explorer still uses the earlier three arms and does not silently substitute sparse-review outputs. Both data sets remain separately source-bound within economics-data.json.

To use a downloaded publication bundle, run this inside the extracted publication/ folder, then open http://127.0.0.1:8000/:

python3 -m http.server 8000 --bind 127.0.0.1

No model requests are made by the report. Opening the HTML directly, disabling JavaScript or losing a script leaves the static results readable.

Local reproduction

From the full study checkout, inspect either frozen plan without inference:

jq '{label, split, arms, policy, panel, concurrency}' article-study/runs/experiments/heldout-2/plan.json
jq '{label, split, arms, policy, concurrency}' article-study/runs/experiments/heldout-completion-1/plan.json

Score retained outputs and audit the decisions:

ARTICLE_BUDGET_CONFIG=article-study/budget/extension-1/config.json bun article-study/report.ts --label heldout-2
bun article-study/audit-decisions.ts --label heldout-2
bun article-study/audit-decisions.ts --label heldout-completion-1

The study requires its private inputs, pinned native image and source dependencies. All four funding phases are sealed against further paid inference. The research's report-input manifest, confirmation freeze and exact-input verification are retained under runs/jev-only-study-1/ in the private checkout; the public ledger exposes their hashes, not their contents. Private benchmark inputs are not distributed with the article. The standalone example is portable: bun example.ts --dry-run needs neither credentials nor benchmark files.

See TypeSafe Choice, state, probabilities and confidence, and the version-matched Hankweave models and harnesses reference.