Southbridge
Engineering study · draft for review

We rewrote one question. Jev went from missing every household link to finding all of them.

An Ohio entity-resolution study: the cost of typed decisions, the cost of reconsidering them, and the clocks that distinguish a fast model from a fast workflow.

Jev alone handled the first pass for a few cents. A small, signal-selected review queue bought more agreement; a larger queue did not eliminate the remaining disagreement. The latest breakdown puts the bill beside output coverage and separates a measured elapsed clock from the sum of time spent in individual calls.

This experiment uses the same second panel as the earlier comparison below. The Jev-only candidate and family-review ranking were frozen before this research pass inspected its outcomes; the Fable and cheap-reasoning controls are retained runs, not new calls. Every agreement figure is against the model-built reference, not human-established truth.

The latest cost and speed breakdown

Jev alone, then a little review

First corrected Jev run on confirmation panel · 200 families · 205 cases · 52 batches · 279 strict bundles · 903 original rows.

3.37¢Jev-only model estimate · 195/200 full family matches
28.64 swhole Jev invocation · two batch workers · 203/205 valid cases
4.35¢Jev plus five reviews · 199/200 full matches · elapsed not separately measured

The predeclared whole-family queue ranks invalid outputs first, then explicit uncertainty, then minimum selected-choice probability. It is not a fixed confidence cutoff. Source-only Luna sees the complete selected family, not Jev proposals or reference answers.

The sparse-review bill, on its own scale

Each candidate includes the entire first corrected Jev run. Bars show priced subtotals, not invoices; review costs are whole-family calls, never prorated.

Jev baseCompleted original reviewsFailed original reviewsExplicit recovery
Jev only$0.0337
Full agreement 195/200 · core 197/200 · valid 203/205 cases · 0 families / 0 cases assigned review
1-family review$0.0362
Full agreement 196/200 · core 198/200 · valid 204/205 cases · 1 family / 1 case assigned review
2-family review$0.0379
Full agreement 197/200 · core 199/200 · valid 205/205 cases · 2 families / 2 cases assigned review
5-family review$0.0435
Full agreement 199/200 · core 199/200 · valid 205/205 cases · 5 families / 5 cases assigned review
Ten-family review · first pass$0.0573
Full agreement 199/200 · core 199/200 · valid 205/205 cases · 10 families / 11 cases assigned review
Ten-family review · with recovery$0.0611
Full agreement 199/200 · core 199/200 · valid 205/205 cases · 10 families / 11 cases assigned review

Every bar starts at zero. Labels are rounded; the detailed accounting below retains the full estimates.

Full family agreement is agreement with a model-adjudicated reference, not human accuracy. Luna helped construct that reference: source-only review is not reference-independent. All original families and cases remain in the denominators.

Inspect every cost component and call count
Cost and call accounting · USD priced subtotal, with attempts beneath
CandidateJev baseCompleted original reviewsFailed original reviewsExplicit recoveryTotal
Jev only$0.03367018266 calls · 0 unpriced$0.00000 calls · 0 unpriced$0.00000 calls · 0 unpriced$0.00000 calls · 0 unpriced$0.03367018266 calls · 0 unpriced
1-family review$0.03367018266 calls · 0 unpriced$0.002501151 call · 0 unpriced$0.00000 calls · 0 unpriced$0.00000 calls · 0 unpriced$0.03617133267 calls · 0 unpriced
2-family review$0.03367018266 calls · 0 unpriced$0.00425722 calls · 0 unpriced$0.00000 calls · 0 unpriced$0.00000 calls · 0 unpriced$0.03792738268 calls · 0 unpriced
5-family review$0.03367018266 calls · 0 unpriced$0.00979025 calls · 0 unpriced$0.00000 calls · 0 unpriced$0.00000 calls · 0 unpriced$0.04346038271 calls · 0 unpriced
Ten-family review · first pass$0.03367018266 calls · 0 unpriced$0.019721759 calls · 0 unpriced$0.00394171 call · 0 unpriced$0.00000 calls · 0 unpriced$0.05733363276 calls · 0 unpriced
Ten-family review · with recovery$0.03367018266 calls · 0 unpriced$0.019721759 calls · 0 unpriced$0.00394171 call · 0 unpriced$0.00371851 call · 0 unpriced$0.06105213277 calls · 0 unpriced

Rows are alternative policies over one shared Jev base, not additive study spending. Recorded model estimates, not provider invoices. Whole-family calls are never prorated. Failed original reviews and the explicit recovery remain charged separately. Research overhead and unpriced reservations are separate from deployment-shaped points; the research ceiling was an internal allocation, not new authorization.

Elapsed time is not accumulated call time

Measured elapsed is one whole invocation. Summed service adds awaited component-call durations, including overhead and unsuccessful included attempts; it is not elapsed time or provider-only latency. Summed stage clocks adds separately recorded invocations; it is not an observed end-to-end deployment.

Candidate timing · seconds
CandidateMeasured elapsedSummed serviceSummed stage clocks
Jev only28.64 s56.2 sNo recorded stage sum
1-family reviewNot separately measured69.31 sNo recorded stage sum
2-family reviewNot separately measured80.03 sNo recorded stage sum
5-family reviewNot separately measured114.89 sNo recorded stage sum
Ten-family review · first passNot separately measured177.18 s92.5 s
Ten-family review · with recoveryNot separately measured188.57 s103.89 s
Inspect the separate stage clocks and recovery
Separate stage invocations · seconds
CandidateJev baseTen-family reviewRecoveryArithmetic sum
Ten-family review · first pass28.64 s63.86 sNot included92.5 s
Ten-family review · with recovery28.64 s63.86 s11.39 s103.89 s

Ten-family review · first pass: 5 base misses repaired; 1 formerly matching family regressed. Combined end-to-end elapsed time was not separately measured. Stage sum adds the separately recorded Jev and ten-family review clocks.

Ten-family review · with recovery: 5 base misses repaired; 1 formerly matching family regressed. Combined end-to-end elapsed time was not separately measured. Stage sum adds recorded Jev, ten-family review and explicit recovery clocks; the original failed call remains included.

The corrected repetitions, not a latency promise

The original policy is shown separately from the three corrected repetitions. Answers and call counts varied, as did shared-machine conditions. These are individual observations, not a confidence interval or a provider SLA.

Whole-panel Jev runs · same 200 families / 205 cases
Run / policyCost / workMeasured elapsedSummed serviceBatch service median / p90Full / core agreementValid cases
Original Jev wordingOriginal wording$0.03072098461 calls · 709 questions · 731,452 input tokens · 0 unpriced25.86 s50.64 s0.73 / 1.41 s166/200 / 171/200202/205
Corrected Jev · run 1Corrected wording$0.03367018266 calls · 716 questions · 801,671 input tokens · 0 unpriced28.64 s56.2 s0.81 / 1.7 s195/200 / 197/200203/205
Corrected Jev · run 2Corrected wording$0.03464340667 calls · 717 questions · 824,843 input tokens · 0 unpriced33.14 s65.27 s0.93 / 1.85 s194/200 / 196/200203/205
Corrected Jev · run 3Corrected wording$0.03367018266 calls · 716 questions · 801,671 input tokens · 0 unpriced31.14 s61.25 s0.88 / 1.69 s195/200 / 196/200203/205

Batch medians and p90s describe summed call time within a batch. They must not be compared to another program’s whole-panel elapsed clock.

Compare programs on the same panel

These columns use the same timing definitions. The retained controls ran interleaved in an earlier execution wave; no independent whole-program elapsed clock is claimed for them. Their unpriced attempts remain outside the priced subtotal.

Same panel · different execution waves
ProgramFull family agreementPriced model estimateSummed serviceMedian batch service
Corrected Jev only · first run195/200$0.033670182All included attempts priced56.2 s0.81 s
Fable · direct200/200$9.831787All included attempts priced844.64 s14.61 s
Cheap reasoning · no Jev199/200$0.301246558 unpriced attempts also retained715.08 s13.27 s
Earlier gated Jev + reasoning198/200$0.1123732284 unpriced attempts also retained290.22 s6.36 s
The research bill and the Jev work behind it

The research phase priced $0.564449114 of model work, with $0.426862 held for 4 unpriced attempts. The held amount is a reservation, not a measured charge; it does not turn the priced subtotal into an exact bill.

$10.00 was allocated internally, not newly authorized. Authorization increased by $0.000000; the protected reserve is $0.5000. Candidate costs above reuse one run and cannot be summed to recover the study bill, which also includes other recorded research work.

The shared corrected base used 66 Jev calls for 716 typed questions and 801,671 input tokens.

Typed questions in the shared Jev base
Question kindQuestions
type293
unit293
identity118
relation12
Total716
Source receipts and content hashes

Public provenance uses descriptive labels and SHA-256 hashes, not private source paths or family identifiers.

Accepted research result
3362808ebd76f34898051a88e6f552ae6d820c8f6b9ec699cdd0b4c99ccad29a
Pinned input manifest
88d29ac86552b28c44d3e7e9c2445ff0cbd00fb380afcc13700b3857cf7774c6
Acceptance verification
c0116c2a1e51d0160952267ba85b9ab4429f1f3f185307073363e565deafef0a
Accepted Jev research results
3362808ebd76f34898051a88e6f552ae6d820c8f6b9ec699cdd0b4c99ccad29a
Research input manifest
88d29ac86552b28c44d3e7e9c2445ff0cbd00fb380afcc13700b3857cf7774c6
Accepted input verification
c0116c2a1e51d0160952267ba85b9ab4429f1f3f185307073363e565deafef0a
Predeclared confirmation policy and prefixes
ca151ca289862e5bf123b78e2e00a9b9aac8cd7e436111ab955ecdea26c8ef5b
Original Jev wording · report
8ece99387b3cda97c2cd1d6978e0784c172fd52b1f688e9734fba01655bd0043
Corrected Jev · run 1 · report
62b5c65c011142a440b57070ccee81848f1cdf7b2ce7c6e08b559ff31ed92732
Corrected Jev · run 2 · report
1eac586e82844ab619a27f04bea30ce9797dfa29c4b0a8371a52bf267b43fd04
Corrected Jev · run 3 · report
52723534bb0c1557f28c074c72567d44d6d03be6fc57f454a4eb1e304bd074aa
Sparse confirmation prefix report
79d748a49bfb4317c1806459af1f44f7ef05bd816180f47c03a1915a27b1c539
Original ten-family review invocation
50fe704d93d97ebad0c59d3f66e1a9bffb68b35bac505c31ab9cfa9b465ceb1f
Predeclared same-input recovery
00cb2837c8057c94ce03693a2b6ccbaf52114f78aec3609c07033b5b6922165a
Recorded recovery result
dbc38cedf0ec51f4c189d943e819c09fffedc57ae63a66cc729d57e87e3500a1
Sealed study funding summary
f7e94c6b61991f90dcf9b93354a2825e25c0aea31c8711de894c5a9f6419f19a

The earlier experiments explain how we got here. On the first panel, the gated pipeline missed nine of fourteen household links and cost more than the cheap control. We published that result. The next comparison changed the wording of the identity and household questions, not the model or the thresholds. The later Jev-only experiment then asked how much reasoning we could remove. The sections below keep those designs separate.

Inspect the latest bill and clocks, read the case history, compare the earlier gated programs, or follow seven rows through the join.

The rows are not the donors

Ohio's campaign-finance files record reported contributions: a name, an address, an amount, a date, a receiving committee. Our substrate is the candidate-committee and PAC filings for report years 2022 and 2024 — 883,012 rows. To total contributions by donor you must decide which rows belong to the same person, and the strings will not tell you. The names in this article are altered; every field pattern is a real one from the panels. DANIEL R. OKAFOR and DANIEL OKAFOR JR at one address are one man with an inconsistent clerk, or a father and son at the same real-estate firm. WALTER H. BRENNAN SR. and WANDA BRENNAN at one address are two people and one household. A county party's "judicial fund" is an organization until a registry says it is a committee.

That is entity resolution. A false merge adds one donor's money to another's total; a false split scatters one donor across several. Splink and Dedupe are the established probabilistic approaches, and Dedupe's own tutorial uses campaign contributions. We did not benchmark them. Our question was narrower: when the deterministic work is done and a judgment call remains, how much of it can a cheap typed model make, and how do you find out when it is wrong?

The earlier three-arm experiment

We first compared these three programs on the same original evidence, with the same semantic output opportunities and legitimate uncertainty allowed. This is the gated architecture that preceded the budgeted family-review experiment above.

ProgramWhat runs
Fable, directlyThe whole case in one lean Messages API call: concise JSON, adaptive thinking at medium effort, a cacheable prefix, no tools, no container.
Cheap reasoning, no JevHankweave runs Luna on the whole case; Gemini handles anything unresolved. This control separates "cheaper reasoning model" from "Jev."
Jev + selective reasoningJev answers typed questions — donor type, unit coherence, identity per pair, then registry and links — with a probability per option. Code composes the case from the answers that pass a gate and sends the rest to a Luna hank; a Gemini hank adjudicates what still will not compose.

You should know how the reference was made. It is the accepted D28 resolution set: a swarm of cheaper models proposed each case and accepted it on agreement, with Luna proposing on every case and DeepSeek adjudicating most. So the cheap control is graded partly against its own handwriting, and Fable is the only arm that is not. The reference never abstains, so an "unknown" always scores as a miss. Every agreement number here is agreement with that reference, not with a human.

The first panel

Panel 1 was 204 cases in 200 families, frozen before any model ran. Here is what we published on September 17, and what happened when we re-read the retained answers.

ProgramFullCorePriced estimateMedian batchNo reasoning modelLinks found
Fable 5.1 · direct197/200197/200$17.3616.10 s0/20415/15
Cheap · no Jev197/200198/200$0.310412.05 s0/20414/15
Jev-first · v1 wording190/200200/200$0.3639 + 1 unpriced7.07 s132/2045/15
Jev-first · v2 wording, post hoc200/200200/200$0.23896.54 s158/20415/15

The three original rows are the frozen first comparison. The revised row re-ran only the Jev-first arm after the question criteria were rewritten using that panel's own retained answers, so it is a development result, not a held-out score. Panel 2 is the held-out test of the revised design.

The headline had been "200/200 on the core task." The audit said something less flattering. Of the 132 cases Jev handled alone, 123 had one donor in them; the fast path was mostly labelling a single row person. On identity, Jev gave a definite answer for 54 of 87 pairs and 53 agreed with the reference — but the 0.70 gate admitted 13, and Luna did the rest. On households, Jev was asked ten times and answered no ten times, nine at 0.90 or above; five of those nine pairs had byte-identical street addresses and the same surname. Every household link the pipeline did find came from Luna or Gemini. And the three arms were within noise on core (197, 198 and 200 of 200): Fable's three misses were unknown on two suffix conflicts and committee for the judicial fund, calls a careful analyst might make the same way.

None of this was hidden by the numbers. It was hidden by reading them as a scoreboard instead of a receipt.

The fix: say what suffices

We had written the household criteria the way you write a warning. A shared employer or mailing address is not by itself a household. True, and useless: it told Jev what fails and never what passes. The reference's own reasoning was explicit — a shared residential street address plus a surname is household evidence — and we had never put that sentence in front of the model. The identity criteria had the same defect. Of Jev's 33 unknown identity answers, 27 were pairs with the same first, middle and last name at the same street address or postal code — some differing in employer, some in nothing but punctuation. The reference merges every one of them. Jev sat at fifty-fifty because nobody had told it that was enough.

So we rewrote two questions. Not the thresholds. Not the models. The words.

Questionv1 saidv2 adds
Household, yes"supported between distinct compatible entities"Same house number and street on both records is the accepted evidence; DR against DRIVE or CT against COURT does not break it; employer, business and PO-box addresses do not count
Identity, same"supports the same donor identity, including corroborated clerical or name-form variants"Same surname and given name at the same street address is sufficient; a middle initial or suffix on only one record is a filing variant, not a conflict
Identity, different"a genuine generation/name conflict can distinguish people"Different given names, or JR against SR on both records, are sufficient; different addresses or employers alone are not

Then we asked Jev the same 87 identity and 13 household questions again, on the same evidence, and compared.

Panel 1 probev1 wordingv2 wording
Identity: definite answers / agreeing56 / 5687 / 86
Identity: unknown310
Household: answered yes0 of 1313 of 13
Household: at or above the untouched 0.90 gate012

The one identity disagreement under v2 was KAREN L. against a first-name field holding KAREN L, at p = 0.60, below the gate. It went to Luna, which is where it belonged.

Two false starts are worth your time. Our first probe passed the reference resolution into the model's state as the "upstream proposal," and Jev said yes 13 of 13 times — under the old wording. We had leaked the answer: first the relations, and when we stripped those, the reference's prose explanation. Only when the proposal was composed by the pipeline's own code, with the generic basis the real run uses, did v1 reproduce the failure exactly. The lesson is not about our probe. Jev reads everything in state as evidence, including narrative you labelled untrusted. Put a reviewer's summary in the state and you are asking the model whether the reviewer sounded sure.

With the words fixed and nothing else touched, we re-ran the Jev-first arm on panel 1. It matched 200 of 200 families on the full contract at $0.24 with nothing unpriced, and 158 of 204 cases never saw a reasoning model. That number is real, but it is a development result — the panel's own answers chose the wording. For the gated design, the held-out test was panel 2.

The second panel: the earlier gated comparison

Panel 2 is 205 cases in 200 families drawn from the same frozen population with a new seed, after excluding every atom, family and source case in panel 1. The Jev-first design was frozen with the v2 wording before any outcome. Fable and the cheap control ran unchanged.

2.09×median-batch service rate versus cheap reasoning
80.49%of cases had no reasoning-model assignment
87.49×smaller priced subtotal than Fable; whole-pipeline comparison

The comparison’s model bill

Priced components for the entire 205-case comparison; not an invoice.

Fable · direct$9.83
Cheap reasoning · no Jev$0.3012
Jev + selective reasoning$0.1124

Lower is better · every bar starts at zero

Median batch service time

Sum of awaited component-call durations per input batch; not per-arm makespan.

Fable · direct14.61 s
Cheap reasoning · no Jev13.27 s
Jev + selective reasoning6.36 s

Lower is better · every bar starts at zero

2.3× the median-batch service rate of Fable, at a 87.49× smaller priced subtotal. That cost gap belongs to the whole pipeline: the cheap no-Jev control also removes most of Fable’s bill.

Cheap reasoning · no Jev has 8 unpriced calls with $2.80 held in commitments; Jev + selective reasoning has 4 unpriced calls with $1.15 held in commitments. These were transport or request-limit failures whose fallbacks completed; a plotted subtotal is not that program’s exact total cost.

Same 200 familiesCore agreementFull agreement
Fable · direct200/200200/200
Cheap reasoning · no Jev200/200199/200
Jev + selective reasoning198/200198/200

Core = typing, identity and registry mapping. Full adds supported relationships and real-donor uncertainty. These are agreements with a model reference, not human-established accuracy.

What exactly counts as a byte or a second?

UTF-8 bytes of the canonical, escaped model-visible source state, counted once per original batch, including supplied references and repeated per-batch policy. Re-derived states match all 52 retained Fable-primary deliveries. This is a common logical-workload denominator, not raw CSV, distinct knowledge, total network traffic or disk throughput.

Sum of full awaited component-call durations, including priced/unpriced attempts and native startup/accounting where applicable. Not independent wall-clock makespan; the arms were interleaved with two batch workers.

The shared numerator is 1,204,763 bytes across 52 batches. For Jev-first: 1,204,763 ÷ 1,000 ÷ 290.22 = 4.15 kB per service second. $0.1124 ÷ 1,204.76 kB = $0.000093 per evidence kB, before the unpriced call.

Download the derived ledger

The measured default view is available without JavaScript.

Jev handled 165 cases alone — 122 with one donor and 43 with two or more — and its identity gate admitted 116 of 118 pairs, every one agreeing with the reference. It found eleven of the twelve household links: nine straight from Jev at the gate, two from Luna after one Jev request exceeded the model's token limit. The programs differed by at most two full-family matches. The bill and the recorded batch-service clocks describe a larger operational difference; they do not establish a provider-latency guarantee.

The two families Jev-first missed are the honest part of the story, because in both Jev had the reference's answer and the gate did not trust it. M LOUISE HARGREAVE and MARCUS J. HARGREAVE are two people at one address; Jev said different at 0.65, under the 0.70 gate, and the Luna fallback merged them — taking the household link with it. A county party organization drew organization from Jev at 0.62, under the 0.80 gate, and the fallback chain relabelled it. The safety net was the error both times. That is not an argument for lowering gates on the strength of two families. It is a reminder that the fallback model is not an oracle either, and that the retained answers are what let you see which one was wrong.

Two boundaries on this result. By the time we drew panel 2 the fresh population had run out of committee, joint, placeholder and registry-mapping families, so this panel tests typing, identity and households and nothing rarer; the twelve household links are its relationship test. And dispatch paused for twelve failed calls — eleven Luna connection errors, absorbed by Gemini adjudication or a later fallback, and the one oversized Jev request. Each pause was inspected and the identical frozen plan resumed with its completed batches; the failed calls kept their held commitments and are not counted as free.

Inspect all 200 anonymous family outcomes

Compare the measured programs

Full agreement is the default. Core checks typing, assigned-atom identity and registry mapping. Full also checks supported relationships and real-donor uncertainty. These are model-reference agreements, not human-accuracy estimates.

Full agreement · entire 200-family cohort. The graph is never filtered.

Fable 5.1 · direct

Full200/200 · 100%

Cheap reasoning · no Jev

Full199/200 · 99.5%

Jev-first · selective reasoning

Full198/200 · 99%
Whole-benchmark cost and timing
Fable 5.1 · direct
$9.8318 priced model estimate14.61 s median component-sum batch time
Cheap reasoning · no Jev
$0.3012 priced model estimate + 8 unpriced calls13.27 s median component-sum batch time
Jev-first · selective reasoning
$0.1124 priced model estimate + 4 unpriced calls6.36 s median component-sum batch time

Global figures never change with the family filter. Model estimates are not invoices; development, reference construction and local compute are excluded. A program with unpriced calls has a priced subtotal that is not its exact total cost. Timing sums component-call durations per paired batch, not independent program runtime.

Inspect the anonymous families

A family is a scoring unit of dependent cases, not a household. Numbers are display ordinals, never source identifiers. Jev-only and reasoning filters describe the Jev-first program’s routing.

200 of 200 families shown · Jev-first · Full agreement

Tile key: = match · disagreement.

123456789101112131415161718192021222324252627282930313233343536373839404142434445464748495051525354555657585960616263646566676869707172737475767778798081828384858687888990919293949596979899100101102103104105106107108109110111112113114115116117118119120121122123124125126127128129130131132133134135136137138139140141142143144145146147148149150151152153154155156157158159160161162163164165166167168169170171172173174175176177178179180181182183184185186187188189190191192193194195196197198199200

Family 1

1 case · 0/1 cases on the Jev-only path · 1 with a reasoning assignment

Fable 5.1 · direct

Core: Match · Full: Match

All case outputs schema/consumer valid.

Issue categories: None recorded.

Cheap reasoning · no Jev

Core: Match · Full: Match

All case outputs schema/consumer valid.

Issue categories: None recorded.

Jev-first · selective reasoning

Core: Match · Full: Match

All case outputs schema/consumer valid.

Issue categories: None recorded.

This is the static full-cohort view. To filter results or choose a family, serve this folder over HTTP with JavaScript enabled. No live model calls are made.

How the earlier gated program is built

Ordinary code keeps the original rows and their identifiers, parses the file-specific field meanings (PAC_REG_NO is the donor's registration in CAN_CON files and the recipient's in PAC_CON files), and checks that every input survives every transformation. None of that is model work.

Jev-first program and conditional reasoning hanks Complete source views feed typed Jev questions over HTTP. Chosen-label probability and unresolved labels route flagged questions to a targeted Luna reasoning hank under Hankweave and Pi. Oversized source views bypass Jev intact for full-case Luna reasoning. Deterministic composition checks structural consistency; incomplete, invalid or unresolved cases can go to a Gemini adjudication hank. Final deterministic checks preserve uncertainty and output IDs before the case-local source-row join. Structural validity is not semantic correctness or global truth admission. Identity/type questions and secondary link questions run in phases with provisional composition between them. Ordinary code orchestrates the fast path; Hankweave runs reasoning steps. This is not a measured petabyte deployment. Coverage limits: Panel 2 was drawn from the same frozen seven-release population as panel 1 after excluding every panel-1 atom, family and source case; the remaining fresh families contain no committee, joint, placeholder, registry-mapping or filer strata, so panel 2 tests typing, identity and household links only. No fresh explicit-uncertainty case exists in either panel; the reference never abstains, so any model abstention scores as disagreement. The reference was accepted by a swarm of cheaper models with Luna proposing on every case; the cheap arm and the Jev-first arm's reasoning fallbacks use that same model, and Fable is the only arm independent of reference construction. A purposive, decision-diverse sample from accepted references, not a population prevalence or human-accuracy estimate. Two held-out comparisons on 200 families each; no petabyte workload was executed. Jev first. Reason where needed. Ordinary code + HTTP orchestrate the fast path; Hankweave runs reasoning hanks. CODE / HTTP HANKWEAVE / PI Complete source viewsOriginal rows, observation IDs and supplied references. No source truncation to fit the Jev request. Typed Jev questionsIdentity, type and unit coherence; then registry / links. Finite choices with probabilities, bound to source IDs. Provisional composition separates the two phases. Probability / consistency gateChosen-label probability and unresolved labels route work. Composition below checks cross-answer consistency. Optional Luna hankReason over flagged questions. Keep the original source view. Oversized views: full-case path. Oversized source view bypasses Jev intact Gate-passed answers Deterministic composition / checksCompose identity clusters, types and predicted links. Check transitivity, type fences and completeness. Unresolved / invalid cases may need adjudication. Optional Gemini hankAdjudicate the unresolved case. Original sources + proposals; proposals are not evidence. Structurally valid result Final deterministic checksNormalize the resolution; retain explicit uncertainty. Validate IDs, coverage and the output contract. Case-local join to source rowsPreserve case namespaces and every original row ID. This is not a new global Ohio identity database. Dashed routes are conditional; passing cases bypass the reasoning hanks. This is the measured program’s control flow—not a measured petabyte deployment.
The measured Jev-first program: code and typed questions on the fast path; Hankweave runs conditional Luna and Gemini reasoning.

Jev gets a complete source view — up to four cases, under 64 kB — and up to 32 questions at once. Each question names its evidence by path (cases[0].observations[1]) and defines its answers in the criteria. What comes back is a label and a probability for every option. Code accepts the label if its probability clears the kind's gate — 0.80 for type, 0.75 for coherence, 0.70 for identity, 0.90 for registry and links — and composes the case: types, identity clusters, then links. Anything that fails a gate, or composes inconsistently, is batched as a short list of questions for a Luna hank; anything that still will not compose goes whole to a Gemini hank.

Hankweave runs those reasoning hanks: a saved program whose steps host model sessions, with the prompt, the response, the session log and the accounting retained per attempt. That retention is what made the audit above possible, what let us recover two Gemini answers from a Markdown fence without buying them again, and what let eleven Luna connection errors fall through to Gemini and keep every batch valid instead of killing the run.

The runnable version of the join is small enough to read in one sitting. Seven synthetic rows, six questions, seven rows back:

Recorded synthetic example · no live model calls

A category answer, joined back to its rows

Follow the repeated label through a shared question, then back to separate output rows. Every name here is invented.

  1. 7original rows
  2. 6unique evidence units / questions
  3. 6recorded Choice answers
  4. 7joined output rows

1 category question saved by exact-label reuse. No rows removed. Keys E1–E6 below are display aliases; exact keys are available inside each trace.

Repeated Mara Quill pair · E1

Original rows

  • synthetic-001
  • synthetic-002

“Contribution received from Mara Quill”

Same exact label → shared question key.

Recorded Choice

E1 evidence / question / answer key

Category
person_donor
P(selected option)
0.97
Confidence
0.97

One answer reused for every row with this key.

Joined back

  • synthetic-001 retained
  • synthetic-002 retained

classified

Category attached; original row IDs preserved.

Exact source-label equality reuses this category answer. It does not establish that the rows identify the same person, and it does not authorize deleting either row.

Inspect E1: exact key and all option probabilities

E1 is a display alias, not a shortened hash. The exact evidence-unit ID is also the question ID and answer lookup key:

eu_bf5677e4335a2905b9297103f1fe2d53a32557c1497afba4f74e7b884ff17b44
person_donor
0.97 selected
organization_donor
0
joint_donors
0
unnamed_donor
0
recipient
0.03
unknown
0
Organization donor · E2

Original rows

  • synthetic-003

“Contribution received from Copper Kite Manufacturing Ltd.”

This exact label has its own question key.

Recorded Choice

E2 evidence / question / answer key

Category
organization_donor
P(selected option)
1
Confidence
1

One answer reused for every row with this key.

Joined back

  • synthetic-003 retained

classified

Category attached; original row IDs preserved.

The label names the organization as the contributor. This category answer does not resolve it to a registry entity.

Inspect E2: exact key and all option probabilities

E2 is a display alias, not a shortened hash. The exact evidence-unit ID is also the question ID and answer lookup key:

eu_e3179acae28859bf6c759c1c2475640d2259eeccd3f35cb067462c35f1522fe2
person_donor
0
organization_donor
1 selected
joint_donors
0
unnamed_donor
0
recipient
0
unknown
0
Joint donors · E3

Original rows

  • synthetic-004

“Joint contribution from Ivo Finch and Nella Finch”

This exact label has its own question key.

Recorded Choice

E3 evidence / question / answer key

Category
joint_donors
P(selected option)
1
Confidence
1

One answer reused for every row with this key.

Joined back

  • synthetic-004 retained

classified

Category attached; original row IDs preserved.

The joint contribution stays joint. The program neither splits this row into individual donors nor invents a division of the contribution.

Inspect E3: exact key and all option probabilities

E3 is a display alias, not a shortened hash. The exact evidence-unit ID is also the question ID and answer lookup key:

eu_2a75601f42569dad570b9a84c5dd3db0a9b81fdeddbb5edfb3fff06be2973adf
person_donor
0
organization_donor
0
joint_donors
1 selected
unnamed_donor
0
recipient
0
unknown
0
Anonymous donor · E4

Original rows

  • synthetic-005

“Contribution received from an anonymous donor”

This exact label has its own question key.

Recorded Choice

E4 evidence / question / answer key

Category
unnamed_donor
P(selected option)
1
Confidence
1

One answer reused for every row with this key.

Joined back

  • synthetic-005 retained

classified

Category attached; original row IDs preserved.

An explicit anonymous donor is classifiable without an identity. No name is inferred.

Inspect E4: exact key and all option probabilities

E4 is a display alias, not a shortened hash. The exact evidence-unit ID is also the question ID and answer lookup key:

eu_679fa8182e5d843baa1c6827b70e188c416b3ad8484a8ab1f2586f284f3ec7f8
person_donor
0
organization_donor
0
joint_donors
0
unnamed_donor
1 selected
recipient
0
unknown
0
Recipient, not donor · E5

Original rows

  • synthetic-006

“Contribution paid to Lantern Orchard Fund (recipient)”

This exact label has its own question key.

Recorded Choice

E5 evidence / question / answer key

Category
recipient
P(selected option)
1
Confidence
1

One answer reused for every row with this key.

Joined back

  • synthetic-006 retained

classified

Category attached; original row IDs preserved.

The label says the fund receives the contribution. Recipient role takes precedence over organization type; this is not an organization_donor answer.

Inspect E5: exact key and all option probabilities

E5 is a display alias, not a shortened hash. The exact evidence-unit ID is also the question ID and answer lookup key:

eu_49d3ac37d1cdfda3a648878a32bc7654452745c142e1e25c7f99a13d238d9cd7
person_donor
0
organization_donor
0
joint_donors
0
unnamed_donor
0
recipient
1 selected
unknown
0
Ambiguous M. Quill · E6

Original rows

  • synthetic-007

“M. Quill — role not recorded”

This exact label has its own question key.

Recorded Choice

E6 evidence / question / answer key

Category
unknown
P(selected option)
0.97
Confidence
0.96

One answer reused for every row with this key.

Joined back

  • synthetic-007 retained

needs_review

Reason: unknown_category. P(selected) = 0.97 does not override an unknown category.

The label records no role. A high probability for unknown supports abstaining; it is not evidence that M. Quill is Mara Quill. These labels have different evidence keys.

Inspect E6: exact key and all option probabilities

E6 is a display alias, not a shortened hash. The exact evidence-unit ID is also the question ID and answer lookup key:

eu_74cd992a40ec5bf49ca5dc1bd66b2d902e2f297797435e411f4ddd7543ed2770
person_donor
0.03
organization_donor
0
joint_donors
0
unnamed_donor
0
recipient
0
unknown
0.97 selected

Why “unknown” still goes to review

The demonstration in example.ts requires both confidence ≥ 0.8 and selected-option probability ≥ 0.9, and always routes unknown to needs_review. These are example policy thresholds, not measured accuracy or the benchmark’s routing policy. Confidence and P(selected option) are separate returned fields.

The join preserves the input rows

rows.map(row => {
  const unit = byLabel.get(row.source_label);
  const judgment = answers.get(unit.evidence_unit_id);
  return { ...row, evidence_unit_id: unit.evidence_unit_id,
           judgment, decision: route(judgment) };
});

Abbreviated from example.ts. byLabel uses exact string equality; answers is keyed by the full evidence-unit ID. No fuzzy name match, aggregation or row deletion occurs.

Every original row returns. Selection marks a shared answer; unmarked rows remain distinct.
Original → output rowExact source labelKey aliasRecorded category
P(selected option)
Route
synthetic-001Contribution received from Mara QuillE1person_donor0.97classified
synthetic-002Contribution received from Mara QuillE1person_donor0.97classified
synthetic-003Contribution received from Copper Kite Manufacturing Ltd.E2organization_donor1classified
synthetic-004Joint contribution from Ivo Finch and Nella FinchE3joint_donors1classified
synthetic-005Contribution received from an anonymous donorE4unnamed_donor1classified
synthetic-006Contribution paid to Lantern Orchard Fund (recipient)E5recipient1classified
synthetic-007M. Quill — role not recordedE6unknown0.97needs_review

The earlier program's bill, and the whole study's spending

The work behind the bill

715typed questions in 65 Jev requests
$0.0320Jev’s priced estimate for those questions
47.35 sof accumulated Jev component time

Reasoning accounts for 71.49% of this workflow’s priced subtotal. Jev is cheap here; the exceptions still dominate the bill.

ModelAttemptsPriced estimateService time
Jev65 (1 unpriced)$0.032047.35 s
Luna31 (3 unpriced)$0.0480209.36 s
Gemini 3.7 Flash3$0.032433.51 s
Total · Jev + selective reasoning99$0.1124290.22 s

The unpriced attempt retains a $1.15 commitment, separate from this subtotal. A commitment is neither a known charge nor a guaranteed invoice ceiling. These figures include the arm’s referrals, adjudication and recorded failures.

What passed the gates, and what still needed reasoning?

663 of 715 Jev answers passed their routing gates; 52 did not. Across whole cases, 165 avoided reasoning and 40 had a reasoning assignment. Cases and questions are different denominators.

Question kindAskedPassed gate
type293273
unit293265
identity118116
relation119

A gate pass is not a correctness verdict. Nine accepted household negatives later disagreed with the reference. The tenth omitted relationship was on the oversized-source path, which bypassed Jev.

What did the entire article study cost?
Priced comparison · all three arms$10.25
Other recorded article-study model work$21.74
Total recorded model estimates$31.99
Separate unpriced commitments$9.58
Protected reserves · not spent$3.70
Authorized envelope$85.00

New article study only, including development and failed/superseded model attempts. Unknown charges are not zero. Catalogue estimates are not invoices; D28/prior-reference construction, local compute and 4a authoring costs are excluded.

The “other work” row is the balance after subtracting the three-arm comparison from recorded study estimates. Do not add unspent reserves to model spend or treat unknown charges as zero.

The default Jev-first receipt is available without JavaScript.

This receipt belongs to the earlier gated comparison, not the new review prefixes. Its Jev line covers 715 questions across 65 calls on panel 2. The new fast pass asks its own questions and has its own bill above. Which model dominates the total depends on the amount and size of the reasoning work; cheap judgments do not make an exception queue free.

Scaling the earlier comparison

The calculator below retains the earlier three-arm measurements. It does not substitute the sparse-review agreement result, and it does not invent an elapsed clock for a smaller review prefix. The latest breakdown above instead keeps the measured calls, actual stage clocks and missing timing measurements visible.

883,012original contribution rows
283,349strict evidence bundles, before identity resolution

3.12 rows per strict bundle. The four held candidate-committee and PAC contribution CSVs for report years 2022 and 2024, captured in June 2026. Not all historical Ohio filings; report year is not necessarily transaction year. Cover registries, later PARTY files and DIME are separate references. Strict bundles group identical evidence under the checked normalization contract; they are not resolved donor identities.

SCENARIO · NOT A MEASURED RUN

883,012 rows · 283,349 strict bundles · 189.91 MB of raw contribution CSV

1,015.59× the comparison’s model work, assuming its case mix, evidence density, batching, referrals and errors repeat.

ProgramPriced model subtotalService effortIdeal 1-lane floor
Fable · direct$9,985238.28 h238.28 h
Cheap reasoning · no Jev$306+ $2,844 commitment scenario201.73 h201.73 h
Jev + selective reasoning$114+ $1,166 commitment scenario81.87 h81.87 h

The commitment column scales the held allowance for missing cost evidence; it is not an observed charge or an upper bound. Service effort is accumulated call time. The ideal-lane floor merely divides it by 1; provider quotas, dependent steps, queues, storage, joins and deployment overhead can make elapsed time longer.

The panel was selected for decision diversity, not population representativeness. It cannot establish full-Ohio quality, error prevalence or actual production cost. Row-equivalent, bundle-equivalent and evidence-byte scenarios are alternative normalizations—not interchangeable facts about storage.

Inspect the full-snapshot byte census
Raw contribution fileData rowsBytes
CAN_CON_2022.csv193,90842,849,560
CAN_CON_2024.csv86,09318,579,375
PAC_CON_2022.csv339,88872,574,196
PAC_CON_2024.csv263,12355,905,307
Total883,012189,908,438

Raw CSV bytes and serialized model-visible evidence bytes differ. The comparison used 1.2 MB of evidence, including references and context. Dividing 189.91 MB of raw CSV by the evidence rate would mix denominators.

Original CR-only files were counted and hashed against the retained, quote-aware transformation receipts. Per-file hashes are in the derived data.

Beyond Ohio: size a Jev input budget, not a storage bill

Large data becomes affordable when repeated bytes do not become repeated judgments. Change the number and size of the evidence units below; the calculator makes the assumption explicit.

$84.00of Jev input, if 1,000,000 questions suffice

1,000,000 questions × 2,000 input tokens × $0.042/million = $84.00

If all 1,000 TB were text at four bytes per token, feeding every byte to Jev would imply $10,500,000 in input cost. These are two different input plans, not a measured compression result.

This is Jev input cost only. You must establish that the selected evidence preserves the task. Reasoning referrals, preparation, storage, scans, joins, transfer and verification are additional. Output-token price was zero in the cited 2026-09-17 price record.

The default Ohio scenario and token-budget example are available without JavaScript.

The full Ohio snapshot has 883,012 contribution rows and 283,349 strict evidence bundles. The calculator scales the measured panel-2 workload to that population under assumptions you can change; it is arithmetic, not a run, and panel 2 is more ordinary than the state. The larger claim is structural: repeated bytes should stop becoming repeated model work. If ordinary code can reduce a much larger source to a million useful 2,000-token questions, Jev's published input price implies $84 for the judgments. Nothing here establishes that such a reduction preserves a task; preparation, referrals, storage and checking each carry their own bill.

What we would tell you to do

Put the exact work in code and let it own the identifiers. Give the model the evidence, not a summary of it. Write each question's criteria to state what suffices under the standard your consumer will accept, then measure each kind of question separately — a gate that was right for typing was wrong for households in the same pipeline. Keep every attempt, so that when the model is confidently wrong you can distinguish the question, the state, the parser and the model. The later experiment adds a budget decision: make the review queue explicit, count unsuccessful reviews, and do not assume a reasoning model can only improve an answer.

Start with the example. Its dry run needs no key. Read one question, one returned probability and one join before deciding which part of your pipeline should think.


Sources and accompanying material

  1. Ohio campaign-finance search and downloads (archived source). The held substrate is the four 2022/2024 report-year candidate/PAC contribution files.
  2. Splink and Dedupe: background approaches, not measured arms.
  3. TypeSafe Choice, state, confidence and model pricing; Hankweave. Prices are the retained September 17, 2026 record.
  4. Methods — reference provenance, both panel draws, the original gates and the later family-review policy; earlier comparison aggregates; cost, clock, byte and corpus ledger, including the latest research and source hashes; earlier anonymous family outcomes.
  5. Runnable source, guide and recorded synthetic output. Raw donor records, private model logs and benchmark labels are not publication material.