We rewrote one question. Jev went from missing every household link to finding all of them.
An Ohio entity-resolution study: the cost of typed decisions, the cost of reconsidering them, and the clocks that distinguish a fast model from a fast workflow.
Jev alone handled the first pass for a few cents. A small, signal-selected review queue bought more agreement; a larger queue did not eliminate the remaining disagreement. The latest breakdown puts the bill beside output coverage and separates a measured elapsed clock from the sum of time spent in individual calls.
This experiment uses the same second panel as the earlier comparison below. The Jev-only candidate and family-review ranking were frozen before this research pass inspected its outcomes; the Fable and cheap-reasoning controls are retained runs, not new calls. Every agreement figure is against the model-built reference, not human-established truth.
The latest cost and speed breakdown
Jev alone, then a little review
First corrected Jev run on confirmation panel · 200 families · 205 cases · 52 batches · 279 strict bundles · 903 original rows.
The predeclared whole-family queue ranks invalid outputs first, then explicit uncertainty, then minimum selected-choice probability. It is not a fixed confidence cutoff. Source-only Luna sees the complete selected family, not Jev proposals or reference answers.
The sparse-review bill, on its own scale
Each candidate includes the entire first corrected Jev run. Bars show priced subtotals, not invoices; review costs are whole-family calls, never prorated.
Every bar starts at zero. Labels are rounded; the detailed accounting below retains the full estimates.
Full family agreement is agreement with a model-adjudicated reference, not human accuracy. Luna helped construct that reference: source-only review is not reference-independent. All original families and cases remain in the denominators.
Inspect every cost component and call count
| Candidate | Jev base | Completed original reviews | Failed original reviews | Explicit recovery | Total |
|---|---|---|---|---|---|
| Jev only | $0.03367018266 calls · 0 unpriced | $0.00000 calls · 0 unpriced | $0.00000 calls · 0 unpriced | $0.00000 calls · 0 unpriced | $0.03367018266 calls · 0 unpriced |
| 1-family review | $0.03367018266 calls · 0 unpriced | $0.002501151 call · 0 unpriced | $0.00000 calls · 0 unpriced | $0.00000 calls · 0 unpriced | $0.03617133267 calls · 0 unpriced |
| 2-family review | $0.03367018266 calls · 0 unpriced | $0.00425722 calls · 0 unpriced | $0.00000 calls · 0 unpriced | $0.00000 calls · 0 unpriced | $0.03792738268 calls · 0 unpriced |
| 5-family review | $0.03367018266 calls · 0 unpriced | $0.00979025 calls · 0 unpriced | $0.00000 calls · 0 unpriced | $0.00000 calls · 0 unpriced | $0.04346038271 calls · 0 unpriced |
| Ten-family review · first pass | $0.03367018266 calls · 0 unpriced | $0.019721759 calls · 0 unpriced | $0.00394171 call · 0 unpriced | $0.00000 calls · 0 unpriced | $0.05733363276 calls · 0 unpriced |
| Ten-family review · with recovery | $0.03367018266 calls · 0 unpriced | $0.019721759 calls · 0 unpriced | $0.00394171 call · 0 unpriced | $0.00371851 call · 0 unpriced | $0.06105213277 calls · 0 unpriced |
Rows are alternative policies over one shared Jev base, not additive study spending. Recorded model estimates, not provider invoices. Whole-family calls are never prorated. Failed original reviews and the explicit recovery remain charged separately. Research overhead and unpriced reservations are separate from deployment-shaped points; the research ceiling was an internal allocation, not new authorization.
Elapsed time is not accumulated call time
Measured elapsed is one whole invocation. Summed service adds awaited component-call durations, including overhead and unsuccessful included attempts; it is not elapsed time or provider-only latency. Summed stage clocks adds separately recorded invocations; it is not an observed end-to-end deployment.
| Candidate | Measured elapsed | Summed service | Summed stage clocks |
|---|---|---|---|
| Jev only | 28.64 s | 56.2 s | No recorded stage sum |
| 1-family review | Not separately measured | 69.31 s | No recorded stage sum |
| 2-family review | Not separately measured | 80.03 s | No recorded stage sum |
| 5-family review | Not separately measured | 114.89 s | No recorded stage sum |
| Ten-family review · first pass | Not separately measured | 177.18 s | 92.5 s |
| Ten-family review · with recovery | Not separately measured | 188.57 s | 103.89 s |
Inspect the separate stage clocks and recovery
| Candidate | Jev base | Ten-family review | Recovery | Arithmetic sum |
|---|---|---|---|---|
| Ten-family review · first pass | 28.64 s | 63.86 s | Not included | 92.5 s |
| Ten-family review · with recovery | 28.64 s | 63.86 s | 11.39 s | 103.89 s |
Ten-family review · first pass: 5 base misses repaired; 1 formerly matching family regressed. Combined end-to-end elapsed time was not separately measured. Stage sum adds the separately recorded Jev and ten-family review clocks.
Ten-family review · with recovery: 5 base misses repaired; 1 formerly matching family regressed. Combined end-to-end elapsed time was not separately measured. Stage sum adds recorded Jev, ten-family review and explicit recovery clocks; the original failed call remains included.
The corrected repetitions, not a latency promise
The original policy is shown separately from the three corrected repetitions. Answers and call counts varied, as did shared-machine conditions. These are individual observations, not a confidence interval or a provider SLA.
| Run / policy | Cost / work | Measured elapsed | Summed service | Batch service median / p90 | Full / core agreement | Valid cases |
|---|---|---|---|---|---|---|
| Original Jev wordingOriginal wording | $0.03072098461 calls · 709 questions · 731,452 input tokens · 0 unpriced | 25.86 s | 50.64 s | 0.73 / 1.41 s | 166/200 / 171/200 | 202/205 |
| Corrected Jev · run 1Corrected wording | $0.03367018266 calls · 716 questions · 801,671 input tokens · 0 unpriced | 28.64 s | 56.2 s | 0.81 / 1.7 s | 195/200 / 197/200 | 203/205 |
| Corrected Jev · run 2Corrected wording | $0.03464340667 calls · 717 questions · 824,843 input tokens · 0 unpriced | 33.14 s | 65.27 s | 0.93 / 1.85 s | 194/200 / 196/200 | 203/205 |
| Corrected Jev · run 3Corrected wording | $0.03367018266 calls · 716 questions · 801,671 input tokens · 0 unpriced | 31.14 s | 61.25 s | 0.88 / 1.69 s | 195/200 / 196/200 | 203/205 |
Batch medians and p90s describe summed call time within a batch. They must not be compared to another program’s whole-panel elapsed clock.
Compare programs on the same panel
These columns use the same timing definitions. The retained controls ran interleaved in an earlier execution wave; no independent whole-program elapsed clock is claimed for them. Their unpriced attempts remain outside the priced subtotal.
| Program | Full family agreement | Priced model estimate | Summed service | Median batch service |
|---|---|---|---|---|
| Corrected Jev only · first run | 195/200 | $0.033670182All included attempts priced | 56.2 s | 0.81 s |
| Fable · direct | 200/200 | $9.831787All included attempts priced | 844.64 s | 14.61 s |
| Cheap reasoning · no Jev | 199/200 | $0.301246558 unpriced attempts also retained | 715.08 s | 13.27 s |
| Earlier gated Jev + reasoning | 198/200 | $0.1123732284 unpriced attempts also retained | 290.22 s | 6.36 s |
The research bill and the Jev work behind it
The research phase priced $0.564449114 of model work, with $0.426862 held for 4 unpriced attempts. The held amount is a reservation, not a measured charge; it does not turn the priced subtotal into an exact bill.
$10.00 was allocated internally, not newly authorized. Authorization increased by $0.000000; the protected reserve is $0.5000. Candidate costs above reuse one run and cannot be summed to recover the study bill, which also includes other recorded research work.
The shared corrected base used 66 Jev calls for 716 typed questions and 801,671 input tokens.
| Question kind | Questions |
|---|---|
| type | 293 |
| unit | 293 |
| identity | 118 |
| relation | 12 |
| Total | 716 |
Source receipts and content hashes
Public provenance uses descriptive labels and SHA-256 hashes, not private source paths or family identifiers.
- Accepted research result
3362808ebd76f34898051a88e6f552ae6d820c8f6b9ec699cdd0b4c99ccad29a- Pinned input manifest
88d29ac86552b28c44d3e7e9c2445ff0cbd00fb380afcc13700b3857cf7774c6- Acceptance verification
c0116c2a1e51d0160952267ba85b9ab4429f1f3f185307073363e565deafef0a- Accepted Jev research results
3362808ebd76f34898051a88e6f552ae6d820c8f6b9ec699cdd0b4c99ccad29a- Research input manifest
88d29ac86552b28c44d3e7e9c2445ff0cbd00fb380afcc13700b3857cf7774c6- Accepted input verification
c0116c2a1e51d0160952267ba85b9ab4429f1f3f185307073363e565deafef0a- Predeclared confirmation policy and prefixes
ca151ca289862e5bf123b78e2e00a9b9aac8cd7e436111ab955ecdea26c8ef5b- Original Jev wording · report
8ece99387b3cda97c2cd1d6978e0784c172fd52b1f688e9734fba01655bd0043- Corrected Jev · run 1 · report
62b5c65c011142a440b57070ccee81848f1cdf7b2ce7c6e08b559ff31ed92732- Corrected Jev · run 2 · report
1eac586e82844ab619a27f04bea30ce9797dfa29c4b0a8371a52bf267b43fd04- Corrected Jev · run 3 · report
52723534bb0c1557f28c074c72567d44d6d03be6fc57f454a4eb1e304bd074aa- Sparse confirmation prefix report
79d748a49bfb4317c1806459af1f44f7ef05bd816180f47c03a1915a27b1c539- Original ten-family review invocation
50fe704d93d97ebad0c59d3f66e1a9bffb68b35bac505c31ab9cfa9b465ceb1f- Predeclared same-input recovery
00cb2837c8057c94ce03693a2b6ccbaf52114f78aec3609c07033b5b6922165a- Recorded recovery result
dbc38cedf0ec51f4c189d943e819c09fffedc57ae63a66cc729d57e87e3500a1- Sealed study funding summary
f7e94c6b61991f90dcf9b93354a2825e25c0aea31c8711de894c5a9f6419f19a
The earlier experiments explain how we got here. On the first panel, the gated pipeline missed nine of fourteen household links and cost more than the cheap control. We published that result. The next comparison changed the wording of the identity and household questions, not the model or the thresholds. The later Jev-only experiment then asked how much reasoning we could remove. The sections below keep those designs separate.
Inspect the latest bill and clocks, read the case history, compare the earlier gated programs, or follow seven rows through the join.
The rows are not the donors
Ohio's campaign-finance files record reported contributions: a name, an address, an amount, a date, a receiving committee. Our substrate is the candidate-committee and PAC filings for report years 2022 and 2024 — 883,012 rows. To total contributions by donor you must decide which rows belong to the same person, and the strings will not tell you. The names in this article are altered; every field pattern is a real one from the panels. DANIEL R. OKAFOR and DANIEL OKAFOR JR at one address are one man with an inconsistent clerk, or a father and son at the same real-estate firm. WALTER H. BRENNAN SR. and WANDA BRENNAN at one address are two people and one household. A county party's "judicial fund" is an organization until a registry says it is a committee.
That is entity resolution. A false merge adds one donor's money to another's total; a false split scatters one donor across several. Splink and Dedupe are the established probabilistic approaches, and Dedupe's own tutorial uses campaign contributions. We did not benchmark them. Our question was narrower: when the deterministic work is done and a judgment call remains, how much of it can a cheap typed model make, and how do you find out when it is wrong?
The earlier three-arm experiment
We first compared these three programs on the same original evidence, with the same semantic output opportunities and legitimate uncertainty allowed. This is the gated architecture that preceded the budgeted family-review experiment above.
| Program | What runs |
|---|---|
| Fable, directly | The whole case in one lean Messages API call: concise JSON, adaptive thinking at medium effort, a cacheable prefix, no tools, no container. |
| Cheap reasoning, no Jev | Hankweave runs Luna on the whole case; Gemini handles anything unresolved. This control separates "cheaper reasoning model" from "Jev." |
| Jev + selective reasoning | Jev answers typed questions — donor type, unit coherence, identity per pair, then registry and links — with a probability per option. Code composes the case from the answers that pass a gate and sends the rest to a Luna hank; a Gemini hank adjudicates what still will not compose. |
You should know how the reference was made. It is the accepted D28 resolution set: a swarm of cheaper models proposed each case and accepted it on agreement, with Luna proposing on every case and DeepSeek adjudicating most. So the cheap control is graded partly against its own handwriting, and Fable is the only arm that is not. The reference never abstains, so an "unknown" always scores as a miss. Every agreement number here is agreement with that reference, not with a human.
The first panel
Panel 1 was 204 cases in 200 families, frozen before any model ran. Here is what we published on September 17, and what happened when we re-read the retained answers.
| Program | Full | Core | Priced estimate | Median batch | No reasoning model | Links found |
|---|---|---|---|---|---|---|
| Fable 5.1 · direct | 197/200 | 197/200 | $17.36 | 16.10 s | 0/204 | 15/15 |
| Cheap · no Jev | 197/200 | 198/200 | $0.3104 | 12.05 s | 0/204 | 14/15 |
| Jev-first · v1 wording | 190/200 | 200/200 | $0.3639 + 1 unpriced | 7.07 s | 132/204 | 5/15 |
| Jev-first · v2 wording, post hoc | 200/200 | 200/200 | $0.2389 | 6.54 s | 158/204 | 15/15 |
The three original rows are the frozen first comparison. The revised row re-ran only the Jev-first arm after the question criteria were rewritten using that panel's own retained answers, so it is a development result, not a held-out score. Panel 2 is the held-out test of the revised design.
The headline had been "200/200 on the core task." The audit said something less flattering. Of the 132 cases Jev handled alone, 123 had one donor in them; the fast path was mostly labelling a single row person. On identity, Jev gave a definite answer for 54 of 87 pairs and 53 agreed with the reference — but the 0.70 gate admitted 13, and Luna did the rest. On households, Jev was asked ten times and answered no ten times, nine at 0.90 or above; five of those nine pairs had byte-identical street addresses and the same surname. Every household link the pipeline did find came from Luna or Gemini. And the three arms were within noise on core (197, 198 and 200 of 200): Fable's three misses were unknown on two suffix conflicts and committee for the judicial fund, calls a careful analyst might make the same way.
None of this was hidden by the numbers. It was hidden by reading them as a scoreboard instead of a receipt.
The fix: say what suffices
We had written the household criteria the way you write a warning. A shared employer or mailing address is not by itself a household. True, and useless: it told Jev what fails and never what passes. The reference's own reasoning was explicit — a shared residential street address plus a surname is household evidence — and we had never put that sentence in front of the model. The identity criteria had the same defect. Of Jev's 33 unknown identity answers, 27 were pairs with the same first, middle and last name at the same street address or postal code — some differing in employer, some in nothing but punctuation. The reference merges every one of them. Jev sat at fifty-fifty because nobody had told it that was enough.
So we rewrote two questions. Not the thresholds. Not the models. The words.
| Question | v1 said | v2 adds |
|---|---|---|
Household, yes | "supported between distinct compatible entities" | Same house number and street on both records is the accepted evidence; DR against DRIVE or CT against COURT does not break it; employer, business and PO-box addresses do not count |
Identity, same | "supports the same donor identity, including corroborated clerical or name-form variants" | Same surname and given name at the same street address is sufficient; a middle initial or suffix on only one record is a filing variant, not a conflict |
Identity, different | "a genuine generation/name conflict can distinguish people" | Different given names, or JR against SR on both records, are sufficient; different addresses or employers alone are not |
Then we asked Jev the same 87 identity and 13 household questions again, on the same evidence, and compared.
| Panel 1 probe | v1 wording | v2 wording |
|---|---|---|
| Identity: definite answers / agreeing | 56 / 56 | 87 / 86 |
Identity: unknown | 31 | 0 |
Household: answered yes | 0 of 13 | 13 of 13 |
| Household: at or above the untouched 0.90 gate | 0 | 12 |
The one identity disagreement under v2 was KAREN L. against a first-name field holding KAREN L, at p = 0.60, below the gate. It went to Luna, which is where it belonged.
Two false starts are worth your time. Our first probe passed the reference resolution into the model's state as the "upstream proposal," and Jev said yes 13 of 13 times — under the old wording. We had leaked the answer: first the relations, and when we stripped those, the reference's prose explanation. Only when the proposal was composed by the pipeline's own code, with the generic basis the real run uses, did v1 reproduce the failure exactly. The lesson is not about our probe. Jev reads everything in state as evidence, including narrative you labelled untrusted. Put a reviewer's summary in the state and you are asking the model whether the reviewer sounded sure.
With the words fixed and nothing else touched, we re-ran the Jev-first arm on panel 1. It matched 200 of 200 families on the full contract at $0.24 with nothing unpriced, and 158 of 204 cases never saw a reasoning model. That number is real, but it is a development result — the panel's own answers chose the wording. For the gated design, the held-out test was panel 2.
The second panel: the earlier gated comparison
Panel 2 is 205 cases in 200 families drawn from the same frozen population with a new seed, after excluding every atom, family and source case in panel 1. The Jev-first design was frozen with the v2 wording before any outcome. Fable and the cheap control ran unchanged.
The comparison’s model bill
Priced components for the entire 205-case comparison; not an invoice.
Lower is better · every bar starts at zero
Median batch service time
Sum of awaited component-call durations per input batch; not per-arm makespan.
Lower is better · every bar starts at zero
2.3× the median-batch service rate of Fable, at a 87.49× smaller priced subtotal. That cost gap belongs to the whole pipeline: the cheap no-Jev control also removes most of Fable’s bill.
Cheap reasoning · no Jev has 8 unpriced calls with $2.80 held in commitments; Jev + selective reasoning has 4 unpriced calls with $1.15 held in commitments. These were transport or request-limit failures whose fallbacks completed; a plotted subtotal is not that program’s exact total cost.
| Same 200 families | Core agreement | Full agreement |
|---|---|---|
| Fable · direct | 200/200 | 200/200 |
| Cheap reasoning · no Jev | 200/200 | 199/200 |
| Jev + selective reasoning | 198/200 | 198/200 |
Core = typing, identity and registry mapping. Full adds supported relationships and real-donor uncertainty. These are agreements with a model reference, not human-established accuracy.
What exactly counts as a byte or a second?
UTF-8 bytes of the canonical, escaped model-visible source state, counted once per original batch, including supplied references and repeated per-batch policy. Re-derived states match all 52 retained Fable-primary deliveries. This is a common logical-workload denominator, not raw CSV, distinct knowledge, total network traffic or disk throughput.
Sum of full awaited component-call durations, including priced/unpriced attempts and native startup/accounting where applicable. Not independent wall-clock makespan; the arms were interleaved with two batch workers.
The shared numerator is 1,204,763 bytes across 52 batches. For Jev-first: 1,204,763 ÷ 1,000 ÷ 290.22 = 4.15 kB per service second. $0.1124 ÷ 1,204.76 kB = $0.000093 per evidence kB, before the unpriced call.
Download the derived ledgerThe measured default view is available without JavaScript.
Jev handled 165 cases alone — 122 with one donor and 43 with two or more — and its identity gate admitted 116 of 118 pairs, every one agreeing with the reference. It found eleven of the twelve household links: nine straight from Jev at the gate, two from Luna after one Jev request exceeded the model's token limit. The programs differed by at most two full-family matches. The bill and the recorded batch-service clocks describe a larger operational difference; they do not establish a provider-latency guarantee.
The two families Jev-first missed are the honest part of the story, because in both Jev had the reference's answer and the gate did not trust it. M LOUISE HARGREAVE and MARCUS J. HARGREAVE are two people at one address; Jev said different at 0.65, under the 0.70 gate, and the Luna fallback merged them — taking the household link with it. A county party organization drew organization from Jev at 0.62, under the 0.80 gate, and the fallback chain relabelled it. The safety net was the error both times. That is not an argument for lowering gates on the strength of two families. It is a reminder that the fallback model is not an oracle either, and that the retained answers are what let you see which one was wrong.
Two boundaries on this result. By the time we drew panel 2 the fresh population had run out of committee, joint, placeholder and registry-mapping families, so this panel tests typing, identity and households and nothing rarer; the twelve household links are its relationship test. And dispatch paused for twelve failed calls — eleven Luna connection errors, absorbed by Gemini adjudication or a later fallback, and the one oversized Jev request. Each pause was inspected and the identical frozen plan resumed with its completed batches; the failed calls kept their held commitments and are not counted as free.
Inspect all 200 anonymous family outcomes
Compare the measured programs
Full agreement is the default. Core checks typing, assigned-atom identity and registry mapping. Full also checks supported relationships and real-donor uncertainty. These are model-reference agreements, not human-accuracy estimates.
Full agreement · entire 200-family cohort. The graph is never filtered.
Fable 5.1 · direct
Cheap reasoning · no Jev
Jev-first · selective reasoning
Whole-benchmark cost and timing
- Fable 5.1 · direct
- $9.8318 priced model estimate14.61 s median component-sum batch time
- Cheap reasoning · no Jev
- $0.3012 priced model estimate + 8 unpriced calls13.27 s median component-sum batch time
- Jev-first · selective reasoning
- $0.1124 priced model estimate + 4 unpriced calls6.36 s median component-sum batch time
Global figures never change with the family filter. Model estimates are not invoices; development, reference construction and local compute are excluded. A program with unpriced calls has a priced subtotal that is not its exact total cost. Timing sums component-call durations per paired batch, not independent program runtime.
Inspect the anonymous families
A family is a scoring unit of dependent cases, not a household. Numbers are display ordinals, never source identifiers. Jev-only and reasoning filters describe the Jev-first program’s routing.
200 of 200 families shown · Jev-first · Full agreement
Tile key: = match · ≠ disagreement.
Tab into the grid. Use arrow keys to move; Home/End jump to the first/last family. Enter or Space selects a family. Details follow the grid.
No families match this filter. Try All families, or change the program or agreement definition.
Family 1
1 case · 0/1 cases on the Jev-only path · 1 with a reasoning assignment
Fable 5.1 · direct
Core: Match · Full: Match
All case outputs schema/consumer valid.
Issue categories: None recorded.
Cheap reasoning · no Jev
Core: Match · Full: Match
All case outputs schema/consumer valid.
Issue categories: None recorded.
Jev-first · selective reasoning
Core: Match · Full: Match
All case outputs schema/consumer valid.
Issue categories: None recorded.
This is the static full-cohort view. To filter results or choose a family, serve this folder over HTTP with JavaScript enabled. No live model calls are made.
How the earlier gated program is built
Ordinary code keeps the original rows and their identifiers, parses the file-specific field meanings (PAC_REG_NO is the donor's registration in CAN_CON files and the recipient's in PAC_CON files), and checks that every input survives every transformation. None of that is model work.
Jev gets a complete source view — up to four cases, under 64 kB — and up to 32 questions at once. Each question names its evidence by path (cases[0].observations[1]) and defines its answers in the criteria. What comes back is a label and a probability for every option. Code accepts the label if its probability clears the kind's gate — 0.80 for type, 0.75 for coherence, 0.70 for identity, 0.90 for registry and links — and composes the case: types, identity clusters, then links. Anything that fails a gate, or composes inconsistently, is batched as a short list of questions for a Luna hank; anything that still will not compose goes whole to a Gemini hank.
Hankweave runs those reasoning hanks: a saved program whose steps host model sessions, with the prompt, the response, the session log and the accounting retained per attempt. That retention is what made the audit above possible, what let us recover two Gemini answers from a Markdown fence without buying them again, and what let eleven Luna connection errors fall through to Gemini and keep every batch valid instead of killing the run.
The runnable version of the join is small enough to read in one sitting. Seven synthetic rows, six questions, seven rows back:
Recorded synthetic example · no live model calls
A category answer, joined back to its rows
Follow the repeated label through a shared question, then back to separate output rows. Every name here is invented.
- 7original rows
- 6unique evidence units / questions
- 6recorded Choice answers
- 7joined output rows
1 category question saved by exact-label reuse. No rows removed. Keys E1–E6 below are display aliases; exact keys are available inside each trace.
Repeated Mara Quill pair · E1
Original rows
synthetic-001synthetic-002
“Contribution received from Mara Quill”
Same exact label → shared question key.
Recorded Choice
E1 evidence / question / answer key
- Category
person_donor- P(selected option)
- 0.97
- Confidence
- 0.97
One answer reused for every row with this key.
Joined back
synthetic-001retainedsynthetic-002retained
classified
Category attached; original row IDs preserved.
Exact source-label equality reuses this category answer. It does not establish that the rows identify the same person, and it does not authorize deleting either row.
Inspect E1: exact key and all option probabilities
E1 is a display alias, not a shortened hash. The exact evidence-unit ID is also the question ID and answer lookup key:
eu_bf5677e4335a2905b9297103f1fe2d53a32557c1497afba4f74e7b884ff17b44person_donor- 0.97 selected
organization_donor- 0
joint_donors- 0
unnamed_donor- 0
recipient- 0.03
unknown- 0
Organization donor · E2
Original rows
synthetic-003
“Contribution received from Copper Kite Manufacturing Ltd.”
This exact label has its own question key.
Recorded Choice
E2 evidence / question / answer key
- Category
organization_donor- P(selected option)
- 1
- Confidence
- 1
One answer reused for every row with this key.
Joined back
synthetic-003retained
classified
Category attached; original row IDs preserved.
The label names the organization as the contributor. This category answer does not resolve it to a registry entity.
Inspect E2: exact key and all option probabilities
E2 is a display alias, not a shortened hash. The exact evidence-unit ID is also the question ID and answer lookup key:
eu_e3179acae28859bf6c759c1c2475640d2259eeccd3f35cb067462c35f1522fe2person_donor- 0
organization_donor- 1 selected
joint_donors- 0
unnamed_donor- 0
recipient- 0
unknown- 0
Joint donors · E3
Original rows
synthetic-004
“Joint contribution from Ivo Finch and Nella Finch”
This exact label has its own question key.
Recorded Choice
E3 evidence / question / answer key
- Category
joint_donors- P(selected option)
- 1
- Confidence
- 1
One answer reused for every row with this key.
Joined back
synthetic-004retained
classified
Category attached; original row IDs preserved.
The joint contribution stays joint. The program neither splits this row into individual donors nor invents a division of the contribution.
Inspect E3: exact key and all option probabilities
E3 is a display alias, not a shortened hash. The exact evidence-unit ID is also the question ID and answer lookup key:
eu_2a75601f42569dad570b9a84c5dd3db0a9b81fdeddbb5edfb3fff06be2973adfperson_donor- 0
organization_donor- 0
joint_donors- 1 selected
unnamed_donor- 0
recipient- 0
unknown- 0
Anonymous donor · E4
Original rows
synthetic-005
“Contribution received from an anonymous donor”
This exact label has its own question key.
Recorded Choice
E4 evidence / question / answer key
- Category
unnamed_donor- P(selected option)
- 1
- Confidence
- 1
One answer reused for every row with this key.
Joined back
synthetic-005retained
classified
Category attached; original row IDs preserved.
An explicit anonymous donor is classifiable without an identity. No name is inferred.
Inspect E4: exact key and all option probabilities
E4 is a display alias, not a shortened hash. The exact evidence-unit ID is also the question ID and answer lookup key:
eu_679fa8182e5d843baa1c6827b70e188c416b3ad8484a8ab1f2586f284f3ec7f8person_donor- 0
organization_donor- 0
joint_donors- 0
unnamed_donor- 1 selected
recipient- 0
unknown- 0
Recipient, not donor · E5
Original rows
synthetic-006
“Contribution paid to Lantern Orchard Fund (recipient)”
This exact label has its own question key.
Recorded Choice
E5 evidence / question / answer key
- Category
recipient- P(selected option)
- 1
- Confidence
- 1
One answer reused for every row with this key.
Joined back
synthetic-006retained
classified
Category attached; original row IDs preserved.
The label says the fund receives the contribution. Recipient role takes precedence over organization type; this is not an organization_donor answer.
Inspect E5: exact key and all option probabilities
E5 is a display alias, not a shortened hash. The exact evidence-unit ID is also the question ID and answer lookup key:
eu_49d3ac37d1cdfda3a648878a32bc7654452745c142e1e25c7f99a13d238d9cd7person_donor- 0
organization_donor- 0
joint_donors- 0
unnamed_donor- 0
recipient- 1 selected
unknown- 0
Ambiguous M. Quill · E6
Original rows
synthetic-007
“M. Quill — role not recorded”
This exact label has its own question key.
Recorded Choice
E6 evidence / question / answer key
- Category
unknown- P(selected option)
- 0.97
- Confidence
- 0.96
One answer reused for every row with this key.
Joined back
synthetic-007retained
needs_review
Reason: unknown_category. P(selected) = 0.97 does not override an unknown category.
The label records no role. A high probability for unknown supports abstaining; it is not evidence that M. Quill is Mara Quill. These labels have different evidence keys.
Inspect E6: exact key and all option probabilities
E6 is a display alias, not a shortened hash. The exact evidence-unit ID is also the question ID and answer lookup key:
eu_74cd992a40ec5bf49ca5dc1bd66b2d902e2f297797435e411f4ddd7543ed2770person_donor- 0.03
organization_donor- 0
joint_donors- 0
unnamed_donor- 0
recipient- 0
unknown- 0.97 selected
Why “unknown” still goes to review
The demonstration in example.ts requires both confidence ≥ 0.8 and selected-option probability ≥ 0.9, and always routes unknown to needs_review. These are example policy thresholds, not measured accuracy or the benchmark’s routing policy. Confidence and P(selected option) are separate returned fields.
The join preserves the input rows
rows.map(row => {
const unit = byLabel.get(row.source_label);
const judgment = answers.get(unit.evidence_unit_id);
return { ...row, evidence_unit_id: unit.evidence_unit_id,
judgment, decision: route(judgment) };
});Abbreviated from example.ts. byLabel uses exact string equality; answers is keyed by the full evidence-unit ID. No fuzzy name match, aggregation or row deletion occurs.
| Original → output row | Exact source label | Key alias | Recorded category P(selected option) | Route |
|---|---|---|---|---|
synthetic-001Selected trace | Contribution received from Mara Quill | E1 | person_donor0.97 | classified |
synthetic-002Selected trace | Contribution received from Mara Quill | E1 | person_donor0.97 | classified |
synthetic-003Selected trace | Contribution received from Copper Kite Manufacturing Ltd. | E2 | organization_donor1 | classified |
synthetic-004Selected trace | Joint contribution from Ivo Finch and Nella Finch | E3 | joint_donors1 | classified |
synthetic-005Selected trace | Contribution received from an anonymous donor | E4 | unnamed_donor1 | classified |
synthetic-006Selected trace | Contribution paid to Lantern Orchard Fund (recipient) | E5 | recipient1 | classified |
synthetic-007Selected trace | M. Quill — role not recorded | E6 | unknown0.97 | needs_review |
The earlier program's bill, and the whole study's spending
The work behind the bill
Reasoning accounts for 71.49% of this workflow’s priced subtotal. Jev is cheap here; the exceptions still dominate the bill.
| Model | Attempts | Priced estimate | Service time |
|---|---|---|---|
| Jev | 65 (1 unpriced) | $0.0320 | 47.35 s |
| Luna | 31 (3 unpriced) | $0.0480 | 209.36 s |
| Gemini 3.7 Flash | 3 | $0.0324 | 33.51 s |
| Total · Jev + selective reasoning | 99 | $0.1124 | 290.22 s |
The unpriced attempt retains a $1.15 commitment, separate from this subtotal. A commitment is neither a known charge nor a guaranteed invoice ceiling. These figures include the arm’s referrals, adjudication and recorded failures.
What passed the gates, and what still needed reasoning?
663 of 715 Jev answers passed their routing gates; 52 did not. Across whole cases, 165 avoided reasoning and 40 had a reasoning assignment. Cases and questions are different denominators.
| Question kind | Asked | Passed gate |
|---|---|---|
| type | 293 | 273 |
| unit | 293 | 265 |
| identity | 118 | 116 |
| relation | 11 | 9 |
A gate pass is not a correctness verdict. Nine accepted household negatives later disagreed with the reference. The tenth omitted relationship was on the oversized-source path, which bypassed Jev.
What did the entire article study cost?
| Priced comparison · all three arms | $10.25 |
|---|---|
| Other recorded article-study model work | $21.74 |
| Total recorded model estimates | $31.99 |
| Separate unpriced commitments | $9.58 |
| Protected reserves · not spent | $3.70 |
| Authorized envelope | $85.00 |
New article study only, including development and failed/superseded model attempts. Unknown charges are not zero. Catalogue estimates are not invoices; D28/prior-reference construction, local compute and 4a authoring costs are excluded.
The “other work” row is the balance after subtracting the three-arm comparison from recorded study estimates. Do not add unspent reserves to model spend or treat unknown charges as zero.
The default Jev-first receipt is available without JavaScript.
This receipt belongs to the earlier gated comparison, not the new review prefixes. Its Jev line covers 715 questions across 65 calls on panel 2. The new fast pass asks its own questions and has its own bill above. Which model dominates the total depends on the amount and size of the reasoning work; cheap judgments do not make an exception queue free.
Scaling the earlier comparison
The calculator below retains the earlier three-arm measurements. It does not substitute the sparse-review agreement result, and it does not invent an elapsed clock for a smaller review prefix. The latest breakdown above instead keeps the measured calls, actual stage clocks and missing timing measurements visible.
3.12 rows per strict bundle. The four held candidate-committee and PAC contribution CSVs for report years 2022 and 2024, captured in June 2026. Not all historical Ohio filings; report year is not necessarily transaction year. Cover registries, later PARTY files and DIME are separate references. Strict bundles group identical evidence under the checked normalization contract; they are not resolved donor identities.
SCENARIO · NOT A MEASURED RUN
883,012 rows · 283,349 strict bundles · 189.91 MB of raw contribution CSV
1,015.59× the comparison’s model work, assuming its case mix, evidence density, batching, referrals and errors repeat.
| Program | Priced model subtotal | Service effort | Ideal 1-lane floor |
|---|---|---|---|
| Fable · direct | $9,985 | 238.28 h | 238.28 h |
| Cheap reasoning · no Jev | $306+ $2,844 commitment scenario | 201.73 h | 201.73 h |
| Jev + selective reasoning | $114+ $1,166 commitment scenario | 81.87 h | 81.87 h |
The commitment column scales the held allowance for missing cost evidence; it is not an observed charge or an upper bound. Service effort is accumulated call time. The ideal-lane floor merely divides it by 1; provider quotas, dependent steps, queues, storage, joins and deployment overhead can make elapsed time longer.
The panel was selected for decision diversity, not population representativeness. It cannot establish full-Ohio quality, error prevalence or actual production cost. Row-equivalent, bundle-equivalent and evidence-byte scenarios are alternative normalizations—not interchangeable facts about storage.
Inspect the full-snapshot byte census
| Raw contribution file | Data rows | Bytes |
|---|---|---|
| CAN_CON_2022.csv | 193,908 | 42,849,560 |
| CAN_CON_2024.csv | 86,093 | 18,579,375 |
| PAC_CON_2022.csv | 339,888 | 72,574,196 |
| PAC_CON_2024.csv | 263,123 | 55,905,307 |
| Total | 883,012 | 189,908,438 |
Raw CSV bytes and serialized model-visible evidence bytes differ. The comparison used 1.2 MB of evidence, including references and context. Dividing 189.91 MB of raw CSV by the evidence rate would mix denominators.
Original CR-only files were counted and hashed against the retained, quote-aware transformation receipts. Per-file hashes are in the derived data.
Beyond Ohio: size a Jev input budget, not a storage bill
Large data becomes affordable when repeated bytes do not become repeated judgments. Change the number and size of the evidence units below; the calculator makes the assumption explicit.
1,000,000 questions × 2,000 input tokens × $0.042/million = $84.00
If all 1,000 TB were text at four bytes per token, feeding every byte to Jev would imply $10,500,000 in input cost. These are two different input plans, not a measured compression result.
This is Jev input cost only. You must establish that the selected evidence preserves the task. Reasoning referrals, preparation, storage, scans, joins, transfer and verification are additional. Output-token price was zero in the cited 2026-09-17 price record.
The default Ohio scenario and token-budget example are available without JavaScript.
The full Ohio snapshot has 883,012 contribution rows and 283,349 strict evidence bundles. The calculator scales the measured panel-2 workload to that population under assumptions you can change; it is arithmetic, not a run, and panel 2 is more ordinary than the state. The larger claim is structural: repeated bytes should stop becoming repeated model work. If ordinary code can reduce a much larger source to a million useful 2,000-token questions, Jev's published input price implies $84 for the judgments. Nothing here establishes that such a reduction preserves a task; preparation, referrals, storage and checking each carry their own bill.
What we would tell you to do
Put the exact work in code and let it own the identifiers. Give the model the evidence, not a summary of it. Write each question's criteria to state what suffices under the standard your consumer will accept, then measure each kind of question separately — a gate that was right for typing was wrong for households in the same pipeline. Keep every attempt, so that when the model is confidently wrong you can distinguish the question, the state, the parser and the model. The later experiment adds a budget decision: make the review queue explicit, count unsuccessful reviews, and do not assume a reasoning model can only improve an answer.
Start with the example. Its dry run needs no key. Read one question, one returned probability and one join before deciding which part of your pipeline should think.
Sources and accompanying material
- Ohio campaign-finance search and downloads (archived source). The held substrate is the four 2022/2024 report-year candidate/PAC contribution files.
- Splink and Dedupe: background approaches, not measured arms.
- TypeSafe Choice, state, confidence and model pricing; Hankweave. Prices are the retained September 17, 2026 record.
- Methods — reference provenance, both panel draws, the original gates and the later family-review policy; earlier comparison aggregates; cost, clock, byte and corpus ledger, including the latest research and source hashes; earlier anonymous family outcomes.
- Runnable source, guide and recorded synthetic output. Raw donor records, private model logs and benchmark labels are not publication material.