Completed evaluation · raw responses and analysis are public
The Escalation Gate
Testing an AI gate meant checking the benchmark, too.
- Problem
- RAG pipelines pay for generation even when retrieval cannot answer the question.
- Decision
- Separate answerability from generation; use easy and hard negatives to expose different failures.
- Result
- 100% → 69.5% with hard negatives
600 SQuAD-derived decisions · ~$0.13. Floor per author audit; enterprise transfer untested.
Check the evidence
Problem
A retrieval system can return a passage that is on topic but missing the requested fact. Generation still costs money and can produce a confident wrong answer. I tested a narrower question: can a small model identify when a passage cannot answer, before paying for generation? The experiment measures answerability classification, not the quality or cost of a complete RAG system.
My contribution
I designed the three-slice experiment, built the seeded dataset and typed-answer client, retained every raw response, and analyzed calibration, routing and API-reported billing. I then reread the apparent false accepts and corrected the interpretation when benchmark labels disagreed with the passage.
Constraints
- Start with dataset labels and construction, rather than my own opinion; audit those labels when errors look suspicious.
- Keep a contamination control because SQuAD may appear in training data; novel question/passage pairings do not eliminate that risk.
- Separate classification from generation, and route shares from end-to-end cost savings.
- Keep the missing large-model baseline, in-sample thresholds and proxy measurement path visible.
Architecture & consequential decisions
- Question + passageOn topic ≠ answerableThree groups of 200
- Typed decisionAnswerability probabilityThe gate does not generate
- Route the judgmentReject · escalate · passProposed routing, evaluated in-sample
Same model · same prompt · different negatives
THREE SLICES EXPOSE DIFFERENT FAILURE MODES
I used three groups of 200: answerable questions with their original passages; hard negatives that appear answerable but lack the requested fact; and easy negatives paired with unrelated articles. The easy group is a contamination control, not a substitute for realistic misses. A single aggregate would hide the difference between topic mismatch and subtle missing information.
A TYPED DECISION, NOT AN ANSWER
Jev returns a yes/no probability, called a noul, for whether the passage can answer the question. I record the numeric result, raw response, latency and billed cost before analysis. The gate never writes the answer. A confident yes still forwards the passage to generation; only the reject path avoids that call outright.
CALIBRATION AGAINST A NOISE FLOOR
On the 400-item answerable-plus-hard-negative mix, expected calibration error was 0.047 against a 0.039 resampled noise floor. I simulated a perfectly calibrated predictor 200 times on the same sample to interpret that aggregate. The top confidence bucket held 164 of 400 items and was dominated by answerable examples; it carried much of the aggregate.
THE MIDDLE BUCKETS CHANGE THE INTERPRETATION
In the 0.95-1.00 confidence bucket, 160 of 164 decisions were correct: 97.6% observed accuracy at about 97.4% stated confidence. That does not validate every probability. In the 0.70-0.75 bucket, stated confidence was about 71.9% but observed accuracy was 42.9%, on only 14 items. The middle buckets were least reliable, precisely where escalation is useful. Small buckets also make those gaps uncertain.
ROUTING WITH THE DENOMINATOR ATTACHED
At a 0.10-0.90 band over the realistic 400-item mix, 58 decisions (14.5%) reject before generation, 170 (42.5%) escalate the judgment, and 172 (43%) pass to generation. The reject group includes one answerable passage wrongly dropped. These thresholds were selected from the same observations, not a held-out validation set. No end-to-end RAG saving or escalated-model outcome was measured.
READING THE ERRORS CHANGED THE STORY
I reread all 61 original-label false accepts. At least 6 of the 10 highest-confidence accepts contained the answer despite SQuAD calling them unanswerable. That is why the repository treats the original 69.5% hard-negative score as a floor, rather than an independently corrected estimate. This was my non-blind review, with model scores visible; it is not independent relabelling. I publish no corrected aggregate accuracy.
NAME THE REMAINING WEAKNESS
Genuine mistakes clustered around swapped names or entities, negation, false premises and altered numbers. Those are more useful test cases than unrelated passages: a retrieved chunk can be relevant while a question subtly changes what it says. The next check described in the repository is blind independent relabelling, with disagreements reported rather than silently replacing labels.
MEASURE THE PATH AND THE BILL
The 2026-09-24 run used jev-1.13.0 for 600 decisions and billed $0.127017, rounded here to $0.13. Median wall-clock latency was 662 ms through a third-party proxy; it is not direct-model latency or a strict bound on a different network route. Calls carried one question each to preserve per-decision timing. I used API-reported billing rather than conflicting published token prices.
THE COMPARISON I STILL OWE
No large-model baseline was run. The experiment compares the small model’s confident and uncertain slices, not its performance against a frontier model. The repository describes running the identical items and instruction text through small and frontier chat models as the next comparison. Without it, I cannot claim the gate improves a production pipeline.
Alternatives considered
Use random negatives as the whole benchmark
Trade-off: They scored 100%, but do not resemble an on-topic retrieval miss. Hard negatives scored 69.5% against the original labels. A trivial always-unanswerable classifier also scores perfectly on negative-only slices, so that score alone cannot establish useful discrimination.
Count every confident yes as a saved call
Trade-off: A yes is an answerability decision, not a generated answer. It still pays for generation. Only the 14.5% reject share avoids that call outright on the measured mix.
Batch questions in each request
Trade-off: Batching could share state cost, but would prevent isolated per-decision latency measurement. I accepted a slower, more expensive request shape for this experiment; it is not a production batching recommendation.
Treat dataset labels as unquestionable
Trade-off: Reading the highest-confidence apparent errors exposed label mistakes. Replacing labels with my own score-informed judgment would introduce another bias, so I retain original-label results and report the non-blind review separately.
What the results show
Experiment, not deployment
600 decisions · ~$0.13
Three groups of 200, SQuAD 2.0, run on 2026-09-24. API-reported total $0.127017; all raw responses are committed.
Original-label slice scores
100% → 69.5%
200 easy versus 200 hard negatives; 200/200 versus 139/200 correct. The hard-negative figure is treated as a floor after the author’s label audit, not independently relabelled accuracy.
Aggregate calibration
0.047 ECE · 0.039 floor
400-item realistic mix. The resampled sampling-noise floor gives context; one large favorable bin carries the aggregate and middle bins remain unreliable.
One confidence bucket
97.6% accuracy
160/164 correct in the 0.95-1.00 confidence bucket, at about 97.4% stated confidence. This is not a calibration score or a production validation.
Observed routing shares
14.5% reject · 42.5% escalate · 43% pass
400 items at the 0.10-0.90 band, selected in-sample. One answerable item was wrongly rejected. Passing still requires generation; these are not end-to-end savings.
Author’s label audit
At least 6 of 10
Highest-confidence apparent false accepts with answers in the passage, according to a non-blind single-author review. Independent relabelling has not been done.
Limitations
- I haven’t compared it with a large-model baseline yet. This does not establish that Jev beats a small or frontier chat model.
- One dataset, task and language. Transfer to enterprise retrieval is untested; the public benchmark may be in training data.
- Routing thresholds were chosen from the same 400 items. They are in-sample observations, not validated production estimates.
- The label audit was not blind or independent. The original-label score remains published; no corrected accuracy is claimed.
- Latency includes a proxy hop and one question per request. It does not isolate model speed or estimate a batched production path.
- Only classification billing was measured. Neither escalated judgments nor downstream generation were run, so no complete RAG saving is established.
Reproduce it yourself
- From a local checkout of the linked repository, Python 3.11+ is sufficient; no third-party dependencies are required.
- For a new billed run, set your key: export JEV_API_KEY=...
- Rebuild the seeded slices: python src/dataset.py --n 200
- Run new decisions: python src/run_experiment.py
- Recompute the committed results without API calls or spend: python src/analyze.py
- Read docs/VERIFICATION.md and docs/ADJUDICATION.md alongside results/report.json; the raw responses remain the measurement inputs.