Hallucination detection benchmark 2026: 50 labelled answers beat a tuned LLM judge
By Calibrated Agents Research Team · 8 Oct 2026 · Updated 9 Oct 2026 · 8 min read
RAGTruth++ · 100 labelled answers · precision first
6 in 10hallucinated sentences found
Most hallucination detection benchmarks compare detectors with whatever thresholds their authors chose. Teams don't work that way: they label a handful of their own answers and tune on those. So that is the test we ran, on RAGTruth++ and LLM-AggreFact, with the same labels for every system.
What is RAGTruth++, and why the same labels for everyone?
RAGTruth++ is a re-annotation of 408 answers from the RAGTruth test set: retrieval-augmented question answering and news summaries written by six models, from Llama 2 7B to GPT-4. Two annotators independently marked 865 hallucinated spans, ten times the original labels. Three answers in four contain at least one.
A team adopting a detector labels a handful of its own answers, sets the detector on them, and runs it. That is the test, and it is the same for every system:
- Same labels: every system gets the same 50 or 100 labelled answers from RAGTruth++, and nothing else. No threshold or model is carried over from another dataset.
- Same freedom: for the LLM judge, the labels set its threshold. For Ground, they set its threshold and choose its model from a fixed list of candidates, using those answers only.
- Same test: everything is scored once on every answer that was not used, 306 to 358 answers. Answers to the same question never sit on both sides.
- Repeated: the draw is repeated 10 times; numbers are means, with the spread across draws.
The judge is Jev, a typed decision model, asked whether each sentence is supported by the source. It is a strong baseline: asked well and calibrated, it ties Ground on whole-answer verdicts across LLM-AggreFact's eleven datasets (Figure 4).

Results: finding the hallucinated sentences
Ground is ahead in all four settings, by 4.0 to 6.2 points of balanced accuracy (Figure 1). The lead is three or more times the spread between draws, and it grows with more labels: Ground keeps learning from the extra answers, while a judge has only its threshold to move.
| System | Labels | Balanced accuracy | Precision | Recall | F1 |
|---|---|---|---|---|---|
| Ground | 50 | 77.6 | 61.4% | 73.0% | 0.666 |
| LLM judge (Jev) | 50 | 73.6 | 52.8% | 73.8% | 0.607 |
| Ground | 100 | 78.6 | 59.1% | 78.1% | 0.672 |
| LLM judge (Jev) | 100 | 73.8 | 53.1% | 73.9% | 0.610 |
Sentence level, both calibrated for balanced accuracy, mean of 10 draws. At the same recall, Ground's precision is 6 to 9 points higher: it flags fewer correct sentences for the same number of catches.
When precision matters most
Most production answers are correct, so false alarms cost more than misses. Calibrated for precision instead (F0.5, which weights precision twice as much as recall), the gap widens: with 100 labelled answers, three in four sentences Ground flags are hallucinated, and it still finds 60% of them. The judge, set the same way, finds 48%.

| Precision-first, 100 labels | Precision | Recall | F1 |
|---|---|---|---|
| Ground | 75.5% | 59.6% | 0.662 |
| LLM judge (Jev) | 69.5% | 47.9% | 0.560 |
That is the trade a team actually faces: tighten a judge and it goes quiet; Ground stays precise without giving up as much recall.
With no labels at all
Before any calibration, Ground's built-in rules alone match the best published answer-level result on RAGTruth++ (F1 0.778 against RT4CHART's 0.776) and lead at locating the hallucinated text (span F1 0.442 against about 0.41). Nothing is tuned on this data, the same condition as the published results.

| System | Span F1 | Answer precision | Answer recall | Answer F1 |
|---|---|---|---|---|
| Ground, no labels | 0.442 | 0.867 | 0.705 | 0.778 |
| RT4CHART (published) | ≈ 0.41 | 0.845 | 0.718 | 0.776 |
| LettuceDetect (published) | 0.235 | — | — | — |
| Vectara HHEM (published) | — | — | — | 0.424 |
Calibrated on 50 answers, Ground's span F1 rises to 0.54–0.57. Published systems were not run under our calibration protocol; their numbers are as reported, on all 408 answers.
Whole-answer verdicts: level with a calibrated judge
Asked only whether a whole answer contains a hallucination, Ground and a calibrated judge are level: 71.6 against 74.3 balanced accuracy with 50 labels, 74.7 against 73.8 with 100, both within the spread between draws. On LLM-AggreFact, eleven public datasets of mostly one- or two-sentence answers, they tie at 79.4.

| Balanced accuracy | Answers | Ground | LLM judge (Jev) |
|---|---|---|---|
| All eleven datasets (average) | 1,397 | 79.4 | 79.4 |
| Summaries of a document | 508 | 76.6 | 76.8 |
| Single claims | 381 | 72.9 | 72.4 |
| Multi-sentence answers | 381 | 85.8 | 86.1 |
| Reasoning chains | 127 | 90.5 | 90.5 |
Level everywhere, and consistent with RAGTruth++: on short answers, a whole-answer judge that is asked well and calibrated is as good a classifier as Ground. Ground's advantage is what the judge cannot do: point at the sentence that is wrong and the passage that proves it, at the precision you choose. For how much that matters, read Hallucination detection in 2026: find the wrong sentence.
A judge calibrated on your labels can move one number: its threshold. Ground learns which of its signals matter for your definition of a wrong sentence.Why calibration on your own labels wins
Why calibrating on your own labels wins
Hallucination is not one fixed thing. The original RAGTruth annotators marked 86 spans in these 408 answers; the RAGTruth++ annotators marked 865 in the same answers. Both are reasonable definitions, and a detector tuned to one under-flags or over-flags the other. Calibration lets your own labels set the definition.
Even how a judge is asked matters. On LLM-AggreFact, changing nothing but the order of the fields in the judge's request (the answer before the source, instead of after it) changed 10% of its verdicts and moved its balanced accuracy by 2.8 points. A judge's number depends on how it is asked and where its threshold sits, and the only way to know is to measure it on your own labelled answers.
How to run this on your own answers
Label 50 to 100 answers your team has already reviewed, marking the sentences that are wrong. Send them with their sources to Ground, and it is calibrated to your definition of a wrong sentence. Then every new answer comes back with a score and, per sentence, what is wrong, why, where and what to do next.
- Today: sign in to the CalibratedAgents console and try your own case in the Playground; the free plan covers 1,000,000 tokens a day.
- This week: send us 20 real conversations with their source documents, and we'll send back what was wrong in each, why, where, and the fix.
- This month: become a design partner and calibrate Ground on a few hundred of your labelled answers.
How we measured
What this does not prove
- One dataset with sentence labels, in English, from news and web sources; results on your documents will differ, which is the point of calibrating on them.
- One LLM judge and one prompt. A different judge gives different numbers; the finding that a judge has only its threshold to learn from should hold.
- Published systems in Figure 3 were not re-run under our protocol; their numbers are as reported.
- Averages over 10 draws; at 50 labels the spread between draws is about a point either way.
Frequently asked questions
What is the RAGTruth++ hallucination detection benchmark?
RAGTruth++ is a re-annotation of 408 retrieval-augmented answers from the RAGTruth test set, in which two annotators marked 865 hallucinated spans. It is used to measure how well detectors find hallucinated sentences and the exact hallucinated text.
How many labelled answers does Ground need?
None to start: with no labels it matches the best published system on RAGTruth++. With 50 of your own labelled answers it leads an equally tuned LLM judge by 4 points; with 100, by 4.8 to 6.2 depending on whether you tune for balance or precision.
How is this different from the LLM-AggreFact leaderboard?
The leaderboard scores whole answers with fixed thresholds. This benchmark gives every system the same small set of labelled answers to tune on, scores it on the rest, and measures sentences, not only answers. On whole answers the two systems tie; on sentences Ground leads.
Does Ground beat LLM-as-a-judge?
At finding which sentence is hallucinated, yes, by 4 to 6 points on RAGTruth++ with the same labels. At a single yes/no on a whole answer, a well-asked, calibrated judge ties Ground. The difference is in locating the error and explaining it.
What is balanced accuracy, and why use it?
The average of the share of hallucinated sentences caught and the share of correct sentences left alone. It is not inflated by the majority class, which matters when three sentences in four are correct.
Can I rerun the benchmark?
Yes. The runners, the protocol and the data loaders are in the public calibratedagents/benchmarks repository, and the console's free plan covers 1,000,000 tokens a day.
References
- RAGTruth++ dataset. Blue Guardrails, 2026.
- Yu et al. RT4CHART: span-level hallucination detection on RAGTruth++. 2026.
- Tang, Laban, Durrett. MiniCheck: efficient fact-checking of LLMs on grounding documents (introduces LLM-AggreFact). EMNLP 2024.
- Niu et al. RAGTruth: a hallucination corpus for developing trustworthy retrieval-augmented language models. ACL 2024.
- Kovács, Recski. LettuceDetect: a hallucination detection framework for RAG applications.
- Vectara. HHEM-2.1-Open model card.
- CalibratedAgents. Benchmark runners and protocol (public repository).
Cite this
CalibratedAgents Research (2026). Hallucination detection benchmark 2026: 50 labelled answers beat a tuned LLM judge. CalibratedAgents. https://calibratedagents.com/blog/hallucination-detection-benchmark