CalibratedAgents
Try in console
← Blog · Benchmark

Hallucination detection benchmark 2026: 50 labelled answers beat a tuned LLM judge

By Calibrated Agents Research Team · 8 Oct 2026 · Updated 9 Oct 2026 · 8 min read

RAGTruth++ · 100 labelled answers · precision first

6 in 10hallucinated sentences found

Ground60%
LLM judge48%
Share of hallucinated sentences found; three in four of Ground's flags are real

Most hallucination detection benchmarks compare detectors with whatever thresholds their authors chose. Teams don't work that way: they label a handful of their own answers and tune on those. So that is the test we ran, on RAGTruth++ and LLM-AggreFact, with the same labels for every system.

What is RAGTruth++, and why the same labels for everyone?

RAGTruth++ is a re-annotation of 408 answers from the RAGTruth test set: retrieval-augmented question answering and news summaries written by six models, from Llama 2 7B to GPT-4. Two annotators independently marked 865 hallucinated spans, ten times the original labels. Three answers in four contain at least one.

A team adopting a detector labels a handful of its own answers, sets the detector on them, and runs it. That is the test, and it is the same for every system:

  • Same labels: every system gets the same 50 or 100 labelled answers from RAGTruth++, and nothing else. No threshold or model is carried over from another dataset.
  • Same freedom: for the LLM judge, the labels set its threshold. For Ground, they set its threshold and choose its model from a fixed list of candidates, using those answers only.
  • Same test: everything is scored once on every answer that was not used, 306 to 358 answers. Answers to the same question never sit on both sides.
  • Repeated: the draw is repeated 10 times; numbers are means, with the spread across draws.

The judge is Jev, a typed decision model, asked whether each sentence is supported by the source. It is a strong baseline: asked well and calibrated, it ties Ground on whole-answer verdicts across LLM-AggreFact's eleven datasets (Figure 4).

Dumbbell chart of balanced accuracy on hallucinated sentences, RAGTruth++. Balanced, 50 labels: judge 73.6, Ground 77.6 (+4.0). Balanced, 100 labels: 73.8 vs 78.6 (+4.8). Precision-first, 50 labels: 71.3 vs 75.8 (+4.5). Precision-first, 100 labels: 69.7 vs 75.9 (+6.2).
Figure 1. Figure 1. Balanced accuracy on hallucinated sentences, RAGTruth++ (2,334 sentences). Each pair: an LLM judge and Ground calibrated on the same 50 or 100 labelled answers, scored on the 306 to 358 answers not used. Mean of 10 random draws; spread ±0.7 to ±1.4 points for Ground, ±0.8 to ±0.9 for the judge. Axis starts at 67.

Results: finding the hallucinated sentences

Ground is ahead in all four settings, by 4.0 to 6.2 points of balanced accuracy (Figure 1). The lead is three or more times the spread between draws, and it grows with more labels: Ground keeps learning from the extra answers, while a judge has only its threshold to move.

SystemLabelsBalanced accuracyPrecisionRecallF1
Ground5077.661.4%73.0%0.666
LLM judge (Jev)5073.652.8%73.8%0.607
Ground10078.659.1%78.1%0.672
LLM judge (Jev)10073.853.1%73.9%0.610

Sentence level, both calibrated for balanced accuracy, mean of 10 draws. At the same recall, Ground's precision is 6 to 9 points higher: it flags fewer correct sentences for the same number of catches.

When precision matters most

Most production answers are correct, so false alarms cost more than misses. Calibrated for precision instead (F0.5, which weights precision twice as much as recall), the gap widens: with 100 labelled answers, three in four sentences Ground flags are hallucinated, and it still finds 60% of them. The judge, set the same way, finds 48%.

Paired dots for precision, recall and F1 on hallucinated sentences with 100 labelled answers and precision-first calibration: Ground 75.5, 59.6, 66.2; LLM judge 69.5, 47.9, 56.0.
Figure 2. Figure 2. 100 labelled answers from RAGTruth++, both systems calibrated for precision (F0.5), scored on the unused answers; mean of 10 draws. Axis starts at 40.
Precision-first, 100 labelsPrecisionRecallF1
Ground75.5%59.6%0.662
LLM judge (Jev)69.5%47.9%0.560

That is the trade a team actually faces: tighten a judge and it goes quiet; Ground stays precise without giving up as much recall.

With no labels at all

Before any calibration, Ground's built-in rules alone match the best published answer-level result on RAGTruth++ (F1 0.778 against RT4CHART's 0.776) and lead at locating the hallucinated text (span F1 0.442 against about 0.41). Nothing is tuned on this data, the same condition as the published results.

Ranked bars on all 408 RAGTruth++ answers with no labels. Span F1: Ground 0.442, RT4CHART 0.407, LettuceDetect 0.235. Answer-level F1: Ground 0.778, RT4CHART 0.776, Vectara HHEM 0.424.
Figure 3. Figure 3. All 408 answers. Ground with no labels and no tuning; other systems as reported by their authors on RAGTruth++. Span F1 is the RAGTruth standard: overlap of the characters marked hallucinated. Bars start at 0.
SystemSpan F1Answer precisionAnswer recallAnswer F1
Ground, no labels0.4420.8670.7050.778
RT4CHART (published)≈ 0.410.8450.7180.776
LettuceDetect (published)0.235———
Vectara HHEM (published)———0.424

Calibrated on 50 answers, Ground's span F1 rises to 0.54–0.57. Published systems were not run under our calibration protocol; their numbers are as reported, on all 408 answers.

Whole-answer verdicts: level with a calibrated judge

Asked only whether a whole answer contains a hallucination, Ground and a calibrated judge are level: 71.6 against 74.3 balanced accuracy with 50 labels, 74.7 against 73.8 with 100, both within the spread between draws. On LLM-AggreFact, eleven public datasets of mostly one- or two-sentence answers, they tie at 79.4.

Dot plot of balanced accuracy on LLM-AggreFact test answers, Ground against a calibrated LLM judge: all eleven datasets 79.4 vs 79.4; reasoning chains 90.5 vs 90.5; multi-sentence answers 85.8 vs 86.1; summaries 76.6 vs 76.8; single claims 72.9 vs 72.4.
Figure 4. Figure 4. LLM-AggreFact, 1,397 test answers (127 per dataset), each system's threshold set on 1,197 separate labelled answers. Every difference is within half a point; none is beyond chance on these answers. Axis starts at 68.
Balanced accuracyAnswersGroundLLM judge (Jev)
All eleven datasets (average)1,39779.479.4
Summaries of a document50876.676.8
Single claims38172.972.4
Multi-sentence answers38185.886.1
Reasoning chains12790.590.5

Level everywhere, and consistent with RAGTruth++: on short answers, a whole-answer judge that is asked well and calibrated is as good a classifier as Ground. Ground's advantage is what the judge cannot do: point at the sentence that is wrong and the passage that proves it, at the precision you choose. For how much that matters, read Hallucination detection in 2026: find the wrong sentence.

A judge calibrated on your labels can move one number: its threshold. Ground learns which of its signals matter for your definition of a wrong sentence.Why calibration on your own labels wins

Why calibrating on your own labels wins

Hallucination is not one fixed thing. The original RAGTruth annotators marked 86 spans in these 408 answers; the RAGTruth++ annotators marked 865 in the same answers. Both are reasonable definitions, and a detector tuned to one under-flags or over-flags the other. Calibration lets your own labels set the definition.

Even how a judge is asked matters. On LLM-AggreFact, changing nothing but the order of the fields in the judge's request (the answer before the source, instead of after it) changed 10% of its verdicts and moved its balanced accuracy by 2.8 points. A judge's number depends on how it is asked and where its threshold sits, and the only way to know is to measure it on your own labelled answers.

How to run this on your own answers

Label 50 to 100 answers your team has already reviewed, marking the sentences that are wrong. Send them with their sources to Ground, and it is calibrated to your definition of a wrong sentence. Then every new answer comes back with a score and, per sentence, what is wrong, why, where and what to do next.

  • Today: sign in to the CalibratedAgents console and try your own case in the Playground; the free plan covers 1,000,000 tokens a day.
  • This week: send us 20 real conversations with their source documents, and we'll send back what was wrong in each, why, where, and the fix.
  • This month: become a design partner and calibrate Ground on a few hundred of your labelled answers.
Try Ground freeSend us 20 conversations

How we measured

What this does not prove

  • One dataset with sentence labels, in English, from news and web sources; results on your documents will differ, which is the point of calibrating on them.
  • One LLM judge and one prompt. A different judge gives different numbers; the finding that a judge has only its threshold to learn from should hold.
  • Published systems in Figure 3 were not re-run under our protocol; their numbers are as reported.
  • Averages over 10 draws; at 50 labels the spread between draws is about a point either way.

Frequently asked questions

What is the RAGTruth++ hallucination detection benchmark?

RAGTruth++ is a re-annotation of 408 retrieval-augmented answers from the RAGTruth test set, in which two annotators marked 865 hallucinated spans. It is used to measure how well detectors find hallucinated sentences and the exact hallucinated text.

How many labelled answers does Ground need?

None to start: with no labels it matches the best published system on RAGTruth++. With 50 of your own labelled answers it leads an equally tuned LLM judge by 4 points; with 100, by 4.8 to 6.2 depending on whether you tune for balance or precision.

How is this different from the LLM-AggreFact leaderboard?

The leaderboard scores whole answers with fixed thresholds. This benchmark gives every system the same small set of labelled answers to tune on, scores it on the rest, and measures sentences, not only answers. On whole answers the two systems tie; on sentences Ground leads.

Does Ground beat LLM-as-a-judge?

At finding which sentence is hallucinated, yes, by 4 to 6 points on RAGTruth++ with the same labels. At a single yes/no on a whole answer, a well-asked, calibrated judge ties Ground. The difference is in locating the error and explaining it.

What is balanced accuracy, and why use it?

The average of the share of hallucinated sentences caught and the share of correct sentences left alone. It is not inflated by the majority class, which matters when three sentences in four are correct.

Can I rerun the benchmark?

Yes. The runners, the protocol and the data loaders are in the public calibratedagents/benchmarks repository, and the console's free plan covers 1,000,000 tokens a day.

References

  1. RAGTruth++ dataset. Blue Guardrails, 2026.
  2. Yu et al. RT4CHART: span-level hallucination detection on RAGTruth++. 2026.
  3. Tang, Laban, Durrett. MiniCheck: efficient fact-checking of LLMs on grounding documents (introduces LLM-AggreFact). EMNLP 2024.
  4. Niu et al. RAGTruth: a hallucination corpus for developing trustworthy retrieval-augmented language models. ACL 2024.
  5. Kovács, Recski. LettuceDetect: a hallucination detection framework for RAG applications.
  6. Vectara. HHEM-2.1-Open model card.
  7. CalibratedAgents. Benchmark runners and protocol (public repository).
CACalibratedAgents Research is the applied ML team behind Ground, the sentence-level verification layer for AI answers. We publish our benchmarks with their protocols and code so they can be checked and rerun. Questions and corrections: [email protected].

Cite this

CalibratedAgents Research (2026). Hallucination detection benchmark 2026: 50 labelled answers beat a tuned LLM judge. CalibratedAgents. https://calibratedagents.com/blog/hallucination-detection-benchmark
  • Hallucination detection benchmark
  • RAGTruth++
  • LLM-AggreFact
  • Span-level detection
  • LLM-as-a-judge
  • Calibration

Written by the Calibrated Agents Research Team. Questions or data requests: [email protected]