CalibratedAgents
Try in console
← Blog · Benchmark

Hallucination detection benchmark (2026): calibrate on 50 labelled answers, find the sentence that's wrong

By Calibrated Agents Research Team · 8 Oct 2026 · Updated 8 Oct 2026 · 7 min read

RAGTruth++ · 100 labelled answers · precision first

6 in 10hallucinated sentences found

Ground59.6%
LLM judge47.9%
Share of hallucinated sentences found · precision 75.5% against 69.5%

On RAGTruth++, given the same 50 or 100 labelled answers, Ground finds hallucinated sentences 4 to 6 points more accurately than an LLM judge calibrated on the same labels, and with no labels it matches the best published system. On whole-answer verdicts, a calibrated judge is level with Ground on both RAGTruth++ and LLM-AggreFact.

The test: same labels for everyone

RAGTruth++ is a re-annotation of 408 answers from the RAGTruth test set: retrieval-augmented question answering and news summaries, written by six models from Llama 2 7B to GPT-4. Two annotators reviewed every answer independently and marked 865 hallucinated spans, ten times as many as the original labels. Three answers in four contain at least one.

Most benchmarks compare detectors with whatever thresholds their authors chose, on some other dataset. That is not how a team uses a detector. A team labels a handful of its own answers, sets the detector on them, and runs it. So that is the test.

  • Every system gets the same 50 or 100 labelled answers from RAGTruth++, and nothing else. No threshold or model comes from another dataset.
  • For the LLM judge, the labels set its threshold. For Ground, they set its threshold and choose its model, from a fixed list of candidates, using those answers only.
  • Everything is then scored once on every answer that was not used: 306 to 354 answers.
  • Answers to the same question never sit on both sides. The draw is repeated 10 times; numbers are means, with the spread across draws.

The judge is Jev, a decision model from TypeSafe, asked whether each sentence is supported by the source. It is a strong baseline: asked well, it matches Ground\'s whole-answer accuracy on LLM-AggreFact, 79.4 balanced accuracy each.

Results: finding the hallucinated sentences

Balanced accuracy on hallucinated sentences, Ground against Jev as a judge, both calibrated on the same labelled answers, in four settings.
Figure 1. Balanced accuracy on hallucinated sentences. Both systems are calibrated on the same labelled answers; the number above each pair is Ground\'s lead. The axis starts at 60.

Ground is ahead in all four settings, by 4.0 to 6.2 points of balanced accuracy. The lead is three or more times the spread between draws, and it grows with more labels: Ground keeps learning from the extra answers, while a judge has only its threshold to move.

SystemLabelsBalanced accuracyPrecisionRecallF1
Ground5077.661.4%73.0%0.666
Jev as a judge5073.652.8%73.8%0.607
Ground10078.659.1%78.1%0.672
Jev as a judge10073.853.1%73.9%0.610

NOTE · Sentence level, both calibrated for balanced accuracy · mean of 10 draws · spread ±0.7 to ±1.4 points for Ground, ±0.8 to ±0.9 for the judge.

When precision matters most

Most production answers are correct, so a detector\'s false alarms matter more than its misses: the arithmetic is in why LLM-as-a-judge fails in production. Calibrated for precision instead (F0.5, which weights precision twice as much as recall), the gap widens.

Precision, recall and F1 on hallucinated sentences with 100 labelled answers and precision-first calibration: Ground 75.5, 59.6, 66.2; Jev as a judge 69.5, 47.9, 56.0.
Figure 2. 100 labelled answers, both calibrated for precision. Ground is more precise and still finds more of the hallucinated sentences.

With 100 labelled answers, three in four sentences Ground flags really are hallucinated, and it still finds 60% of them. The judge, set the same way, finds 48%.

That is the trade a team actually faces: tighten a judge and it goes quiet; Ground stays precise without giving up as much recall.

With no labels at all

Before any calibration, the same comparison against systems published on RAGTruth++. Here Ground uses its built-in rules only: nothing is tuned on this data, the same condition as the published results.

Zero-shot results on all 408 RAGTruth++ answers: span F1 Ground 0.442, RT4CHART 0.407, LettuceDetect 0.235; answer-level F1 Ground 0.778, RT4CHART 0.776, Vectara HHEM 0.424.
Figure 3. All 408 answers. Ground with no labels and no tuning; other systems as reported by Yu et al., 2026. Span F1 is the RAGTruth standard: overlap of the characters marked hallucinated.
SystemSpan F1Answer precisionAnswer recallAnswer F1
Ground, no labels0.4420.8670.7050.778
RT4CHART (published)≈ 0.410.8450.7180.776
LettuceDetect (published)0.235
Vectara HHEM (published)0.424

Without a single label, Ground matches the best published answer-level result and is ahead at locating the hallucinated text. Calibrated on 50 answers, its span F1 rises to 0.54–0.57.

NOTE · Published systems were not run under our calibration protocol; their numbers are as reported, on all 408 answers.

Whole-answer verdicts

Asked only whether a whole answer contains a hallucination, Ground and a calibrated judge are level: 71.6 against 74.3 balanced accuracy with 50 labels, 74.7 against 73.8 with 100, both within the spread between draws. Where Ground pulls ahead is telling you which sentence is wrong, and that is where it matters: an answer that is three-quarters right is useful once the bad sentence is removed.

On LLM-AggreFact too

LLM-AggreFact gathers eleven public datasets of AI-written claims and summaries labelled against their sources, mostly one or two sentences each, so it measures whole-answer verdicts. We ran Ground and Jev as a judge on 1,397 test answers, 127 per dataset, with each system\'s threshold set on 1,197 separate labelled answers.

Balanced accuracyAnswersGroundJev as a judge
All eleven datasets (average)1,39779.479.4
Summaries of a document50876.676.8
Single claims38172.972.4
Multi-sentence answers38185.886.1
Reasoning chains12790.590.5

Level everywhere, and consistent with RAGTruth++: on short answers, a whole-answer judge that is asked well and calibrated is as good a classifier as Ground. Ground\'s advantage is what the judge cannot do: point at the sentence that is wrong and the passage that proves it, at the precision you choose.

NOTE · Each kind of answer within about half a point; none of the differences is beyond chance on these answers.

Why calibrating on your own labels wins

Hallucination is not one fixed thing. The original RAGTruth annotators marked 86 spans in these 408 answers; the RAGTruth++ annotators marked 865 in the same answers. Both are reasonable definitions, and a detector tuned to one under-flags or over-flags the other. Every fixed system, whether a prompted judge, a published pipeline or a trained tagger, carries one definition. Calibration lets a team\'s own labels set it.

Even how a judge is asked matters. On LLM-AggreFact, changing nothing but the order of the fields in the judge\'s request (the answer before the source, instead of after it) changed 10% of its verdicts and moved its balanced accuracy by 2.8 points. A judge\'s number depends on how it is asked and where its threshold sits, and the only way to know is to measure it on your own labelled answers.

A judge calibrated on your labels can move one number, its threshold. Ground\'s checks give a calibrated model many independent signals about each sentence to weigh, so the same 50 answers teach it more.

Test it yourself

  1. 1Label 50 of your own answersMark which answers, or which sentences, are not supported by their sources. Keep answers to the same question together.
  2. 2Calibrate bothSet a judge\'s threshold on those 50, and calibrate Ground on the same 50.
  3. 3Score on the restCompare on answers neither system saw, at the operating point you need: balanced, or precision-first.
  4. 4Repeat with a new drawOne draw of 50 can be lucky. Several draws show whether a lead is real.

The data: RAGTruth++ on Hugging Face (Blue Guardrails). Our benchmark scripts are on GitHub.

How to use this with CalibratedAgents Ground

Ground checks every sentence of an AI answer against its sources, and calibrates on your own labelled answers: tell it what a hallucination means for you, and it returns the sentence that is wrong and the passage that proves it.

Try it on an answer How Ground works API docs

Frequently asked questions

What is RAGTruth++?

A re-annotation by Blue Guardrails of 408 answers from the RAGTruth test set (question answering and summaries). Two annotators reviewed every answer and marked 865 hallucinated spans, against 86 in the original labels.

How was the comparison made fair?

Every system got the same 50 or 100 labelled answers and nothing else; each set its threshold, and Ground chose its model, on those answers only. All were then scored on the remaining 306 to 354 answers, with 10 different draws.

How much better is Ground than an LLM judge?

At finding hallucinated sentences, 4.0 to 6.2 points of balanced accuracy with the same labels. With 100 labels and precision-first calibration, 75.5% precision against 69.5%, with 59.6% recall against 47.9%.

Does Ground need labelled data?

No. With no labels at all it reached answer-level F1 0.778 on RAGTruth++, level with the best published system, and span F1 0.442. Calibrating on 50 of your own answers improves sentence-level accuracy further.

Is Ground better at whole-answer verdicts too?

Not on this benchmark: Ground and a calibrated judge are level at deciding whether a whole answer contains a hallucination. Ground's advantage is in locating the hallucinated sentence.

How does Ground do on LLM-AggreFact?

On 1,397 test answers from its eleven datasets, Ground and Jev as a calibrated judge both score 79.4 balanced accuracy, level on every kind of answer. LLM-AggreFact answers are mostly one or two sentences, so it measures whole-answer verdicts, where a well-calibrated judge is a strong baseline.

Does it matter how a judge is prompted?

Yes. On LLM-AggreFact, changing only the order of the fields in the judge's request changed 10% of its verdicts and its balanced accuracy by 2.8 points. Measure a judge on your own labelled answers before trusting its number.

Why does calibration matter so much?

Because labels define hallucination. The original RAGTruth annotators marked 86 spans in these answers and the RAGTruth++ annotators 865. A detector tuned to one definition misjudges the other; calibrating on your own labels sets the definition you need.

Sources

  1. Blue Guardrails, RAGTruth++ dataset and how it was built: 408 answers, 865 spans, two independent annotators.
  2. Niu et al., RAGTruth (2024): the original corpus and its span-level protocol.
  3. Yu et al., RT4CHART (2026): published RAGTruth++ results for RT4CHART, LettuceDetect and Vectara HHEM.
  4. Kovács and Recski, LettuceDetect (2025): a supervised span-level hallucination detector.
  5. LLM-AggreFact leaderboard and paper (Tang, Laban and Durrett, 2024): the eleven datasets and their test split.
  6. TypeSafe, Jev: the decision model used as the judge.

Cite this

CalibratedAgents Research (2026). Hallucination detection benchmark (2026): calibrate on 50 labelled answers, find the sentence that's wrong. CalibratedAgents. https://calibratedagents.com/blog/hallucination-detection-benchmark-ragtruth-plus-plus-llm-aggrefact
  • RAGTruth++
  • LLM-AggreFact
  • RAGTruth
  • hallucination detection benchmark
  • span-level detection
  • LLM-as-a-judge
  • calibration
  • RT4CHART
  • LettuceDetect
  • Jev

Written by the Calibrated Agents Research Team. Questions or data requests: [email protected]