CalibratedAgents
Try in console
← Blog · Research

Hallucination detection in 2026: find the wrong sentence, not just the wrong answer

By Calibrated Agents Research Team · 9 Oct 2026 · Updated 9 Oct 2026 · 9 min read

RAGTruth++ · finding the hallucinated sentence

0.877AUROC for Ground, ahead of every system tested

Ground0.877
Microsoft-D10.848
LLM judge0.830
5,916 human-labelled sentences across RAGTruth++ and FaithBench · bars start at 0.70

An AI answer is rarely all wrong. Usually one or two sentences are, and an answer-level hallucination detection check can't tell you which. We measured that gap on 5,916 human-labelled sentences from two public benchmarks, and how far checking each sentence closes it.

Why answer-level hallucination detection isn't enough

Because a hallucinated answer is mostly correct. On RAGTruth++, a flagged answer has about two wrong sentences in five or six; on FaithBench, under two in five. A single verdict for the whole answer can only block it, send all of it to review, or regenerate it blind.

Most hallucination detection tools return one groundedness or faithfulness score per answer. That worked when answers were one line. It stops working when a RAG assistant or an agent writes a paragraph from a policy document and the only mistake is a refund percentage in the third sentence. The same five-sentence answer looks like this under each kind of check:

Three things go wrong with the left-hand card:

  • You throw away good answers. Block every answer with a hallucination and you discard 7 in 10 answers on these benchmarks, most of whose content was right.
  • Your reviewers do the detective work. A flagged answer tells them to distrust it, not where to look. They reread the whole answer and the whole source to find one or two sentences.
  • Regeneration is blind. Ask the model to try again without saying what was wrong, and it can repeat the error or introduce a new one.

Retrieval does not make this go away. With RAG, hallucinations shift from obvious inventions to subtle mismatches between the evidence and the answer [7], and a high average groundedness score can hide a much higher rate of individual wrong statements [6]. The unit that matters is the sentence.

A flag on the whole answer tells you to distrust it. A flag on the sentence tells you what to fix.The case for sentence-level hallucination detection
Waffle chart, each square 1%. RAGTruth++: 75% of answers contain a hallucination, but only 28% of sentences are wrong. FaithBench: 68% of summaries contain one, only 25% of sentences are wrong.
Figure 1. Figure 1. Share of answers with at least one hallucinated sentence against the share of sentences that are wrong. Human labels from RAGTruth++ (408 answers, 2,334 sentences) and FaithBench (750 summaries, 3,582 sentences).

What sentence-level hallucination detection gives you

For every sentence its source doesn't support, Ground returns four things: what is wrong (it contradicts the source, the source doesn't say it, or it rests on an earlier error), why, where the source shows it, and what next: the corrected sentence. The whole answer also gets one score, so you can route it.

Here is what that looks like for one sentence of an AI tutor's answer about TLS 1.3, checked against the course material:

Sentence“…data is encrypted with a symmetric cipher such as AES-GCM, which supports keys of up to 512 bits.”
WhatContradicts the source
WhyThe number is different: the source gives AES key sizes of 128, 192 and 256 bits.
Where“AES (Advanced Encryption Standard) is the most widely used symmetric cipher and supports key sizes of 128, 192 and 256 bits.”
What nextCorrect it to what the source says: AES supports keys of 128, 192 or 256 bits.

That changes what each team can do with a flag. A reviewer reads one sentence and one passage instead of the whole answer. An agent regenerates one sentence with the fix as a hint. A product team keeps the four good sentences and corrects the fifth.

Benchmark: Ground vs LLM-as-a-judge vs open-source hallucination detectors

On two human-labelled benchmarks, Ground finds hallucinated sentences better than every system we tested: an LLM judge asked about each sentence, and five open-source detectors, including Vectara HHEM, LettuceDetect and IBM Granite Guardian. Ground scores 0.877 AUROC on RAGTruth++ and 0.741 on FaithBench.

Dot plot of sentence-level AUROC per detector, both benchmarks on one axis. RAGTruth++: Ground 0.877, LLM-as-a-judge 0.830, LettuceDetect large 0.743, LettuceDetect v2 0.740, HHEM 0.718, DeBERTa NLI 0.660, Granite Guardian 0.653. FaithBench: Ground 0.741, judge 0.701, LettuceDetect v2 0.686, LettuceDetect large 0.675, HHEM 0.615, Granite Guardian 0.570, DeBERTa NLI 0.559.
Figure 2. Figure 2. Sentence-level AUROC on all 2,334 RAGTruth++ sentences and all 3,582 FaithBench sentences; 0.5 is chance. Ground is calibrated on each dataset's labels with grouped 5-fold cross-validation; the LLM judge and the open-source models are scored as released. Axes start at 0.5.
Hallucination detectorRAGTruth++FaithBenchType
CalibratedAgents Ground0.8770.741Sentence-level verification
LLM-as-a-judge (Jev, per sentence)0.8300.701LLM judge
LettuceDetect large (ModernBERT)0.7430.675Open source, token-level
LettuceDetect v2 (mmBERT)0.7400.686Open source, token-level
Vectara HHEM-2.1-Open0.7180.615Open source, consistency score
DeBERTa-v3 NLI0.6600.559Open source, entailment
IBM Granite Guardian 3.3 (groundedness)0.6530.570Open source, guardian model

Two notes for reading Figure 2. FaithBench is the hardest public benchmark for this task; its authors report that the best detectors sit near 50% accuracy on it [2], so every score there is low. And Granite Guardian is designed to judge a whole response in context; scored one sentence at a time, it over-flags.

When precision matters most

Tuned for precision on 100 labelled answers from RAGTruth++, Ground's flags were right 75.5% of the time against 69.5% for an LLM judge tuned on the same answers, while catching 59.6% of hallucinated sentences against 47.9%. Fewer false alarms and more catches, at the same time.

Ground gets better with your data. An LLM judge doesn't.

Give both systems the same labelled answers from your own use case, and the gap widens as you add labels. On FaithBench, Ground goes from level with an LLM judge at 50 labels to +5.6 points at 300; on RAGTruth++, from +3.7 at 50 labels to +6.2 at 200. The judge barely moves.

Left: FaithBench sentence-level balanced accuracy against labelled summaries, Ground rising from 63.5 at 50 labels to 68.2 at 300 while the judge stays near 63. Right: Ground's lead over the judge in points at each label budget, FaithBench +0.4 to +5.6 and RAGTruth++ +3.7 to +6.2.
Figure 3. Figure 3. Sentence-level balanced accuracy when both systems are calibrated on exactly the same labelled answers, drawn by whole question, and scored on every answer they never saw. Averages over 10 random draws at 50 and 100 labels and 5 draws above. Left axis starts at 60; right panel shows Ground minus the judge, in points.
Labelled answersFaithBench: GroundFaithBench: LLM judgeRAGTruth++: GroundRAGTruth++: LLM judge
5063.563.177.373.6
10065.863.078.573.8
20067.362.979.373.1
30068.262.6——
40067.963.8——

Why: calibrating a judge only moves its threshold, and 50 labels already pin that down. Ground learns which of its signals matter for your definition of a wrong sentence, so every extra label your team has already reviewed makes it sharper. Most of the gain arrives by 200 to 300 labels.

Where LLM-as-a-judge misses hallucinations

LLM judges are good at spotting invented content and weak at distortions of what the source does say: a reversed relation, a changed qualifier, a fact moved to the wrong subject. On FaithBench, a judge catches 88% of added information but only 56% of distorted facts.

Lollipop chart of the share of hallucinated sentences an LLM judge catches, with the share that slips through marked in red: obvious invented content 96%, added information 88%, obvious contradiction 84%, subtle invented content 79%, subtle contradiction 68%, distorted source facts 56%.
Figure 4. Figure 4. Share of hallucinated sentences an LLM judge (per sentence, at its best balanced-accuracy threshold) catches, by the benchmarks' own error labels: RAGTruth++ conflict and baseless-information types, FaithBench intrinsic and extrinsic types.

These are the hallucinations that survive a quick read: sentences built from the source's own words, slightly bent. They are also the ones that cost the most in regulated work, where the policy says 50% and the answer says full refund. Checking each sentence against the exact passage it should rest on is how you catch them.

How to add sentence-level hallucination detection to your RAG app or agent

Send each answer with its question and source documents to the Ground API, or paste a case into the console. You get the overall score and, for every sentence that needs it, what, why, where and what next. Then label 100 to 300 of your own answers to calibrate it to your domain.

  • Today: sign in to the CalibratedAgents console with Google and try your own case in the Playground. The free plan covers 1,000,000 tokens a day.
  • This week: send us 20 real conversations with their source documents, and we'll send back what was wrong in each, why, where, and the fix.
  • This month: become a design partner. Label a few hundred answers your team has already reviewed, and Ground is calibrated to your own definition of a wrong answer.
Try Ground freeSend us 20 conversations

How we measured

What this does not prove

  • Two English benchmarks, mostly news and web text. Your documents may behave differently, which is why calibrating on your own labels matters.
  • One LLM judge and one prompt. Other judges will give other numbers, though we expect the shape (flat with more labels) to hold.
  • Averages over random draws; at 50 labels the spread between draws is several points.
  • Part of FaithBench's labels are marked questionable by the annotators themselves; no detector does well there.

Frequently asked questions

What is hallucination detection?

Hallucination detection checks whether an AI answer contains statements that its source documents don't support or that contradict them. Sentence-level hallucination detection goes further: it marks which sentence is wrong, and why.

How do you detect hallucinations in a RAG application?

Check each sentence of the generated answer against the retrieved passages it should rest on. Flag sentences that contradict the passages or that the passages don't mention, and return the passage and a fix so the answer can be reviewed or regenerated.

Is LLM-as-a-judge good enough for hallucination detection?

For a quick yes/no on short answers, a well-calibrated LLM judge is a strong baseline. For long answers it misses distortions of the source and can't learn from your labels: in our benchmark it caught 56% of distorted facts, and tuning it on more labels didn't improve it.

What is the best hallucination detection model in 2026?

On our sentence-level benchmark, CalibratedAgents Ground scored highest on both RAGTruth++ (0.877 AUROC) and FaithBench (0.741), ahead of an LLM judge and of the open-source detectors HHEM-2.1-Open, LettuceDetect, Granite Guardian 3.3 and DeBERTa NLI.

Do I need labelled data to start?

No. Ground works out of the box. Labelling 100 to 300 answers your team has already reviewed calibrates it to your domain, which is where its lead over an LLM judge reached 4 to 6 points.

What does Ground return for a hallucinated sentence?

What is wrong (it contradicts the source, the source doesn't say it, or it rests on an earlier error), why, where the source shows it, and what next: the corrected sentence. The whole answer also gets one score.

Can I try it for free?

Yes. Sign in to the console and check your own cases in the Playground or through the API. The free plan covers 1,000,000 tokens a day.

References

  1. RAGTruth++ dataset. Blue Guardrails, 2026.
  2. Tamber et al. FaithBench: a diverse hallucination benchmark for summarization by modern LLMs. NAACL 2025.
  3. Jev (TypeSafe AI), a typed decision model, used here as an LLM judge asked about each sentence.
  4. Tang, Laban, Durrett. MiniCheck: efficient fact-checking of LLMs on grounding documents (introduces LLM-AggreFact). EMNLP 2024.
  5. Kovács, Recski. LettuceDetect: a hallucination detection framework for RAG applications. 2025–2026.
  6. FutureAGI. LLM hallucination: a 2026 architectural deep dive.
  7. PatSnap. LLM hallucination rate evaluation for engineering. 2026.
  8. Vectara. HHEM-2.1-Open model card.
  9. IBM. Granite Guardian documentation.
CACalibratedAgents Research is the applied ML team behind Ground, the sentence-level verification layer for AI answers. We publish our benchmarks with their protocols and code so they can be checked and rerun. Questions and corrections: [email protected].

Cite this

CalibratedAgents Research (2026). Hallucination detection in 2026: find the wrong sentence, not just the wrong answer. CalibratedAgents. https://calibratedagents.com/blog/hallucination-detection
  • Hallucination detection
  • RAG hallucination
  • LLM-as-a-judge
  • Groundedness
  • AI guardrails
  • RAGTruth++
  • FaithBench

Written by the Calibrated Agents Research Team. Questions or data requests: [email protected]