Hallucination detection in 2026: find the wrong sentence, not just the wrong answer
By Calibrated Agents Research Team · 9 Oct 2026 · Updated 9 Oct 2026 · 9 min read
RAGTruth++ · finding the hallucinated sentence
0.877AUROC for Ground, ahead of every system tested
An AI answer is rarely all wrong. Usually one or two sentences are, and an answer-level hallucination detection check can't tell you which. We measured that gap on 5,916 human-labelled sentences from two public benchmarks, and how far checking each sentence closes it.
Why answer-level hallucination detection isn't enough
Because a hallucinated answer is mostly correct. On RAGTruth++, a flagged answer has about two wrong sentences in five or six; on FaithBench, under two in five. A single verdict for the whole answer can only block it, send all of it to review, or regenerate it blind.
Most hallucination detection tools return one groundedness or faithfulness score per answer. That worked when answers were one line. It stops working when a RAG assistant or an agent writes a paragraph from a policy document and the only mistake is a refund percentage in the third sentence. The same five-sentence answer looks like this under each kind of check:
Three things go wrong with the left-hand card:
- You throw away good answers. Block every answer with a hallucination and you discard 7 in 10 answers on these benchmarks, most of whose content was right.
- Your reviewers do the detective work. A flagged answer tells them to distrust it, not where to look. They reread the whole answer and the whole source to find one or two sentences.
- Regeneration is blind. Ask the model to try again without saying what was wrong, and it can repeat the error or introduce a new one.
Retrieval does not make this go away. With RAG, hallucinations shift from obvious inventions to subtle mismatches between the evidence and the answer [7], and a high average groundedness score can hide a much higher rate of individual wrong statements [6]. The unit that matters is the sentence.
A flag on the whole answer tells you to distrust it. A flag on the sentence tells you what to fix.The case for sentence-level hallucination detection

What sentence-level hallucination detection gives you
For every sentence its source doesn't support, Ground returns four things: what is wrong (it contradicts the source, the source doesn't say it, or it rests on an earlier error), why, where the source shows it, and what next: the corrected sentence. The whole answer also gets one score, so you can route it.
Here is what that looks like for one sentence of an AI tutor's answer about TLS 1.3, checked against the course material:
That changes what each team can do with a flag. A reviewer reads one sentence and one passage instead of the whole answer. An agent regenerates one sentence with the fix as a hint. A product team keeps the four good sentences and corrects the fifth.
Benchmark: Ground vs LLM-as-a-judge vs open-source hallucination detectors
On two human-labelled benchmarks, Ground finds hallucinated sentences better than every system we tested: an LLM judge asked about each sentence, and five open-source detectors, including Vectara HHEM, LettuceDetect and IBM Granite Guardian. Ground scores 0.877 AUROC on RAGTruth++ and 0.741 on FaithBench.

| Hallucination detector | RAGTruth++ | FaithBench | Type |
|---|---|---|---|
| CalibratedAgents Ground | 0.877 | 0.741 | Sentence-level verification |
| LLM-as-a-judge (Jev, per sentence) | 0.830 | 0.701 | LLM judge |
| LettuceDetect large (ModernBERT) | 0.743 | 0.675 | Open source, token-level |
| LettuceDetect v2 (mmBERT) | 0.740 | 0.686 | Open source, token-level |
| Vectara HHEM-2.1-Open | 0.718 | 0.615 | Open source, consistency score |
| DeBERTa-v3 NLI | 0.660 | 0.559 | Open source, entailment |
| IBM Granite Guardian 3.3 (groundedness) | 0.653 | 0.570 | Open source, guardian model |
Two notes for reading Figure 2. FaithBench is the hardest public benchmark for this task; its authors report that the best detectors sit near 50% accuracy on it [2], so every score there is low. And Granite Guardian is designed to judge a whole response in context; scored one sentence at a time, it over-flags.
When precision matters most
Tuned for precision on 100 labelled answers from RAGTruth++, Ground's flags were right 75.5% of the time against 69.5% for an LLM judge tuned on the same answers, while catching 59.6% of hallucinated sentences against 47.9%. Fewer false alarms and more catches, at the same time.
Ground gets better with your data. An LLM judge doesn't.
Give both systems the same labelled answers from your own use case, and the gap widens as you add labels. On FaithBench, Ground goes from level with an LLM judge at 50 labels to +5.6 points at 300; on RAGTruth++, from +3.7 at 50 labels to +6.2 at 200. The judge barely moves.

| Labelled answers | FaithBench: Ground | FaithBench: LLM judge | RAGTruth++: Ground | RAGTruth++: LLM judge |
|---|---|---|---|---|
| 50 | 63.5 | 63.1 | 77.3 | 73.6 |
| 100 | 65.8 | 63.0 | 78.5 | 73.8 |
| 200 | 67.3 | 62.9 | 79.3 | 73.1 |
| 300 | 68.2 | 62.6 | — | — |
| 400 | 67.9 | 63.8 | — | — |
Why: calibrating a judge only moves its threshold, and 50 labels already pin that down. Ground learns which of its signals matter for your definition of a wrong sentence, so every extra label your team has already reviewed makes it sharper. Most of the gain arrives by 200 to 300 labels.
Where LLM-as-a-judge misses hallucinations
LLM judges are good at spotting invented content and weak at distortions of what the source does say: a reversed relation, a changed qualifier, a fact moved to the wrong subject. On FaithBench, a judge catches 88% of added information but only 56% of distorted facts.

These are the hallucinations that survive a quick read: sentences built from the source's own words, slightly bent. They are also the ones that cost the most in regulated work, where the policy says 50% and the answer says full refund. Checking each sentence against the exact passage it should rest on is how you catch them.
How to add sentence-level hallucination detection to your RAG app or agent
Send each answer with its question and source documents to the Ground API, or paste a case into the console. You get the overall score and, for every sentence that needs it, what, why, where and what next. Then label 100 to 300 of your own answers to calibrate it to your domain.
- Today: sign in to the CalibratedAgents console with Google and try your own case in the Playground. The free plan covers 1,000,000 tokens a day.
- This week: send us 20 real conversations with their source documents, and we'll send back what was wrong in each, why, where, and the fix.
- This month: become a design partner. Label a few hundred answers your team has already reviewed, and Ground is calibrated to your own definition of a wrong answer.
How we measured
What this does not prove
- Two English benchmarks, mostly news and web text. Your documents may behave differently, which is why calibrating on your own labels matters.
- One LLM judge and one prompt. Other judges will give other numbers, though we expect the shape (flat with more labels) to hold.
- Averages over random draws; at 50 labels the spread between draws is several points.
- Part of FaithBench's labels are marked questionable by the annotators themselves; no detector does well there.
Frequently asked questions
What is hallucination detection?
Hallucination detection checks whether an AI answer contains statements that its source documents don't support or that contradict them. Sentence-level hallucination detection goes further: it marks which sentence is wrong, and why.
How do you detect hallucinations in a RAG application?
Check each sentence of the generated answer against the retrieved passages it should rest on. Flag sentences that contradict the passages or that the passages don't mention, and return the passage and a fix so the answer can be reviewed or regenerated.
Is LLM-as-a-judge good enough for hallucination detection?
For a quick yes/no on short answers, a well-calibrated LLM judge is a strong baseline. For long answers it misses distortions of the source and can't learn from your labels: in our benchmark it caught 56% of distorted facts, and tuning it on more labels didn't improve it.
What is the best hallucination detection model in 2026?
On our sentence-level benchmark, CalibratedAgents Ground scored highest on both RAGTruth++ (0.877 AUROC) and FaithBench (0.741), ahead of an LLM judge and of the open-source detectors HHEM-2.1-Open, LettuceDetect, Granite Guardian 3.3 and DeBERTa NLI.
Do I need labelled data to start?
No. Ground works out of the box. Labelling 100 to 300 answers your team has already reviewed calibrates it to your domain, which is where its lead over an LLM judge reached 4 to 6 points.
What does Ground return for a hallucinated sentence?
What is wrong (it contradicts the source, the source doesn't say it, or it rests on an earlier error), why, where the source shows it, and what next: the corrected sentence. The whole answer also gets one score.
Can I try it for free?
Yes. Sign in to the console and check your own cases in the Playground or through the API. The free plan covers 1,000,000 tokens a day.
References
- RAGTruth++ dataset. Blue Guardrails, 2026.
- Tamber et al. FaithBench: a diverse hallucination benchmark for summarization by modern LLMs. NAACL 2025.
- Jev (TypeSafe AI), a typed decision model, used here as an LLM judge asked about each sentence.
- Tang, Laban, Durrett. MiniCheck: efficient fact-checking of LLMs on grounding documents (introduces LLM-AggreFact). EMNLP 2024.
- Kovács, Recski. LettuceDetect: a hallucination detection framework for RAG applications. 2025–2026.
- FutureAGI. LLM hallucination: a 2026 architectural deep dive.
- PatSnap. LLM hallucination rate evaluation for engineering. 2026.
- Vectara. HHEM-2.1-Open model card.
- IBM. Granite Guardian documentation.
Cite this
CalibratedAgents Research (2026). Hallucination detection in 2026: find the wrong sentence, not just the wrong answer. CalibratedAgents. https://calibratedagents.com/blog/hallucination-detection