Sentence-level hallucination detection (2026): find the wrong sentence, and get better with your own data
By Calibrated Agents Research Team · 9 Oct 2026 · Updated 9 Oct 2026 · 8 min read
RAGTruth++ · sentence level · 200 labelled answers
+6.2points over an LLM judge
An AI answer is rarely all wrong: one sentence is. On two human-labelled benchmarks, RAGTruth++ and FaithBench, CalibratedAgents Ground finds the hallucinated sentence 4 to 6 points more accurately than an LLM judge calibrated on the same labels, says what is wrong, why, where and what to do next, and keeps improving as you add your own examples.
The short answer
Sentence-level hallucination detection marks the exact sentence in an AI answer that its source does not support, instead of one verdict for the whole answer. CalibratedAgents Ground returns, for each flagged sentence, what is wrong, why, where the source shows it, and a fix. Tuned on 100–300 of your labelled answers, it beats an equally tuned LLM judge by 4–6 points.

What is sentence-level hallucination detection?
It is checking an AI answer one sentence at a time against the documents it should rest on, and flagging only the sentences the documents do not support. The output is a verdict per sentence, plus a reason and the source passage, rather than one score for the whole answer.
Most groundedness and faithfulness checks score the whole answer. That hides the problem teams care about: a long answer that is 90% right can still state one wrong refund amount, one swapped definition or one removed protocol step. As one 2026 guide puts it, a high average groundedness score can hide a much higher rate of individual wrong statements (FutureAGI). And retrieval does not make the problem go away: with RAG, hallucinations shift from outright inventions to subtle mismatches between the evidence and the answer (PatSnap).
A verdict on the whole answer tells you to distrust it. A verdict on the sentence tells you which part to fix, and that is what a reviewer, an agent or a regeneration step can act on.
What does Ground return for each sentence?
For every sentence it flags, Ground returns four things: what is wrong (it contradicts the source, it is not in the source, or it rests on an earlier error), why, where the source shows it, and what next, a corrected sentence.
Here is one sentence from a course answer about TLS 1.3, checked against the course's own material:
The answer also gets one overall score, so you can route it: send it, send it to review, or regenerate the flagged sentence with the fix as a hint.
How we tested it: the same labels for everyone
Every system got the same small set of labelled answers from the benchmark and nothing else, then was scored on every answer it had never seen. The comparison is Ground against an LLM judge (Jev, asked about each sentence) calibrated on exactly those labels.
- Datasets: RAGTruth++ (408 answers, question answering and summaries, 2,334 sentences, every hallucinated span marked by two annotators) and FaithBench (750 summaries from 10 LLMs, 3,582 sentences, expert-labelled spans; its authors report detectors near 50% on it).
- Labels: 50 to 400 answers, drawn by whole question or source passage, so no question appears on both sides.
- Repeats: 10 random draws at 50 and 100 labels, 5 at larger sizes; we report the average.
- Fairness: the judge sets its threshold on the labels; Ground chooses its model and threshold on the same labels, and keeps the judge's verdict unless the labels clearly favour something else. The full rules were written down before the runs.
Results on RAGTruth++
On RAGTruth++, Ground is ahead at every label budget: +3.7 points of balanced accuracy with 50 labelled answers, +4.7 with 100 and +6.2 with 200. The judge stays near 73.5 however many labels it gets; Ground rises to 79.3.

| Labelled answers | Ground | LLM judge | Lead |
|---|---|---|---|
| 50 | 77.3 | 73.6 | +3.7 |
| 100 | 78.5 | 73.8 | +4.7 |
| 200 | 79.3 | 73.1 | +6.2 |
Question answering and summaries
The lead comes mostly from question answering, where the judge struggles most. On summaries the two are close at sentence level.

When precision matters most
Tuned for precision instead of balance, with 100 labelled answers Ground's flags were right 75.5% of the time against 69.5% for the judge, while finding 59.6% of hallucinated sentences against 47.9%. Fewer false alarms and more catches at the same time.
Results on FaithBench: it improves with your data
FaithBench is the hardest public benchmark for this task. With 50 labelled summaries Ground and the judge are level; from there Ground keeps improving, to +4.4 at 200 labels and +5.6 at 300, while the judge stays flat at about 63. Most of the gain arrives by 200 to 300 labels.

| Labelled summaries | Ground | LLM judge | Lead |
|---|---|---|---|
| 50 | 63.5 | 63.1 | +0.4 |
| 100 | 65.8 | 63.0 | +2.8 |
| 200 | 67.3 | 62.9 | +4.4 |
| 300 | 68.2 | 62.6 | +5.6 |
| 350 | 67.6 | 62.7 | +4.9 |
| 400 | 67.9 | 63.8 | +4.1 |
Why the judge does not improve: calibrating a judge only moves its threshold, and 50 labels already pin that down. Ground has more to learn from each extra label: which of its signals matter for your definition of a wrong sentence.
Where does the LLM judge miss, and where does Ground help?
An LLM judge is strong at invented content and weak at distortions of what the source does say: a reversed relation, a changed qualifier, a fact moved to the wrong subject. On RAGTruth++ it catches 96% of evident inventions but 68% of subtle conflicts; on FaithBench, 56% of distorted statements.

Adding Ground's sentence signals to the judge lifts AUROC from 0.830 to 0.877 on RAGTruth++ and from 0.689 to 0.741 on FaithBench. We also tested five open detectors against the same sentences: Vectara HHEM-2.1-Open, two LettuceDetect models, IBM Granite Guardian 3.3 and a DeBERTa NLI model. None beat the judge on its own; the best added 0.012 AUROC to it, and once Ground was in, none added more than 0.005.

What about finding the exact words, and whole-answer verdicts?
Locating the exact wrong words follows the sentence decision: Ground's character-level span F1 is level with the judge at small label budgets and up to 3 points ahead with more labels (0.534 against 0.508 on RAGTruth++ at 200; 0.352 against 0.321 on FaithBench at 300). For one yes/no on the whole answer, a calibrated judge is as good as Ground or slightly better.
| Whole-answer verdict (balanced accuracy) | Ground | LLM judge |
|---|---|---|
| RAGTruth++, 200 labels | 78.2 | 74.8 |
| FaithBench, 300 labels | 63.3 | 63.9 |
| LLM-AggreFact, 11 datasets, 100 labels each | 79.7 | 80.3 |
| LLM-AggreFact, fresh unseen sample, 50 labels each | 77.3 | 78.9 |
So use the whole-answer score to route answers, and the sentence verdicts to fix them. On short, single-statement answers, where the sentence is the answer, the two approaches converge.
How do you get these results on your own data?
Start with no labels: Ground works out of the box, and on RAGTruth++ it already matches the best published system. Then label about 200 of your own answers, and Ground is calibrated to your definition of a wrong sentence in minutes.
- Day one: call the API or paste a case in the console Playground; every answer gets a score and, per sentence, what, why, where and what next.
- Week one: label 100 to 300 answers your team has already reviewed; that is where the lead over a calibrated judge reached 4 to 6 points in our tests.
- Ongoing: route flagged answers to your reviewers, and feed their decisions back as labels.
Limitations
- Two benchmarks with sentence labels, both English and mostly news and web text; your documents may behave differently, which is why calibrating on your own labels matters.
- The comparison judge is one LLM judge (Jev) asked per sentence; other judges and prompts will give other numbers.
- Gains are averages over random draws; at 50 labels the spread between draws is several points.
- On FaithBench, part of the labels are marked questionable by the annotators themselves; no detector does well there.
Frequently asked questions
What is sentence-level hallucination detection?
Checking each sentence of an AI answer against its source documents and flagging only the unsupported or contradicted ones, instead of one verdict for the whole answer. It tells you which part of an answer to fix.
How is it different from LLM-as-a-judge?
An LLM judge gives a verdict and a probability. Ground gives a verdict per sentence with the reason, the source passage and a fix, and it learns from your labels; in our tests a judge calibrated on the same labels stayed flat as labels grew.
How many labelled examples do I need?
None to start. With 100 labelled answers Ground led an equally tuned LLM judge by 4.7 points on RAGTruth++; most of the gain on FaithBench arrived by 200 to 300 labels.
Does it work on summaries as well as question answering?
Yes, with a smaller lead on summaries: on RAGTruth++ at 100 labels, +4.8 points on question answering and +0.6 on summaries. On FaithBench, which is all summaries, the lead grows to 5.6 points at 300 labels.
Which benchmarks did you use?
RAGTruth++ (408 answers re-annotated by two annotators) and FaithBench (750 summaries with expert span labels) at sentence level, and LLM-AggreFact (11 datasets) for whole-answer verdicts.
Do open-source hallucination detectors add anything?
Not on top of Ground in our tests. HHEM-2.1-Open, LettuceDetect, Granite Guardian 3.3 and a DeBERTa NLI model each scored below the LLM judge alone, and none added more than 0.005 AUROC once Ground was combined with the judge.
Can I test it myself?
Yes. Open the console, paste a question, its source documents and an answer, and you get the overall score and the per-sentence what, why, where and what next. The free plan covers 1,000,000 tokens a day.
Sources
- RAGTruth++ dataset (Blue Guardrails)
- FaithBench (Tamber et al., NAACL 2025)
- LLM-AggreFact and MiniCheck (Tang et al., EMNLP 2024)
- Vectara HHEM-2.1-Open
- LettuceDetect (KR Labs)
- IBM Granite Guardian
- LLM hallucination: a 2026 architectural deep dive (FutureAGI)
- LLM hallucination rate evaluation (PatSnap)
Cite this
CalibratedAgents Research (2026). Sentence-level hallucination detection (2026): find the wrong sentence, and get better with your own data. CalibratedAgents. https://calibratedagents.com/blog/sentence-level-hallucination-detection