CalibratedAgents
Try in console
← Blog · Research

Sentence-level hallucination detection (2026): find the wrong sentence, and get better with your own data

By Calibrated Agents Research Team · 9 Oct 2026 · Updated 9 Oct 2026 · 8 min read

RAGTruth++ · sentence level · 200 labelled answers

+6.2points over an LLM judge

Ground79.3
LLM judge73.1
Balanced accuracy on hallucinated sentences · bars start at 65

An AI answer is rarely all wrong: one sentence is. On two human-labelled benchmarks, RAGTruth++ and FaithBench, CalibratedAgents Ground finds the hallucinated sentence 4 to 6 points more accurately than an LLM judge calibrated on the same labels, says what is wrong, why, where and what to do next, and keeps improving as you add your own examples.

The short answer

Sentence-level hallucination detection marks the exact sentence in an AI answer that its source does not support, instead of one verdict for the whole answer. CalibratedAgents Ground returns, for each flagged sentence, what is wrong, why, where the source shows it, and a fix. Tuned on 100–300 of your labelled answers, it beats an equally tuned LLM judge by 4–6 points.

+6.2points over an LLM judge on RAGTruth++, at 200 labelled answers
+5.6points on FaithBench at 300 labelled summaries; the judge does not improve
0.877AUROC when Ground is added to the judge on RAGTruth++, from 0.830
Sentence-level balanced accuracy on FaithBench as labelled examples grow from 50 to 400: Ground rises from 63.5 to about 68, an LLM judge calibrated on the same labels stays near 63.
Figure 1. FaithBench, sentence level. Both systems use the same labelled summaries; the judge stays flat while Ground keeps learning.

What is sentence-level hallucination detection?

It is checking an AI answer one sentence at a time against the documents it should rest on, and flagging only the sentences the documents do not support. The output is a verdict per sentence, plus a reason and the source passage, rather than one score for the whole answer.

Most groundedness and faithfulness checks score the whole answer. That hides the problem teams care about: a long answer that is 90% right can still state one wrong refund amount, one swapped definition or one removed protocol step. As one 2026 guide puts it, a high average groundedness score can hide a much higher rate of individual wrong statements (FutureAGI). And retrieval does not make the problem go away: with RAG, hallucinations shift from outright inventions to subtle mismatches between the evidence and the answer (PatSnap).

A verdict on the whole answer tells you to distrust it. A verdict on the sentence tells you which part to fix, and that is what a reviewer, an agent or a regeneration step can act on.

What does Ground return for each sentence?

For every sentence it flags, Ground returns four things: what is wrong (it contradicts the source, it is not in the source, or it rests on an earlier error), why, where the source shows it, and what next, a corrected sentence.

Here is one sentence from a course answer about TLS 1.3, checked against the course's own material:

Sentence“…data is encrypted with a symmetric cipher such as AES-GCM, which supports keys of up to 512 bits.”
WhatContradicts the source
WhyThe number is different: the source gives AES key sizes of 128, 192 and 256 bits.
Where“AES (Advanced Encryption Standard) is the most widely used symmetric cipher and supports key sizes of 128, 192 and 256 bits.”
What nextCorrect it to what the source says: AES supports keys of 128, 192 or 256 bits.

The answer also gets one overall score, so you can route it: send it, send it to review, or regenerate the flagged sentence with the fix as a hint.

How we tested it: the same labels for everyone

Every system got the same small set of labelled answers from the benchmark and nothing else, then was scored on every answer it had never seen. The comparison is Ground against an LLM judge (Jev, asked about each sentence) calibrated on exactly those labels.

  • Datasets: RAGTruth++ (408 answers, question answering and summaries, 2,334 sentences, every hallucinated span marked by two annotators) and FaithBench (750 summaries from 10 LLMs, 3,582 sentences, expert-labelled spans; its authors report detectors near 50% on it).
  • Labels: 50 to 400 answers, drawn by whole question or source passage, so no question appears on both sides.
  • Repeats: 10 random draws at 50 and 100 labels, 5 at larger sizes; we report the average.
  • Fairness: the judge sets its threshold on the labels; Ground chooses its model and threshold on the same labels, and keeps the judge's verdict unless the labels clearly favour something else. The full rules were written down before the runs.

Results on RAGTruth++

On RAGTruth++, Ground is ahead at every label budget: +3.7 points of balanced accuracy with 50 labelled answers, +4.7 with 100 and +6.2 with 200. The judge stays near 73.5 however many labels it gets; Ground rises to 79.3.

Grouped bars: Ground 77.3, 78.5, 79.3 against an LLM judge 73.6, 73.8, 73.1 at 50, 100 and 200 labelled answers.
Figure 2. Balanced accuracy on hallucinated sentences, RAGTruth++. The axis starts at 65.
Labelled answersGroundLLM judgeLead
5077.373.6+3.7
10078.573.8+4.7
20079.373.1+6.2

Question answering and summaries

The lead comes mostly from question answering, where the judge struggles most. On summaries the two are close at sentence level.

Bars for RAGTruth++ at 100 labels: question answering Ground 78.1 vs judge 73.3; summaries Ground 76.8 vs judge 76.2.
Figure 3. RAGTruth++ split by task, 100 labelled answers per task. The axis starts at 65.

When precision matters most

Tuned for precision instead of balance, with 100 labelled answers Ground's flags were right 75.5% of the time against 69.5% for the judge, while finding 59.6% of hallucinated sentences against 47.9%. Fewer false alarms and more catches at the same time.

Results on FaithBench: it improves with your data

FaithBench is the hardest public benchmark for this task. With 50 labelled summaries Ground and the judge are level; from there Ground keeps improving, to +4.4 at 200 labels and +5.6 at 300, while the judge stays flat at about 63. Most of the gain arrives by 200 to 300 labels.

Line chart: Ground 63.5, 65.8, 67.3, 68.2, 67.6, 67.9 against the judge 63.1, 63.0, 62.9, 62.6, 62.7, 63.8 at 50 to 400 labelled summaries; the lead is written above each point.
Figure 4. Balanced accuracy on hallucinated sentences, FaithBench, at six label budgets. The axis starts at 60.
Labelled summariesGroundLLM judgeLead
5063.563.1+0.4
10065.863.0+2.8
20067.362.9+4.4
30068.262.6+5.6
35067.662.7+4.9
40067.963.8+4.1

Why the judge does not improve: calibrating a judge only moves its threshold, and 50 labels already pin that down. Ground has more to learn from each extra label: which of its signals matter for your definition of a wrong sentence.

Where does the LLM judge miss, and where does Ground help?

An LLM judge is strong at invented content and weak at distortions of what the source does say: a reversed relation, a changed qualifier, a fact moved to the wrong subject. On RAGTruth++ it catches 96% of evident inventions but 68% of subtle conflicts; on FaithBench, 56% of distorted statements.

Horizontal bars: share of hallucinated sentences the LLM judge catches by error type. RAGTruth++: evident baseless info 96%, evident conflict 84%, subtle baseless info 79%, subtle conflict 68%. FaithBench: extrinsic 88%, intrinsic 56%, questionable 47%.
Figure 5. Share of hallucinated sentences the LLM judge catches, by the datasets' own error labels. Red marks the weak spots.

Adding Ground's sentence signals to the judge lifts AUROC from 0.830 to 0.877 on RAGTruth++ and from 0.689 to 0.741 on FaithBench. We also tested five open detectors against the same sentences: Vectara HHEM-2.1-Open, two LettuceDetect models, IBM Granite Guardian 3.3 and a DeBERTa NLI model. None beat the judge on its own; the best added 0.012 AUROC to it, and once Ground was in, none added more than 0.005.

Bars of AUROC: RAGTruth++ judge 0.830, judge plus best open detector 0.842, judge plus Ground 0.877; FaithBench 0.689, 0.701, 0.741.
Figure 6. AUROC at sentence level, five-fold cross-validation grouped by question. The axes do not start at zero.

What about finding the exact words, and whole-answer verdicts?

Locating the exact wrong words follows the sentence decision: Ground's character-level span F1 is level with the judge at small label budgets and up to 3 points ahead with more labels (0.534 against 0.508 on RAGTruth++ at 200; 0.352 against 0.321 on FaithBench at 300). For one yes/no on the whole answer, a calibrated judge is as good as Ground or slightly better.

Whole-answer verdict (balanced accuracy)GroundLLM judge
RAGTruth++, 200 labels78.274.8
FaithBench, 300 labels63.363.9
LLM-AggreFact, 11 datasets, 100 labels each79.780.3
LLM-AggreFact, fresh unseen sample, 50 labels each77.378.9

So use the whole-answer score to route answers, and the sentence verdicts to fix them. On short, single-statement answers, where the sentence is the answer, the two approaches converge.

How do you get these results on your own data?

Start with no labels: Ground works out of the box, and on RAGTruth++ it already matches the best published system. Then label about 200 of your own answers, and Ground is calibrated to your definition of a wrong sentence in minutes.

  • Day one: call the API or paste a case in the console Playground; every answer gets a score and, per sentence, what, why, where and what next.
  • Week one: label 100 to 300 answers your team has already reviewed; that is where the lead over a calibrated judge reached 4 to 6 points in our tests.
  • Ongoing: route flagged answers to your reviewers, and feed their decisions back as labels.

Limitations

  • Two benchmarks with sentence labels, both English and mostly news and web text; your documents may behave differently, which is why calibrating on your own labels matters.
  • The comparison judge is one LLM judge (Jev) asked per sentence; other judges and prompts will give other numbers.
  • Gains are averages over random draws; at 50 labels the spread between draws is several points.
  • On FaithBench, part of the labels are marked questionable by the annotators themselves; no detector does well there.

Frequently asked questions

What is sentence-level hallucination detection?

Checking each sentence of an AI answer against its source documents and flagging only the unsupported or contradicted ones, instead of one verdict for the whole answer. It tells you which part of an answer to fix.

How is it different from LLM-as-a-judge?

An LLM judge gives a verdict and a probability. Ground gives a verdict per sentence with the reason, the source passage and a fix, and it learns from your labels; in our tests a judge calibrated on the same labels stayed flat as labels grew.

How many labelled examples do I need?

None to start. With 100 labelled answers Ground led an equally tuned LLM judge by 4.7 points on RAGTruth++; most of the gain on FaithBench arrived by 200 to 300 labels.

Does it work on summaries as well as question answering?

Yes, with a smaller lead on summaries: on RAGTruth++ at 100 labels, +4.8 points on question answering and +0.6 on summaries. On FaithBench, which is all summaries, the lead grows to 5.6 points at 300 labels.

Which benchmarks did you use?

RAGTruth++ (408 answers re-annotated by two annotators) and FaithBench (750 summaries with expert span labels) at sentence level, and LLM-AggreFact (11 datasets) for whole-answer verdicts.

Do open-source hallucination detectors add anything?

Not on top of Ground in our tests. HHEM-2.1-Open, LettuceDetect, Granite Guardian 3.3 and a DeBERTa NLI model each scored below the LLM judge alone, and none added more than 0.005 AUROC once Ground was combined with the judge.

Can I test it myself?

Yes. Open the console, paste a question, its source documents and an answer, and you get the overall score and the per-sentence what, why, where and what next. The free plan covers 1,000,000 tokens a day.

Sources

Cite this

CalibratedAgents Research (2026). Sentence-level hallucination detection (2026): find the wrong sentence, and get better with your own data. CalibratedAgents. https://calibratedagents.com/blog/sentence-level-hallucination-detection
  • Sentence-level hallucination detection
  • RAGTruth++
  • FaithBench
  • LLM-as-a-judge
  • Groundedness
  • RAG evaluation
  • AI guardrails

Written by the Calibrated Agents Research Team. Questions or data requests: [email protected]