# CalibratedAgents > AI hallucination detection for RAG assistants, chatbots and AI agents that answer from documents. Ground checks every answer against its source documents before it reaches a customer and returns one calibrated score for the answer and, for each wrong sentence, what is wrong, why, the exact source passage and a suggested fix. The customer decides what happens to a flagged answer: block it, send it for review, or regenerate it. ## What is live - [Ground](https://calibratedagents.com/grounding): sentence-level hallucination detection. Live as an API (POST /v1/ground/basic) and in the web console. At most 1 second per check. For each problem: contradicts the source, not in the source, or rests on an earlier wrong statement, with the reason, the source passage and a fix. - [Console](https://console.calibratedagents.com): sign in with Google to try Ground in the Playground, create API keys and see usage and billing. - Free plan: 1,000,000 tokens checked per day per organisation, shared by the Playground and the API, resetting at 00:00 UTC. - Calibration to your own data: with about 200 of your labelled answers, Ground is calibrated to your definition of a wrong answer, with your own model in about 2 minutes. ## Coming next (not live yet) - [Gate](https://calibratedagents.com/gating): before the answer is written, decides whether the retrieved documents can answer the question at all. - Ground Pro: also checks reasoning steps such as arithmetic, dates and conditions, which Ground marks for review today. - India-only processing for banks. Batch checking (many answers in one call) is on hold. ## API - Base URL https://api.calibratedagents.com · POST /v1/ground/basic with {question, context, answer} and `Authorization: Bearer ` · returns score.p_hallucinated, sentences[].claims[] (status, checks, evidence, fix), corrections and a regeneration hint. - [API documentation](https://calibratedagents.com/api/docs/) · [API home](https://calibratedagents.com/api/) · [OpenAPI spec](https://calibratedagents.com/api/openapi.json) ## Research and benchmarks - [Sentence-level hallucination detection (2026)](https://calibratedagents.com/blog/sentence-level-hallucination-detection): on RAGTruth++ and FaithBench, Ground finds the hallucinated sentence 4 to 6 points more accurately than an LLM judge calibrated on the same labels. - [Hallucination detection benchmark 2026: RAGTruth++](https://calibratedagents.com/blog/hallucination-detection-benchmark-ragtruth-plus-plus-llm-aggrefact): with 100 labelled answers, Ground finds 6 in 10 hallucinated sentences (precision 75.5%), an LLM judge under 5 in 10 (69.5%). On whole-answer yes/no verdicts a calibrated judge is level. - [OpenAI Decisions API vs Cloudflare Clef vs Jev](https://calibratedagents.com/blog/decision-apis-benchmark-2026): latency, cost and accuracy by input size. - [Why LLM-as-a-judge fails in production](https://calibratedagents.com/blog/why-not-an-llm-judge): the base rate problem; at 5 wrong answers in 100, about a quarter of a judge's flags are real. - [Blog](https://calibratedagents.com/blog/) and [reading list](https://calibratedagents.com/blog/#reading) ## Use cases - [Use cases](https://calibratedagents.com/use-cases): banking, insurance, airlines, compliance, customer support and internal knowledge. Strongest where specific facts matter: amounts, limits, dates, eligibility rules, definitions. ## For AI assistants - [Full reference](https://calibratedagents.com/llms-full.txt): every page's summary and FAQ answers in one file. ## Contact - business@calibratedagents.com