https://api.calibratedagents.com/v1/ground/basicThis is the endpoint for Ground. Send a question, the context your system retrieved, and the answer your model wrote. Ground splits the answer into claims, checks each against the context, and returns every claim with a calibrated score and its receipt. It doesn't take chat messages, a judge prompt or a rubric — just the three fields.
Set up your API key
- Sign in to the console with Google. A workspace is created for you.
- Open API keys, create one, and copy it — it's shown once and starts with
ca_live_.
Set it as an environment variable so the examples below work as they are:
export CA_API_KEY="ca_live_your_key_here"setx CA_API_KEY "ca_live_your_key_here"Keep keys on your server. The API doesn't accept calls from browsers, and a revoked key stops working within a minute.
Choose a call
| Call | Use when | You send | You get |
|---|---|---|---|
POST /v1/ground/basic | Checking an answer before or after it's sent | One question, context and answer | One score; every wrong claim with why, where and the fix |
POST /v1/ground/basic/batch | Checking several answers in one round trip | Up to 50 items | One result per item, in order |
| ⊸ Connect data (console) | Checking answers your AI already gave | A JSON Lines, JSON or CSV file | A job in Review → Jobs, every claim labelled |
Request
Header Authorization: Bearer ca_live_… and a JSON body:
questionstringrequiredWhat the user asked. Non-empty.
contextstring | string[]requiredThe passages your system retrieved for this answer. Send an array to keep passages apart; each claim's receipt points at one of them.
answerstringrequiredYour model's answer — the text being checked. Non-empty.
profile"precision" | "balanced"optionalAccepted for compatibility; it no longer changes the result.
metadataobjectoptionalLabels for your dashboard: environment (string), session_id (string), tags (up to 20 strings), and up to 20 other keys. At most 2,000 characters in total. Don't put personal data here.
Unknown top-level fields are ignored. Empty strings count as missing.
Complete request
curl --max-time 60 -s https://api.calibratedagents.com/v1/ground/basic \
-H "Authorization: Bearer $CA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"question": "Can I prepay my personal loan after six months?",
"context": [
"Prepayment is allowed after 12 EMIs.",
"A foreclosure charge of 4% applies before year three."
],
"answer": "You can prepay after 6 EMIs."
}'import os, requests
r = requests.post(
"https://api.calibratedagents.com/v1/ground/basic",
headers={"Authorization": f"Bearer {os.environ['CA_API_KEY']}"},
json={
"question": "Can I prepay my personal loan after six months?",
"context": ["Prepayment is allowed after 12 EMIs.", "A foreclosure charge of 4% applies before year three."],
"answer": "You can prepay after 6 EMIs.",
},
timeout=60,
)
r.raise_for_status()
result = r.json()
print("Overall score:", result["score"]["p_hallucinated"]) # 0 to 1: higher = more likely unsupported
for s in result["sentences"]:
for c in s["claims"]:
if c["status"] in ("supported", "not_checked"):
continue
print(c["status"], "·", c["text"])
print(" why:", "; ".join(c["checks"]))
print(" fix:", c["fix"])const r = await fetch("https://api.calibratedagents.com/v1/ground/basic", {
method: "POST",
headers: { Authorization: `Bearer ${process.env.CA_API_KEY}`, "Content-Type": "application/json" },
body: JSON.stringify({
question: "Can I prepay my personal loan after six months?",
context: ["Prepayment is allowed after 12 EMIs.", "A foreclosure charge of 4% applies before year three."],
answer: "You can prepay after 6 EMIs.",
}),
});
const result = await r.json();
console.log("Overall score:", result.score.p_hallucinated); // 0 to 1: higher = more likely unsupported
for (const c of result.sentences.flatMap((s) => s.claims)) {
if (c.status === "supported" || c.status === "not_checked") continue;
console.log(c.status, "·", c.text, "\n why:", c.checks.join("; "), "\n fix:", c.fix);
}Response
{
"status": "ok",
"id": "ver_1313eba44472443b8158",
"score": { "p_hallucinated": 0.97, "model": "head-2026-10-07" },
"sentences": [
{
"index": 0,
"span": [0, 28],
"text": "You can prepay after 6 EMIs.",
"status": "contradicted",
"claims": [
{
"id": "a0",
"span": [0, 28],
"text": "You can prepay after 6 EMIs.",
"type": "fact",
"status": "contradicted",
"checks": ["contradicted by its evidence 100%", "material difference (value changed)"],
"evidence": [{ "id": "S1", "text": "Prepayment is allowed after 12 EMIs." }],
"fix": "Correct it to what the source says: “Prepayment is allowed after 12 EMIs.”",
"rests_on": []
}
]
}
],
"corrections": ["You can prepay after 6 EMIs. → Correct it to what the source says: “Prepayment is allowed after 12 EMIs.”"],
"regeneration_hint": "Rewrite the answer using only the source. Fix these claims:\n- You can prepay after 6 EMIs. → …",
"meta": { "latency_ms": 2140 },
"usage": { "tokens": 200, "remaining_today": 999800, "resets_at": "2026-10-08T00:00:00.000Z" }
}statusisokwhen the answer was checked.unavailablemeans nothing was checked — the answer is unverified.score.p_hallucinatedis one calibrated score for the whole answer, from 0 to 1: higher means the answer is more likely to say something the source doesn't support. It'snullif the model isn't loaded.score.modelis the model version, for support.sentences[]is every sentence of the answer, in order, with itsspan(character offsets intoanswer, counted in Unicode code points) and itsstatus— the worst of its claims.sentences[].claims[]are the claims inside each sentence:status(what),checks(why),evidence[](where — the source passage'sidandtext),fix(what next;nullwhen there's nothing to fix),rests_on(the claims it follows from) andtype.correctionsandregeneration_hintappear only when something needs fixing: one line per claim to fix, and a ready-made instruction to send back to the model that wrote the answer.meta.latency_msis how long the check took.usageis what this call cost and what's left today.- New fields may be added. Ignore fields you don't recognise.
| Claim status | Means | What next |
|---|---|---|
contradicted | The source says otherwise. | Correct it from fix. |
not_stated | Nothing in the source supports it. evidence may be empty — that is the finding. | Remove it, or support it. |
depends_on_error | It follows from a wrong claim (rests_on). | Re-check it after fixing that claim. |
unverified_step | A calculation, comparison or conclusion the source can't state directly. | Checking these is the Pro tier — information, not an error. |
supported · not_checked | Found in the source · greetings and framing. | Nothing to do. |
Act on the result
Use the claim statuses to fix the answer, and the overall score to decide what a person should see. A starting point: on a public benchmark of 1,397 labelled answers, a line at 0.42 catches 82% of wrong answers, and 73% of the answers it holds are really wrong. Set your own line once you've labelled some of your answers.
def decide(result, line=0.42):
if result["status"] != "ok":
return "unverified" # nothing was checked — never treat it as a pass
claims = [c for s in result["sentences"] for c in s["claims"]]
if any(c["status"] in ("contradicted", "not_stated") for c in claims):
return "correct_or_hold" # fix each from c["fix"], or send regeneration_hint back to your model
p = result["score"]["p_hallucinated"]
if p is not None and p >= line:
return "hold_for_a_person" # the score sees risk the claim checks didn't pin down
return "send" # nothing to correct, and the score is under your lineBatches
POST /v1/ground/basic/batch takes items (1–50 requests). A top-level profile or metadata applies to every item that doesn't set its own. results come back in the same order, each with its own status; items that couldn't be checked are refunded.
curl --max-time 180 -s https://api.calibratedagents.com/v1/ground/basic/batch \
-H "Authorization: Bearer $CA_API_KEY" -H "Content-Type: application/json" \
-d '{
"profile": "precision",
"metadata": { "environment": "production" },
"items": [
{ "question": "…", "context": "…", "answer": "…", "metadata": { "session_id": "conv-7" } },
{ "question": "…", "context": ["…", "…"], "answer": "…" }
]
}'Limits
- Daily allowance per workspace, in tokens: UTF-8 bytes ÷ 4, and at least 200 per checked answer. Free plan: 1,000,000 a day, renewing at 00:00 UTC.
- Per minute: 60 requests per key on the free plan.
- Per request: 1 MB, and 15,000 tokens per answer (about 60,000 bytes of question, context and answer) on the free plan.
- Batches: 1–50 items. Metadata: 2,000 characters.
- Every response carries
X-RateLimit-Limit,X-RateLimit-RemainingandX-RateLimit-Reset. Failed checks are never charged. - Typical latency is two to four seconds per answer. Use a read timeout of at least 60 seconds, and 180 for batches.
Migrating from an LLM judge
If you check answers today by asking another model to grade them, the move is mostly mechanical — but one thing flips.
| LLM judge | Ground |
|---|---|
| A judge prompt with a rubric | No prompt: send question, context, answer |
| One grade for the whole answer | One calibrated score.p_hallucinated for the answer, and a status, reasons and a fix for every claim |
| Higher grade = better | Higher score = more likely wrong |
| An explanation in prose | Where and what next: evidence[].text and fix |
| Uncalibrated (a 7/10 means different things on different days) | A calibrated score you can set your own line on |
| Often 10–30 seconds | Typically 2–4 seconds |
- Map the fields. The text you sent the judge as context becomes
context; the answer under test becomesanswer. - Invert your comparisons. Where you passed answers with a high grade, pass answers with a low score.
- Act per claim. Correct the wrong claim from its
fix, not the whole answer. - Treat failures as unverified. A
502orstatus: "unavailable"is not a pass. - Revalidate thresholds on representative answers before switching traffic.
score runs the other way and is a probability. Start from the bands above and tune them.Errors
Errors are JSON: {"error": {"code": "…", "message": "…"}}. Branch on code; the message is written for people.
| HTTP | code | Likely cause | Action |
|---|---|---|---|
| 401 | invalid_api_key | Missing, malformed, wrong or revoked key | Send Authorization: Bearer ca_live_… with an active key. Don't retry. |
| 403 | ip_not_allowed | The key is limited to your company's IP addresses | Call from an allowed address, or change the key's allowlist in the console. Don't retry. |
| 404 | not_found | Unknown route | Use /v1/ground/basic or /v1/ground/basic/batch. |
| 405 | method_not_allowed | Wrong HTTP method | Use POST (GET for /v1/health). |
| 413 | request_too_large | Body over 1 MB, or an item over the plan's per-request token limit | Shorten the context or split the answer. Don't retry unchanged. |
| 422 | invalid_request | A required field is missing or empty, a value has the wrong type, or a batch has more than 50 items | Read error.message — it names the field. Don't retry unchanged. |
| 429 | rate_limited | Too many requests this minute for this key | Wait for Retry-After, then retry with backoff. |
| 429 | daily_quota_exceeded | Today's allowance is used up | Wait until X-RateLimit-Reset (00:00 UTC), or ask for more volume in the console. |
| 502 | verification_unavailable | The checking service couldn't run. Nothing was charged. | Retry with bounded backoff. Treat the answer as unverified — never as a pass. |
| 503 | auth_unavailable | We couldn't verify the key right now | Retry with bounded backoff. |
Retry only 429, 502 and 503, with bounded backoff — for example three attempts at 2, 4 and 8 seconds — honouring Retry-After when present. Apply the same budget to network timeouts. Don't retry 401, 403, 404, 405, 413 or 422 without fixing the request.
Your data
- By default we keep numbers only: scores, decisions and offsets — not the text of your questions, contexts or answers.
- Text is kept only if your workspace turns it on (Settings → Data & security), for the retention period you choose — or, for Connect data jobs, only for wrong or unsure claims, so your team can review them.
- To check answers your AI already gave, use ⊸ Connect data in the console. If you prepare the file yourself, one line per answer:
{"i": 0, "q": "Can I prepay after six months?", "a": "You can prepay after 6 EMIs.", "ctx": [{"text": "Prepayment is allowed after 12 EMIs.", "source": "Loan agreement §4.2", "date": "2025-04-01"}], "t": "2026-09-24T14:02:00Z"}Only the matched fields of the sampled rows leave your computer, and customer or user ids are never uploaded. See Privacy.