CalibratedAgents
Try in console

API documentation

Check an answer with /v1/ground/basic

Authenticate, send the question, the passages your system retrieved and your model's answer, and get one calibrated score for the answer and, for every wrong claim, why, where and the fix.

POSThttps://api.calibratedagents.com/v1/ground/basic

This is the endpoint for Ground. Send a question, the context your system retrieved, and the answer your model wrote. Ground splits the answer into claims, checks each against the context, and returns every claim with a calibrated score and its receipt. It doesn't take chat messages, a judge prompt or a rubric — just the three fields.

Set up your API key

  1. Sign in to the console with Google. A workspace is created for you.
  2. Open API keys, create one, and copy it — it's shown once and starts with ca_live_.

Set it as an environment variable so the examples below work as they are:

macOS / Linux
export CA_API_KEY="ca_live_your_key_here"
Windows PowerShell
setx CA_API_KEY "ca_live_your_key_here"

Keep keys on your server. The API doesn't accept calls from browsers, and a revoked key stops working within a minute.

Choose a call

CallUse whenYou sendYou get
POST /v1/ground/basicChecking an answer before or after it's sentOne question, context and answerOne score; every wrong claim with why, where and the fix
POST /v1/ground/basic/batchChecking several answers in one round tripUp to 50 itemsOne result per item, in order
⊸ Connect data (console)Checking answers your AI already gaveA JSON Lines, JSON or CSV fileA job in Review → Jobs, every claim labelled

Request

Header Authorization: Bearer ca_live_… and a JSON body:

questionstringrequired

What the user asked. Non-empty.

contextstring | string[]required

The passages your system retrieved for this answer. Send an array to keep passages apart; each claim's receipt points at one of them.

answerstringrequired

Your model's answer — the text being checked. Non-empty.

profile"precision" | "balanced"optional

Accepted for compatibility; it no longer changes the result.

metadataobjectoptional

Labels for your dashboard: environment (string), session_id (string), tags (up to 20 strings), and up to 20 other keys. At most 2,000 characters in total. Don't put personal data here.

Unknown top-level fields are ignored. Empty strings count as missing.

Complete request

curl
curl --max-time 60 -s https://api.calibratedagents.com/v1/ground/basic \
  -H "Authorization: Bearer $CA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "question": "Can I prepay my personal loan after six months?",
    "context": [
      "Prepayment is allowed after 12 EMIs.",
      "A foreclosure charge of 4% applies before year three."
    ],
    "answer": "You can prepay after 6 EMIs."
  }'
Python
import os, requests

r = requests.post(
    "https://api.calibratedagents.com/v1/ground/basic",
    headers={"Authorization": f"Bearer {os.environ['CA_API_KEY']}"},
    json={
        "question": "Can I prepay my personal loan after six months?",
        "context": ["Prepayment is allowed after 12 EMIs.", "A foreclosure charge of 4% applies before year three."],
        "answer": "You can prepay after 6 EMIs.",
    },
    timeout=60,
)
r.raise_for_status()
result = r.json()

print("Overall score:", result["score"]["p_hallucinated"])   # 0 to 1: higher = more likely unsupported
for s in result["sentences"]:
    for c in s["claims"]:
        if c["status"] in ("supported", "not_checked"):
            continue
        print(c["status"], "·", c["text"])
        print("  why:", "; ".join(c["checks"]))
        print("  fix:", c["fix"])
JavaScript
const r = await fetch("https://api.calibratedagents.com/v1/ground/basic", {
  method: "POST",
  headers: { Authorization: `Bearer ${process.env.CA_API_KEY}`, "Content-Type": "application/json" },
  body: JSON.stringify({
    question: "Can I prepay my personal loan after six months?",
    context: ["Prepayment is allowed after 12 EMIs.", "A foreclosure charge of 4% applies before year three."],
    answer: "You can prepay after 6 EMIs.",
  }),
});
const result = await r.json();

console.log("Overall score:", result.score.p_hallucinated);   // 0 to 1: higher = more likely unsupported
for (const c of result.sentences.flatMap((s) => s.claims)) {
  if (c.status === "supported" || c.status === "not_checked") continue;
  console.log(c.status, "·", c.text, "\n  why:", c.checks.join("; "), "\n  fix:", c.fix);
}

Response

JSON
{
  "status": "ok",
  "id": "ver_1313eba44472443b8158",
  "score": { "p_hallucinated": 0.97, "model": "head-2026-10-07" },
  "sentences": [
    {
      "index": 0,
      "span": [0, 28],
      "text": "You can prepay after 6 EMIs.",
      "status": "contradicted",
      "claims": [
        {
          "id": "a0",
          "span": [0, 28],
          "text": "You can prepay after 6 EMIs.",
          "type": "fact",
          "status": "contradicted",
          "checks": ["contradicted by its evidence 100%", "material difference (value changed)"],
          "evidence": [{ "id": "S1", "text": "Prepayment is allowed after 12 EMIs." }],
          "fix": "Correct it to what the source says: “Prepayment is allowed after 12 EMIs.”",
          "rests_on": []
        }
      ]
    }
  ],
  "corrections": ["You can prepay after 6 EMIs. → Correct it to what the source says: “Prepayment is allowed after 12 EMIs.”"],
  "regeneration_hint": "Rewrite the answer using only the source. Fix these claims:\n- You can prepay after 6 EMIs. → …",
  "meta": { "latency_ms": 2140 },
  "usage": { "tokens": 200, "remaining_today": 999800, "resets_at": "2026-10-08T00:00:00.000Z" }
}
  • status is ok when the answer was checked. unavailable means nothing was checked — the answer is unverified.
  • score.p_hallucinated is one calibrated score for the whole answer, from 0 to 1: higher means the answer is more likely to say something the source doesn't support. It's null if the model isn't loaded. score.model is the model version, for support.
  • sentences[] is every sentence of the answer, in order, with its span (character offsets into answer, counted in Unicode code points) and its status — the worst of its claims.
  • sentences[].claims[] are the claims inside each sentence: status (what), checks (why), evidence[] (where — the source passage's id and text), fix (what next; null when there's nothing to fix), rests_on (the claims it follows from) and type.
  • corrections and regeneration_hint appear only when something needs fixing: one line per claim to fix, and a ready-made instruction to send back to the model that wrote the answer.
  • meta.latency_ms is how long the check took. usage is what this call cost and what's left today.
  • New fields may be added. Ignore fields you don't recognise.
Claim statusMeansWhat next
contradictedThe source says otherwise.Correct it from fix.
not_statedNothing in the source supports it. evidence may be empty — that is the finding.Remove it, or support it.
depends_on_errorIt follows from a wrong claim (rests_on).Re-check it after fixing that claim.
unverified_stepA calculation, comparison or conclusion the source can't state directly.Checking these is the Pro tier — information, not an error.
supported · not_checkedFound in the source · greetings and framing.Nothing to do.

Act on the result

Use the claim statuses to fix the answer, and the overall score to decide what a person should see. A starting point: on a public benchmark of 1,397 labelled answers, a line at 0.42 catches 82% of wrong answers, and 73% of the answers it holds are really wrong. Set your own line once you've labelled some of your answers.

Python
def decide(result, line=0.42):
    if result["status"] != "ok":
        return "unverified"                    # nothing was checked — never treat it as a pass
    claims = [c for s in result["sentences"] for c in s["claims"]]
    if any(c["status"] in ("contradicted", "not_stated") for c in claims):
        return "correct_or_hold"               # fix each from c["fix"], or send regeneration_hint back to your model
    p = result["score"]["p_hallucinated"]
    if p is not None and p >= line:
        return "hold_for_a_person"             # the score sees risk the claim checks didn't pin down
    return "send"                             # nothing to correct, and the score is under your line

Batches

POST /v1/ground/basic/batch takes items (1–50 requests). A top-level profile or metadata applies to every item that doesn't set its own. results come back in the same order, each with its own status; items that couldn't be checked are refunded.

curl
curl --max-time 180 -s https://api.calibratedagents.com/v1/ground/basic/batch \
  -H "Authorization: Bearer $CA_API_KEY" -H "Content-Type: application/json" \
  -d '{
    "profile": "precision",
    "metadata": { "environment": "production" },
    "items": [
      { "question": "…", "context": "…", "answer": "…", "metadata": { "session_id": "conv-7" } },
      { "question": "…", "context": ["…", "…"], "answer": "…" }
    ]
  }'

Limits

  • Daily allowance per workspace, in tokens: UTF-8 bytes ÷ 4, and at least 200 per checked answer. Free plan: 1,000,000 a day, renewing at 00:00 UTC.
  • Per minute: 60 requests per key on the free plan.
  • Per request: 1 MB, and 15,000 tokens per answer (about 60,000 bytes of question, context and answer) on the free plan.
  • Batches: 1–50 items. Metadata: 2,000 characters.
  • Every response carries X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset. Failed checks are never charged.
  • Typical latency is two to four seconds per answer. Use a read timeout of at least 60 seconds, and 180 for batches.

Migrating from an LLM judge

If you check answers today by asking another model to grade them, the move is mostly mechanical — but one thing flips.

LLM judgeGround
A judge prompt with a rubricNo prompt: send question, context, answer
One grade for the whole answerOne calibrated score.p_hallucinated for the answer, and a status, reasons and a fix for every claim
Higher grade = betterHigher score = more likely wrong
An explanation in proseWhere and what next: evidence[].text and fix
Uncalibrated (a 7/10 means different things on different days)A calibrated score you can set your own line on
Often 10–30 secondsTypically 2–4 seconds
  1. Map the fields. The text you sent the judge as context becomes context; the answer under test becomes answer.
  2. Invert your comparisons. Where you passed answers with a high grade, pass answers with a low score.
  3. Act per claim. Correct the wrong claim from its fix, not the whole answer.
  4. Treat failures as unverified. A 502 or status: "unavailable" is not a pass.
  5. Revalidate thresholds on representative answers before switching traffic.
Don't copy a judge's pass threshold. A judge's “7 out of 10 is fine” has no counterpart here: Ground's score runs the other way and is a probability. Start from the bands above and tune them.

Errors

Errors are JSON: {"error": {"code": "…", "message": "…"}}. Branch on code; the message is written for people.

HTTPcodeLikely causeAction
401invalid_api_keyMissing, malformed, wrong or revoked keySend Authorization: Bearer ca_live_… with an active key. Don't retry.
403ip_not_allowedThe key is limited to your company's IP addressesCall from an allowed address, or change the key's allowlist in the console. Don't retry.
404not_foundUnknown routeUse /v1/ground/basic or /v1/ground/basic/batch.
405method_not_allowedWrong HTTP methodUse POST (GET for /v1/health).
413request_too_largeBody over 1 MB, or an item over the plan's per-request token limitShorten the context or split the answer. Don't retry unchanged.
422invalid_requestA required field is missing or empty, a value has the wrong type, or a batch has more than 50 itemsRead error.message — it names the field. Don't retry unchanged.
429rate_limitedToo many requests this minute for this keyWait for Retry-After, then retry with backoff.
429daily_quota_exceededToday's allowance is used upWait until X-RateLimit-Reset (00:00 UTC), or ask for more volume in the console.
502verification_unavailableThe checking service couldn't run. Nothing was charged.Retry with bounded backoff. Treat the answer as unverified — never as a pass.
503auth_unavailableWe couldn't verify the key right nowRetry with bounded backoff.

Retry only 429, 502 and 503, with bounded backoff — for example three attempts at 2, 4 and 8 seconds — honouring Retry-After when present. Apply the same budget to network timeouts. Don't retry 401, 403, 404, 405, 413 or 422 without fixing the request.

Your data

  • By default we keep numbers only: scores, decisions and offsets — not the text of your questions, contexts or answers.
  • Text is kept only if your workspace turns it on (Settings → Data & security), for the retention period you choose — or, for Connect data jobs, only for wrong or unsure claims, so your team can review them.
  • To check answers your AI already gave, use ⊸ Connect data in the console. If you prepare the file yourself, one line per answer:
JSON Lines
{"i": 0, "q": "Can I prepay after six months?", "a": "You can prepay after 6 EMIs.", "ctx": [{"text": "Prepayment is allowed after 12 EMIs.", "source": "Loan agreement §4.2", "date": "2025-04-01"}], "t": "2026-09-24T14:02:00Z"}

Only the matched fields of the sampled rows leave your computer, and customer or user ids are never uploaded. See Privacy.

Continue learning

Give us a week of your logs. We'll show you what your AI got wrong.