OpenAI Decisions API vs Cloudflare Clef vs Jev: latency, cost and accuracy as the input grows
By Calibrated Agents Research Team · 8 Oct 2026 · Updated 8 Oct 2026 · 8 min read
Median seconds per answer · 6,100-character inputs
6×how much slower OpenAI got from the smallest inputs to the largest
Three decision APIs on the same 280 answer-verification checks, with every number split by input size. OpenAI is fastest on small inputs and slowest on large ones; Jev is flat; Clef-flash is slow and steady.
Setup
Decision models answer typed questions about a text with a probability. They are the cheapest and fastest way to ask many small questions, which is what checking AI answers needs. We compared three: OpenAI's Decisions API (gpt-6-luna), Cloudflare's Clef-flash and TypeSafe's Jev.
The workload: 280 labelled answers from LLM-AggreFact, each verified against its source with several requests to the model, the same requests for every API. We measured the total input of each check (question, source and answer together) and split the answers into five equal groups of 56, from about 600 characters to about 6,100. All three ran on the same answers, four at a time, from one machine, with every cache off, on 7 and 8 October 2026.
For depth, see LLM-AggreFact and its eleven datasets.
Latency by input size

Jev is flat: from the smallest inputs to the largest its median moves from 0.95 to 1.04 seconds, and its 90th percentile from 1.1 to 1.4. OpenAI's Decisions API is the fastest on small inputs and the slowest on large ones: 0.46 seconds at 600 characters, half of Jev's, then 2.7 seconds at 6,100, with a 90th percentile of 7 seconds. Clef-flash is slow and steady, 1.7 to 2.4 seconds.
Launch figures (about 150 ms per call for OpenAI; a 38.8 ms median for Clef-flash against 524 ms for Jev in Cloudflare's post) are per small call. They say where each curve starts, not how steep it is.
Cost by input size

The same work is counted very differently. On the largest inputs Jev counted 16,900 tokens per answer and OpenAI 31,200. Per 1,000 answers, Jev rises from $0.47 to $0.71 across the five sizes, OpenAI from $1.18 to $3.12, and Clef-flash stays between $1.15 and $2.06. Clef-flash's count falls on the two largest groups, which suggests long inputs are shortened before they are counted; we have not confirmed why.
Accuracy as a one-question judge

As judges, OpenAI's Decisions API and Jev are level overall, 76.8 and 76.3, and they move together across input sizes. Clef-flash is 8 points lower at 68.2, and falls behind most on the two largest groups, the same groups where its token counts dropped.
For depth, see where a calibrated judge falls short: finding the sentence that is wrong.
A 0.5 is not a 0.5
The thresholds that worked best on our development answers were 0.25 for Jev, 0.11 for OpenAI and 0.28 for Clef-flash. Their probabilities do not mean the same thing, so a threshold copied from one model to another can be badly wrong.
For depth, see why choosing your own operating point matters.
Analysis
Three findings a single benchmark number would hide. First, the ranking by speed reverses with input size: the model that wins on a short routing prompt loses on a document-sized check, and the 90th percentile moves far more than the median. Second, price per million tokens is not cost: two models with near-identical list prices differed 2.5× in what they counted for the same text. Third, probabilities are not portable: each model needs its own threshold, chosen on labelled examples from the task it will be used for.
Test a decision API yourself
- Bucket every measurement by input size. A model that wins on short inputs can lose on long ones.
- Time whole tasks, not calls, and read the 90th percentile.
- Turn every cache off when you time and count. Our first runs reused cached answers and looked 25% faster and half as expensive as they really were.
- Count tokens on your own text at several sizes, and compare cost per task.
- Check rate limits in tokens per minute against your load: OpenAI's limit for our account was reached with four answers in flight.
- Choose each model's threshold on labelled examples from your own domain.
How to use this with CalibratedAgents Ground
Choosing and testing decision models for verification is what we do, so you don't have to. Ground checks AI answers against their sources and returns a calibrated probability, the sentence that's wrong, and the passage that proves it.
Try it on an answer How Ground works API docsFrequently asked questions
Which decision API is fastest?
It depends on input size. On small inputs (about 600 characters) OpenAI's Decisions API was fastest at 0.45 seconds per answer; on large ones (about 6,100 characters) Jev was fastest at about 1 second, against 2.4 for Clef-flash and 2.7 for OpenAI.
How much does it cost to verify 1,000 answers?
At list prices on our workload: $0.43–$0.71 with Jev, $1.15–$2.06 with Clef-flash, and $1.13–$3.12 with OpenAI's Decisions API, rising with input size.
Why do token counts differ between APIs for the same text?
Different tokenizers and different ways of packing questions into a request. On our largest inputs OpenAI counted 31,200 tokens per answer where Jev counted 16,900.
Which decision model is the most accurate judge?
On our 280 answers, Jev and OpenAI's Decisions API are level (76.3 and 76.8 balanced accuracy); Clef-flash is 8 points lower at 68.2.
Is a 0.5 probability the same across decision models?
No. The best thresholds on our development answers were 0.25 for Jev, 0.11 for OpenAI and 0.28 for Clef-flash. Calibrate each model on your own labelled examples.
How should I benchmark a decision API?
Time whole tasks with caching off, bucket by input size, read the 90th percentile, count tokens on your own text, check rate limits, and score on labelled examples from your domain.
Sources
- OpenAI, Decisions API documentation: $0.10 per million input tokens; launch latency claim about 150 ms.
- Cloudflare, Clef launch post and model page: median latency Clef-flash 38.8 ms, Clef 209.3 ms, Jev 524.1 ms.
- TypeSafe, Jev documentation: $0.042 per million input tokens.
- Perplexity, Decisions API quickstart (not tested).
- Clef-flash price (about $0.09 per million input tokens) from launch coverage; confirm on Cloudflare's pricing page.
- LLM-AggreFact. Our own measurements: medians and 90th percentiles over the answers stated, fresh runs, October 2026.
Cite this
CalibratedAgents Research (2026). OpenAI Decisions API vs Cloudflare Clef vs Jev: latency, cost and accuracy as the input grows. CalibratedAgents. https://calibratedagents.com/blog/decision-apis-benchmark-2026