CalibratedAgents
Try in console
← Blog · Benchmark

Microsoft-Decision-1 vs OpenAI vs Jev vs Cloudflare Clef: decision models benchmarked (2026)

By Calibrated Agents Research Team · 10 Oct 2026 · Updated 10 Oct 2026 · 8 min read

Median milliseconds per judge call · lower is faster

137 msOpenAI's median judge call, the fastest of the four

OpenAI137 ms
Jev329 ms
Clef-flash414 ms
Microsoft-D1624 ms
Accuracy is a three-way tie (AUROC 0.860, 0.852, 0.842); Microsoft is cheapest per call

Microsoft launched Microsoft-Decision-1 on 9 October 2026. The next day, we put it through the same test as the other leading decision APIs: one judge call per answer, on the same LLM-AggreFact answers, measured for speed, accuracy and cost. No single provider wins everywhere.

What are decision models, and why benchmark them?

A decision model takes some content and a fixed set of answer options and returns a calibrated probability for each, instead of generated text. Software can act on that directly: route, block, escalate. Hallucination detection is one of their main jobs: is this answer supported by its source, yes or no?

Microsoft describes decision models as purpose-built to deliver structured outputs that software can immediately act on [1]. OpenAI, Cloudflare and TypeSafe (Jev) ship models of the same kind. Checking AI answers against their sources is a natural job for them, so we measure every new one the same way.

Microsoft's own launch numbers compare Microsoft-Decision-1 with GPT-6 Sol and other decision models on its 36 benchmarks [1]. This is an independent test on one specific job, checking AI answers against their sources, from outside the US, with the method published.

Dot plot of milliseconds per judge call, one call at a time: OpenAI Decisions API median 137 ms (p90 239), Jev 329 (419), Cloudflare Clef-flash 414 (923), Jev via OpenRouter 476 (606), Microsoft-Decision-1 via OpenRouter 624 (726).
Figure 1. Figure 1. Milliseconds per judge call, one call at a time, 100 LLM-AggreFact answers spread across input sizes, uncached, measured from Bengaluru on 10 Oct 2026. Microsoft-Decision-1 is reached through OpenRouter; 'Jev via OpenRouter' shows that route's overhead (about 150 ms).

Which decision API is fastest?

OpenAI's Decisions API, by a wide margin: 137 ms per judge call at the median, against 329 ms for Jev direct, 414 ms for Cloudflare Clef-flash and 624 ms for Microsoft-Decision-1 through OpenRouter. Its slow tail is short too: 239 ms at the 90th percentile.

RouteMedianp90p99Input tokens per call
OpenAI Decisions API · direct137 ms239 ms296 ms1,127
Jev · TypeSafe direct329 ms419 ms602 ms1,246
Cloudflare Clef-flash · Workers AI414 ms923 ms1,168 ms1,107
Jev · OpenRouter476 ms606 ms647 ms1,246
Microsoft-Decision-1 · OpenRouter624 ms726 ms1,724 ms933

Microsoft-Decision-1 is only reachable to us through OpenRouter, so it carries OpenRouter's overhead. Measured on the same calls, that route added about 150 ms to Jev (476 vs 329 ms). Taking the same 150 ms off, Microsoft-Decision-1 direct would sit near 470 ms: still behind Jev and well behind OpenAI, from this location. A direct Azure endpoint closer to India could narrow that; we will test it when we can.

Does input size slow decision models down?

For a single call, barely. From 200 to almost 20,000 characters, OpenAI, Jev and Microsoft-Decision-1 stay flat within a few tens of milliseconds. Cloudflare Clef-flash is the exception: its median grows by 56%, from 343 ms on the shortest inputs to 536 ms on the longest.

Line chart of median milliseconds per judge call by input size, from 0.2–0.9k to 5.1–19.8k characters: OpenAI 132 to 138, Jev 312 to 344, Jev via OpenRouter 468 to 483, Microsoft-Decision-1 612 to 642, Cloudflare Clef-flash 343 to 536.
Figure 2. Figure 2. Median milliseconds per judge call in five equal groups of input size (characters of question + source + answer), one call at a time, 20 answers per group.
Input size (characters)OpenAIJevClef-flashJev · OpenRouterMicrosoft-Decision-1
219–937132312343468612
938–2,003134361390475659
2,004–3,396137332405463617
3,397–5,071147322466483619
5,072–19,759138344536483642

This corrects an impression from our earlier comparison, which timed longer multi-step workloads and found OpenAI slowing on long inputs. Per call, today, it does not: the extra time on long answers comes from the number of decisions a long answer needs, not from each call getting slower.

What happens under load?

Every provider's slow tail grows when four calls run at once. Medians move little for OpenAI, Jev and Microsoft-Decision-1, but the 90th percentile roughly doubles. Cloudflare Clef-flash suffers most: its median rises from 414 to 671 ms and its p90 to 1.3 seconds.

Dumbbell chart of 90th-percentile milliseconds per call, one at a time vs four at a time: OpenAI 239 to 570, Jev 419 to 834, Jev via OpenRouter 606 to 897, Microsoft-Decision-1 726 to 1,110, Cloudflare Clef-flash 923 to 1,336.
Figure 3. Figure 3. 90th-percentile milliseconds per call, one call at a time (hollow) and four at a time (filled), same 100 answers, run back to back.
Four calls at a timeMedianp90p99
OpenAI Decisions API165 ms570 ms599 ms
Jev · direct331 ms834 ms1,061 ms
Jev · OpenRouter505 ms897 ms1,090 ms
Microsoft-Decision-1 · OpenRouter666 ms1,110 ms2,112 ms
Cloudflare Clef-flash671 ms1,336 ms1,610 ms

Which decision model is the most accurate judge?

Three are tied. As a one-call judge of whether an answer is supported by its source, Jev, Microsoft-Decision-1 and the OpenAI Decisions API score AUROC 0.860, 0.852 and 0.842 on the same 296 LLM-AggreFact answers, with overlapping 95% intervals. Cloudflare Clef-flash trails at 0.754.

Dot plot of judge AUROC with 95% intervals on 296 LLM-AggreFact answers: Jev 0.860 (0.818–0.899), Microsoft-Decision-1 0.852 (0.808–0.892), OpenAI 0.842 (0.792–0.883), Cloudflare Clef-flash 0.754 (0.698–0.813).
Figure 4. Figure 4. AUROC of the one-call judge on the same 296 LLM-AggreFact test answers; 95% bootstrap intervals. AUROC measures how well each model ranks hallucinated answers above correct ones, independent of any threshold.
JudgeAUROC (95% interval)Balanced accuracyAgrees with Jev
Jev0.860 (0.818–0.899)77.8—
Microsoft-Decision-10.852 (0.808–0.892)73.191%
OpenAI Decisions API0.842 (0.792–0.883)78.189%
Cloudflare Clef-flash0.754 (0.698–0.813)70.978%

Balanced accuracy uses a threshold set on half of the answers and applied to the other half. Microsoft-Decision-1 ranks as well as the others but loses more when its threshold carries over (73.1), a sign its probabilities are less stable around the decision point on this data. With 296 answers, balanced accuracy moves by a few points between splits, so AUROC is the steadier comparison.

What does each decision cost?

At list prices, Microsoft-Decision-1 is the cheapest: about $0.039 per 1,000 judge calls on these inputs, then Jev at $0.052, Cloudflare Clef-flash at $0.100 and OpenAI at $0.113. Microsoft's edge comes from its tokenizer, which counted 933 tokens per call against 1,107 to 1,246 for the others.

Horizontal bars of USD per 1,000 judge calls at list price: Microsoft-Decision-1 $0.039, Jev $0.052, Cloudflare Clef-flash $0.100, OpenAI Decisions API $0.113.
Figure 5. Figure 5. USD per 1,000 judge calls at list price per million input tokens (Microsoft-Decision-1 and Jev $0.042, Clef-flash $0.09, OpenAI $0.10), from the tokens each provider counted on the same 100 answers. Output tokens are free or negligible for all four.

At finding the wrong sentence, Ground leads them all

Judges answer one question about a whole answer. CalibratedAgents Ground checks every sentence against its source and says what is wrong, why, where and how to fix it. On 5,916 human-labelled sentences, Ground scores 0.877 AUROC on RAGTruth++ and 0.741 on FaithBench, ahead of Microsoft-Decision-1 (0.848, 0.715) and Jev (0.830, 0.701) asked the same question per sentence.

Bar charts of sentence-level AUROC. RAGTruth++: Ground 0.877, Microsoft-Decision-1 0.848, Jev 0.830. FaithBench: Ground 0.741, Microsoft-Decision-1 0.715, Jev 0.701.
Figure 6. Figure 6. Sentence-level AUROC on RAGTruth++ (2,334 sentences) and FaithBench (3,582 sentences). The judges are asked about each sentence and its source; Ground is calibrated on each dataset's labels with grouped 5-fold cross-validation.

Microsoft-Decision-1 is a slightly better per-sentence judge than Jev here (+0.014 to +0.018 AUROC), and Ground still leads it by about 0.03. Which decision model sits underneath matters less than checking each sentence and learning from your own labels. For the full story, read Hallucination detection in 2026: find the wrong sentence.

Which decision model should you use?

It depends on what you are optimising. For the lowest latency, OpenAI's Decisions API. For accuracy with a steady slow tail, Jev. For the lowest cost per decision, Microsoft-Decision-1. Avoid Clef-flash where inputs are long or traffic is bursty. And because no single model wins everywhere, routing each kind of check to the model best at it is the next step.

  • Latency-critical paths (an agent waiting on every step): OpenAI Decisions API, 137 ms per call and flat with input size.
  • Accuracy and predictability: Jev, top AUROC and the second-shortest tail, also available through OpenRouter with identical answers.
  • High-volume, cost-sensitive checks: Microsoft-Decision-1, about a third of OpenAI's cost per decision, at a judge accuracy tied with the best.
  • Finding which sentence is wrong: a sentence-level verifier such as Ground on top of any of them, rather than a single whole-answer judge.
Try Ground freeTalk to us about your stack

How we measured

What this does not prove

  • One location (Bengaluru) and one day. Latency depends on where you call from; repeat it from your region before choosing.
  • Microsoft-Decision-1 was measured through OpenRouter, not a direct Azure endpoint; its direct latency is estimated, not measured.
  • For Cloudflare Clef-flash, 183 of the 296 accuracy answers come from an earlier version of our pipeline that asked the judge slightly differently, which may understate it.
  • One task: checking answers against sources. Decision models are used for routing, classification and more; Microsoft's own 36 benchmarks cover those.

Frequently asked questions

Is Microsoft-Decision-1 faster than the OpenAI Decisions API?

Not in our test. From Bengaluru, one call at a time, OpenAI's Decisions API took 137 ms at the median and Microsoft-Decision-1 took 624 ms through OpenRouter, or about 470 ms with OpenRouter's overhead removed. Location and route matter, so test from your own region.

How accurate is Microsoft-Decision-1 as an LLM judge?

As accurate as the best we tested: AUROC 0.852 on 296 LLM-AggreFact answers, statistically tied with Jev (0.860) and OpenAI (0.842), and ahead of Cloudflare Clef-flash (0.754). Per sentence on RAGTruth++ it scored 0.848, slightly ahead of Jev.

How much does Microsoft-Decision-1 cost?

$0.042 per million input tokens with free output, through OpenRouter and Microsoft Foundry. On our inputs that came to about $0.039 per 1,000 judge calls, the cheapest of the four, because its tokenizer counts fewer tokens.

What is a decision model?

A model that returns a calibrated probability for each of a fixed set of answer options instead of generated text, so software can act on it directly: route, classify, approve or reject. Microsoft-Decision-1, OpenAI's Decisions API, Jev and Cloudflare Clef are all decision models.

Does input length slow decision models down?

Barely, per call: from 200 to 20,000 characters, OpenAI, Jev and Microsoft-Decision-1 stayed within a few tens of milliseconds. Cloudflare Clef-flash slowed by 56% on the longest inputs.

Which decision model is best for hallucination detection?

For a whole-answer yes/no, Jev, Microsoft-Decision-1 and OpenAI are tied on accuracy. To find which sentence is wrong, a sentence-level verifier does better: CalibratedAgents Ground beat all of them on RAGTruth++ and FaithBench.

References

  1. Microsoft. Introducing Microsoft-Decision-1, our model for fast decision-making. 9 Oct 2026.
  2. OpenRouter. Microsoft-Decision-1: pricing, providers and the Decisions API.
  3. OpenRouter. Jev, the TypeSafe decision model, on OpenRouter.
  4. Tang, Laban, Durrett. MiniCheck: efficient fact-checking of LLMs on grounding documents (introduces LLM-AggreFact). EMNLP 2024.
  5. OpenAI. Decisions API guide.
CACalibrated Agents Research Team builds and benchmarks Ground, the sentence-level verification layer for AI answers. We publish our benchmarks with their methods so they can be checked and rerun. Questions and corrections: [email protected].

Cite this

Calibrated Agents Research Team (2026). Microsoft-Decision-1 vs OpenAI vs Jev vs Cloudflare Clef: decision models benchmarked (2026). CalibratedAgents. https://calibratedagents.com/blog/microsoft-decision-1-benchmark
  • Microsoft-Decision-1
  • OpenAI Decisions API
  • Jev
  • Cloudflare Clef
  • Decision models
  • LLM-as-a-judge
  • Latency benchmark

Written by the Calibrated Agents Research Team. Questions or data requests: [email protected]