Microsoft-Decision-1 vs OpenAI vs Jev vs Cloudflare Clef: decision models benchmarked (2026)
By Calibrated Agents Research Team · 10 Oct 2026 · Updated 10 Oct 2026 · 8 min read
Median milliseconds per judge call · lower is faster
137 msOpenAI's median judge call, the fastest of the four
Microsoft launched Microsoft-Decision-1 on 9 October 2026. The next day, we put it through the same test as the other leading decision APIs: one judge call per answer, on the same LLM-AggreFact answers, measured for speed, accuracy and cost. No single provider wins everywhere.
What are decision models, and why benchmark them?
A decision model takes some content and a fixed set of answer options and returns a calibrated probability for each, instead of generated text. Software can act on that directly: route, block, escalate. Hallucination detection is one of their main jobs: is this answer supported by its source, yes or no?
Microsoft describes decision models as purpose-built to deliver structured outputs that software can immediately act on [1]. OpenAI, Cloudflare and TypeSafe (Jev) ship models of the same kind. Checking AI answers against their sources is a natural job for them, so we measure every new one the same way.
Microsoft's own launch numbers compare Microsoft-Decision-1 with GPT-6 Sol and other decision models on its 36 benchmarks [1]. This is an independent test on one specific job, checking AI answers against their sources, from outside the US, with the method published.

Which decision API is fastest?
OpenAI's Decisions API, by a wide margin: 137 ms per judge call at the median, against 329 ms for Jev direct, 414 ms for Cloudflare Clef-flash and 624 ms for Microsoft-Decision-1 through OpenRouter. Its slow tail is short too: 239 ms at the 90th percentile.
| Route | Median | p90 | p99 | Input tokens per call |
|---|---|---|---|---|
| OpenAI Decisions API · direct | 137 ms | 239 ms | 296 ms | 1,127 |
| Jev · TypeSafe direct | 329 ms | 419 ms | 602 ms | 1,246 |
| Cloudflare Clef-flash · Workers AI | 414 ms | 923 ms | 1,168 ms | 1,107 |
| Jev · OpenRouter | 476 ms | 606 ms | 647 ms | 1,246 |
| Microsoft-Decision-1 · OpenRouter | 624 ms | 726 ms | 1,724 ms | 933 |
Microsoft-Decision-1 is only reachable to us through OpenRouter, so it carries OpenRouter's overhead. Measured on the same calls, that route added about 150 ms to Jev (476 vs 329 ms). Taking the same 150 ms off, Microsoft-Decision-1 direct would sit near 470 ms: still behind Jev and well behind OpenAI, from this location. A direct Azure endpoint closer to India could narrow that; we will test it when we can.
Does input size slow decision models down?
For a single call, barely. From 200 to almost 20,000 characters, OpenAI, Jev and Microsoft-Decision-1 stay flat within a few tens of milliseconds. Cloudflare Clef-flash is the exception: its median grows by 56%, from 343 ms on the shortest inputs to 536 ms on the longest.

| Input size (characters) | OpenAI | Jev | Clef-flash | Jev · OpenRouter | Microsoft-Decision-1 |
|---|---|---|---|---|---|
| 219–937 | 132 | 312 | 343 | 468 | 612 |
| 938–2,003 | 134 | 361 | 390 | 475 | 659 |
| 2,004–3,396 | 137 | 332 | 405 | 463 | 617 |
| 3,397–5,071 | 147 | 322 | 466 | 483 | 619 |
| 5,072–19,759 | 138 | 344 | 536 | 483 | 642 |
This corrects an impression from our earlier comparison, which timed longer multi-step workloads and found OpenAI slowing on long inputs. Per call, today, it does not: the extra time on long answers comes from the number of decisions a long answer needs, not from each call getting slower.
What happens under load?
Every provider's slow tail grows when four calls run at once. Medians move little for OpenAI, Jev and Microsoft-Decision-1, but the 90th percentile roughly doubles. Cloudflare Clef-flash suffers most: its median rises from 414 to 671 ms and its p90 to 1.3 seconds.

| Four calls at a time | Median | p90 | p99 |
|---|---|---|---|
| OpenAI Decisions API | 165 ms | 570 ms | 599 ms |
| Jev · direct | 331 ms | 834 ms | 1,061 ms |
| Jev · OpenRouter | 505 ms | 897 ms | 1,090 ms |
| Microsoft-Decision-1 · OpenRouter | 666 ms | 1,110 ms | 2,112 ms |
| Cloudflare Clef-flash | 671 ms | 1,336 ms | 1,610 ms |
Which decision model is the most accurate judge?
Three are tied. As a one-call judge of whether an answer is supported by its source, Jev, Microsoft-Decision-1 and the OpenAI Decisions API score AUROC 0.860, 0.852 and 0.842 on the same 296 LLM-AggreFact answers, with overlapping 95% intervals. Cloudflare Clef-flash trails at 0.754.

| Judge | AUROC (95% interval) | Balanced accuracy | Agrees with Jev |
|---|---|---|---|
| Jev | 0.860 (0.818–0.899) | 77.8 | — |
| Microsoft-Decision-1 | 0.852 (0.808–0.892) | 73.1 | 91% |
| OpenAI Decisions API | 0.842 (0.792–0.883) | 78.1 | 89% |
| Cloudflare Clef-flash | 0.754 (0.698–0.813) | 70.9 | 78% |
Balanced accuracy uses a threshold set on half of the answers and applied to the other half. Microsoft-Decision-1 ranks as well as the others but loses more when its threshold carries over (73.1), a sign its probabilities are less stable around the decision point on this data. With 296 answers, balanced accuracy moves by a few points between splits, so AUROC is the steadier comparison.
What does each decision cost?
At list prices, Microsoft-Decision-1 is the cheapest: about $0.039 per 1,000 judge calls on these inputs, then Jev at $0.052, Cloudflare Clef-flash at $0.100 and OpenAI at $0.113. Microsoft's edge comes from its tokenizer, which counted 933 tokens per call against 1,107 to 1,246 for the others.

At finding the wrong sentence, Ground leads them all
Judges answer one question about a whole answer. CalibratedAgents Ground checks every sentence against its source and says what is wrong, why, where and how to fix it. On 5,916 human-labelled sentences, Ground scores 0.877 AUROC on RAGTruth++ and 0.741 on FaithBench, ahead of Microsoft-Decision-1 (0.848, 0.715) and Jev (0.830, 0.701) asked the same question per sentence.

Microsoft-Decision-1 is a slightly better per-sentence judge than Jev here (+0.014 to +0.018 AUROC), and Ground still leads it by about 0.03. Which decision model sits underneath matters less than checking each sentence and learning from your own labels. For the full story, read Hallucination detection in 2026: find the wrong sentence.
Which decision model should you use?
It depends on what you are optimising. For the lowest latency, OpenAI's Decisions API. For accuracy with a steady slow tail, Jev. For the lowest cost per decision, Microsoft-Decision-1. Avoid Clef-flash where inputs are long or traffic is bursty. And because no single model wins everywhere, routing each kind of check to the model best at it is the next step.
- Latency-critical paths (an agent waiting on every step): OpenAI Decisions API, 137 ms per call and flat with input size.
- Accuracy and predictability: Jev, top AUROC and the second-shortest tail, also available through OpenRouter with identical answers.
- High-volume, cost-sensitive checks: Microsoft-Decision-1, about a third of OpenAI's cost per decision, at a judge accuracy tied with the best.
- Finding which sentence is wrong: a sentence-level verifier such as Ground on top of any of them, rather than a single whole-answer judge.
How we measured
What this does not prove
- One location (Bengaluru) and one day. Latency depends on where you call from; repeat it from your region before choosing.
- Microsoft-Decision-1 was measured through OpenRouter, not a direct Azure endpoint; its direct latency is estimated, not measured.
- For Cloudflare Clef-flash, 183 of the 296 accuracy answers come from an earlier version of our pipeline that asked the judge slightly differently, which may understate it.
- One task: checking answers against sources. Decision models are used for routing, classification and more; Microsoft's own 36 benchmarks cover those.
Frequently asked questions
Is Microsoft-Decision-1 faster than the OpenAI Decisions API?
Not in our test. From Bengaluru, one call at a time, OpenAI's Decisions API took 137 ms at the median and Microsoft-Decision-1 took 624 ms through OpenRouter, or about 470 ms with OpenRouter's overhead removed. Location and route matter, so test from your own region.
How accurate is Microsoft-Decision-1 as an LLM judge?
As accurate as the best we tested: AUROC 0.852 on 296 LLM-AggreFact answers, statistically tied with Jev (0.860) and OpenAI (0.842), and ahead of Cloudflare Clef-flash (0.754). Per sentence on RAGTruth++ it scored 0.848, slightly ahead of Jev.
How much does Microsoft-Decision-1 cost?
$0.042 per million input tokens with free output, through OpenRouter and Microsoft Foundry. On our inputs that came to about $0.039 per 1,000 judge calls, the cheapest of the four, because its tokenizer counts fewer tokens.
What is a decision model?
A model that returns a calibrated probability for each of a fixed set of answer options instead of generated text, so software can act on it directly: route, classify, approve or reject. Microsoft-Decision-1, OpenAI's Decisions API, Jev and Cloudflare Clef are all decision models.
Does input length slow decision models down?
Barely, per call: from 200 to 20,000 characters, OpenAI, Jev and Microsoft-Decision-1 stayed within a few tens of milliseconds. Cloudflare Clef-flash slowed by 56% on the longest inputs.
Which decision model is best for hallucination detection?
For a whole-answer yes/no, Jev, Microsoft-Decision-1 and OpenAI are tied on accuracy. To find which sentence is wrong, a sentence-level verifier does better: CalibratedAgents Ground beat all of them on RAGTruth++ and FaithBench.
References
- Microsoft. Introducing Microsoft-Decision-1, our model for fast decision-making. 9 Oct 2026.
- OpenRouter. Microsoft-Decision-1: pricing, providers and the Decisions API.
- OpenRouter. Jev, the TypeSafe decision model, on OpenRouter.
- Tang, Laban, Durrett. MiniCheck: efficient fact-checking of LLMs on grounding documents (introduces LLM-AggreFact). EMNLP 2024.
- OpenAI. Decisions API guide.
Cite this
Calibrated Agents Research Team (2026). Microsoft-Decision-1 vs OpenAI vs Jev vs Cloudflare Clef: decision models benchmarked (2026). CalibratedAgents. https://calibratedagents.com/blog/microsoft-decision-1-benchmark