Jev, Clef, Haiku, or nothing?
What a decision model is, who sells one, and what happened when I put four of them in front of an agent.
Updated 8 October 2026. OpenAI opened its own decision model, Luna Decisions, on 6 October. I added it below, tested on its own next to Jev, and ran both on a lending queue: Jev against Luna. The process runs, the grid and the charts are unchanged; Luna did not run inside the process.
Written on 5 October 2026, three weeks after Jev launched and four days after Clef. These models, their prices and their APIs are changing fast, so treat every number here as true on that date. Message me if you are working on this too.
You've probably heard about Jev by now. I like to run my own evals before I say much about a model, and by the time I started, more had shown up: Cloudflare's Clef and Clef-flash, with open weights, Liquid's d1 and Laya. So I tested them on the same job, side by side, and wrote down what each one is and who makes it.
What a decision model is
A decision model answers questions you give it about a situation, and only those questions. You send the situation as text or JSON, plus a short list of typed questions: pick one of these options, answer yes or no, score this from 0 to 3. It sends back a probability for every option of every question, from one pass through the model. There is no paragraph to read and nothing to parse. Your code reads the numbers and acts.
The job is old. Spam filters and intent routers have sorted text into labels for decades. Two things are new. The questions arrive with each request, so you don't train a classifier per label set. And the models are built on LLMs: Clef is Qwen3.8-27B with a scoring head on top.
The reason they matter now is cost. When an agent handles a customer turn, its first model call often only works out what she is asking. In this study, Opus 5.5 working alone spent $0.021 a turn. One Jev call cost $0.000037, about 570 times less. A decision model in front can make that first call unnecessary.
The ones I tested, and who makes them
Three companies shipped decision models between 15 September and 1 October 2026, and OpenAI followed on 6 October. These are the ones I put through the test.
TypeSafe AI is a new AI lab founded by Diogo Almeida. It calls Jev the first of a class it names System One models, after Daniel Kahneman's fast, intuitive System 1 thinking; the name Jev comes from the economist William Stanley Jevons. TypeSafe says Jev is trained with a method it calls Reinforcement Learning for Calibrated Decisions, so its probabilities mean what they say, and it claims 70 to 500 ms per request. Output is free. TypeSafe opened its own API as early access at launch; Jev has also been on Vercel's AI Gateway since 16 September, which is how I reached it.
In my test, on its own: the right card on 98% of messages, about 150 ms inside TypeSafe's servers, $0.000037 a call.
Source: TypeSafe's launch post, 15 Sep 2026.
Cloudflare built Clef on top of open Qwen models: Clef is Qwen3.8-27B, Clef-flash is Qwen3.5-9B, each frozen, with a small adapter and a scoring head that rates every option of every question at once. The weights are on Hugging Face under Apache 2.0, so you can run them yourself; Cloudflare serves both on Workers AI. Its launch post says it already uses Clef to classify websites for its threat-intelligence team, and that its API is fully compatible with Jev's. The name is the musical symbol, with "CF" for Cloudflare.
In my test, on its own: Clef chose the right card on 100% of messages, Clef-flash on 97%. Because the weights are open, I also ran Clef-flash on my laptop, free: the right card on 97% of messages there too, at about 3 s a call.
Sources: Cloudflare's launch post, model card, Workers AI pricing.
Liquid AI also makes the LFM family of language models. It describes d1 as "purpose-built for structured decisions" that returns probabilities "in a single call with zero generated tokens", using the same three question types as Jev. Liquid offers a free tier on its own API, and d1 is on Vercel's AI Gateway and OpenRouter. Liquid says d1 is the first model to outperform Jev on the community Decision Index; I could not find d1 in that index's published data.
In my test, on its own: the right card on 95% of messages, $0.000039 a call. I tested d1 on its own only, not inside the full process.
Source: Liquid's d1 docs.
OpenAI's Decisions API runs on gpt-6-luna, one of the two reasoning models it released on 22 September. You send text, images or both, with questions of three types: a predicate (the probability that a condition is true), a choice (one of your options) and a score (ordered levels). OpenAI says it answers "about 10x faster than the Responses API" and charges for input tokens only. It is in public beta, with general availability expected "in the coming weeks". Vercel's AI Gateway serves it on the same endpoint as Jev, which is how I reached it, and it took Jev's request unchanged.
In my test, on its own (added 8 October): the right card on all 95 calls it answered, $0.000099 a call, against Jev's 93 of 96 calls and $0.000037 in a same-day rerun. I tested Luna on its own only, not inside the full process.
Sources: OpenAI's Decisions guide, API changelog, 6 Oct 2026.
Two more exist: Laya from Convai Innovations, open and free on Vercel, which picked the right card on only 31% of my messages, and Kev 9B from Jared Palmer, open and self-host only, which I did not test. The benchmark tables on these models are mostly run by the vendors themselves; Cloudflare's own leaderboard marks Clef "self-reported".
What I tried to do, and why
When an AI agent answers a customer, a large model does two jobs in one go. It works out what she is asking, and it writes the answer. Getting the first job wrong sends her the wrong answer, and it does not need a large model. Decision models claim to do it in about a second for a fraction of a cent. I wanted to know whether that holds up in a real workflow, with real prices, behind cheap models as well as expensive ones, and which one to pick.
So I ran an eval: a repeatable test with fixed inputs, a written rule for what counts as a right answer, and the same measurements for every setup. Same 32 customer messages, same agent, same tools, same rule. The only thing I changed was what sat in front of the agent, and I recorded the time, the cost and whether the customer got the right answer, on every run.
The test
One turn in a wealth-onboarding chat. A customer types a message. The turn has to show one of eight prepared cards (her allocation, a growth projection, an advisor call and so on) and write one or two sentences with facts from tools. I wrote 32 messages and the cards each one should get, and ran every message through six agent models, from Amazon Nova Micro at $0.035 per million input tokens to Claude Fable 5.1 at $10.
Each model ran five ways: alone, deciding everything itself; and with Jev, Clef, Clef-flash or Claude Haiku 4.5 deciding the card first. Behind each classifier sat the same rule. Show the card when the classifier is at least 50% sure, ask the customer to choose between 30 and 50%, hand her to a person when it is 80% sure she wants one. The agent then only writes the message.
Every combination, in one grid
Six models, five setups each, the same 32 messages in every cell. Pick what to compare.
First, each classifier on its own
Before the full runs, I sent every message straight to each classifier, three times, from my laptop. "Right card" is whether its top pick was one the message should get. "Right after the rule" applies the thresholds above, so a pick it was unsure of becomes a question back to the customer.
Times are a laptop's round trip, not the model's own speed. The first six rows are 672 calls from 2 October. The last two ran on 8 October for the OpenAI update, Jev again so Luna is compared with a same-day Jev; Luna's 95 of 96 calls is one timeout.
Update, 8 October: Jev against Luna
OpenAI's Decisions API went to public beta on 6 October, so I ran it next to Jev on a second job: a bank's lending intake queue. 40 items (uploads, broker emails, chats and system events), each routed to one of seven teams, three times each, through the same gateway. Both models got the same request, byte for byte; only the model name changed.
| Jev | Luna Decisions | Luna ÷ Jev | |
|---|---|---|---|
| List price, per million input tokens | $0.042 | $0.10 | 2.4× |
| Lending: tokens it counted, per call | 592 | 490 | 0.83× |
| Lending: billed per call | $0.000025 | $0.000049 | 2.0× |
| Lending: right route | 117 of 120 | 114 of 120 | |
| Lending: median round trip | 244 ms | 214 ms | |
| Wealth: tokens it counted, per call | 885 | 992 | 1.12× |
| Wealth: billed per call | $0.000037 | $0.000099 | 2.7× |
| Both sets: billed per call | $0.000030 | $0.000071 | 2.3× |
Accuracy was close and changed with the job: Jev led on lending, Luna on the wealth messages. Cost was the difference. Across both sets Luna billed 2.3× Jev per call. Part of that is price, $0.10 per million input tokens against $0.042. The rest is counting: for the same request, Luna counted fewer tokens than Jev on the two-question lending request and more on the five-question wealth one, so the gap depends on your questions. At 10 million lending calls a month that is $249 on Jev and $490 on Luna. Luna was a little faster.
Costs are what Vercel's AI Gateway billed for each call on 8 October 2026. The lending items and their labels are mine. Luna ran on its own only, so the grid and the charts in this guide do not include it.
Which classifier made the turn cheaper
This is the cost of a whole turn with each classifier in front, against the same model working alone. Left of the line is cheaper.
Cost per turn against the agent alone
Table view
On the Claude models, each of the three decision models cut the cost of a turn by about half. Jev cut 45 to 47% and kept its answers close to the agent alone; Clef cut 43 to 44%. On the two cheapest models the picture flips. One Clef call costs $0.00023, more than Nova Micro's whole turn of $0.00013, so putting Clef in front made that turn 127% more expensive. Jev, at a sixth of Clef's price, stayed cheaper than the model alone on all six.
On frontier models
This is where most teams feel the bill. On Opus 5.5 and Fable 5.1, every classifier in front made the turn cheaper and faster. Jev took about 45% off the cost and 1.5 to 2.8 seconds off the median turn, for two or three points of right answers. The Haiku router took 35 to 39% off and lost nothing on Opus, five points on Fable. Clef-flash was the cheapest and asked the customer most.
Opus 5.5, cost a turn
Fable 5.1, cost a turn
Try it for your own stack
Measured, not modelled: each number is this study's mean over 32 messages. Your prompts and messages will move them.
Right answers, questions back, and misses
A turn counts as right when the customer sees an accepted card and the message carries the fact that card needs, for example the 0.85% fee or her advisor's name. A question back to her ("would you like the comparison or the risk profile?") never counts as right, because she still has to answer it.
What the customer got, averaged over six models
Clef-flash picked the right card 97% of the time when I tested it on its own. In the process it was right on only 76% of turns, because it asked the customer on 22%. Its answers are fine; its confidence runs lower than Jev's. When it was right it was about 57% sure, where Jev was about 92% sure, and a rule written around Jev's numbers reads 57% as "not sure enough to show". Clef, the larger model, asked on 6% and landed within a few points of Jev.
Benchmark speed and speed in a process
Cloudflare's own benchmark table, run by Cloudflare, puts Clef-flash at 38.8 ms and Jev at 524.1 ms. Inside the process, each classifier step took about a second, start to finish.
Classifier time, milliseconds
Most of that second is the trip out and back: the call out of the process, the network, the provider's queue. The model's own share is a rounding error. That is fine. Saving one frontier-model call saves several seconds, and that is where the speed came from: with Jev in front, Fable 5.1 answered 3.1 s faster than alone.
Three things I'd tell a team choosing one
A classifier earns its place by replacing a model call that costs much more than it does. Behind Fable or Opus, every one I tried paid. Behind Nova Micro, Clef and Haiku cost more than the turn they were meant to shorten.
Probabilities are only useful when your rule reads them on the right scale. Jev's 0.5 is not Clef-flash's 0.5. Tune the rule on a held-out set for each classifier you try, and test a swap the way you would test a new model.
Clef's API matches Jev's own. It does not match Jev as Vercel serves it, where the yes-or-no type is called boolean and the answer field probability. On Clef the field is noul. Swap one for the other with only a type change, and the needs-a-person probability reads as zero: the handoff never fires, and nothing errors.
The work behind the numbers
Jev came through Vercel's AI Gateway, because TypeSafe's own API was in early access when I started; I bought a dollar of gateway credits when the free tier closed. Clef ran on Cloudflare's free plan, which allows 10,000 neurons a day, about 480 Clef calls. That ran out twice, mid-run. 236 turns fell back to the agent; I voided them and reran them when the allowance came back. Before any of that, every message went to each classifier on its own: 672 calls, including Clef-flash on my laptop. All in, 3,264 scored turns over one weekend, for $36.87 in model spend.
One thing made the volume manageable. Each setup is a BPMN process I drew in Camunda rather than code I wrote. Moving the classifier, adding the rule or adding the handoff meant changing a diagram, and I could replay any of the 3,264 runs step by step, with its timings, straight from Camunda, instead of building tracing over code first. I'm grateful for that.
How I ran it, and its limits
Each turn ran as a BPMN process in Camunda, and every time and cost here comes from the engine's own records. I wrote down five predictions before any Clef call. Three held, the one on speed was wrong, and the one on right answers held for Clef but not for Clef-flash. 32 messages from one demo is a small, single-task test. I hand-labelled the cards on 26 September and reviewed the four checks that looked wrong on 5 October; three changed, and I rescored every saved answer against the review.
The one-paragraph version
A decision model returns probabilities for typed questions in one pass, which lets a rule decide before an agent spends a frontier-model call. On 32 messages and six models, Jev in front cut the cost of a turn by 9 to 47% with answers close to the agent alone. Clef matched it on big models and more than doubled the cost on the cheapest. Clef-flash was cheaper still but asked the customer too often under thresholds tuned for Jev. The classifier mattered less than its price next to the agent, the thresholds around it, and the process that holds them.
References
TypeSafe, Introducing System One models and Jev · Cloudflare, Clef decision models · Clef model card · Workers AI pricing · Vercel AI Gateway, evaluation · Decision Index
Get the next one
New field guides and episodes, straight to your inbox. No noise, unsubscribe anytime.