~/letmereviewyourcode
Guide · decision models

Jev, Clef, Haiku, or nothing?

What a decision model is, who sells one, and what happened when I put four of them in front of an agent.

Zishan Ali Khan · 5 October 2026 · 13 min read · full guide

Updated 8 October 2026. OpenAI opened its own decision model, Luna Decisions, on 6 October. I added it below, tested on its own next to Jev, and ran both on a lending queue: Jev against Luna. The process runs, the grid and the charts are unchanged; Luna did not run inside the process.

Written on 5 October 2026, three weeks after Jev launched and four days after Clef. These models, their prices and their APIs are changing fast, so treat every number here as true on that date. Message me if you are working on this too.

You've probably heard about Jev by now. I like to run my own evals before I say much about a model, and by the time I started, more had shown up: Cloudflare's Clef and Clef-flash, with open weights, Liquid's d1 and Laya. So I tested them on the same job, side by side, and wrote down what each one is and who makes it.

What a decision model is

A decision model answers questions you give it about a situation, and only those questions. You send the situation as text or JSON, plus a short list of typed questions: pick one of these options, answer yes or no, score this from 0 to 3. It sends back a probability for every option of every question, from one pass through the model. There is no paragraph to read and nothing to parse. Your code reads the numbers and acts.

ONE REQUEST, ONE PASS State "what exactly would my money be invested in?" Questions panel: choice of 8 cards needsHuman: true or false fit_growth: score 0 to 3 fit_balanced: score 0 to 3 954 input tokens on Clef Decision model Answers panel .96 allocation · agent .04 · rest 0 needsHuman .07 fit_balanced expected 1.82 of 3 no text to parse
A real request from this study and Jev's answer: 96% on the allocation card, 7% that she needs a person.

The job is old. Spam filters and intent routers have sorted text into labels for decades. Two things are new. The questions arrive with each request, so you don't train a classifier per label set. And the models are built on LLMs: Clef is Qwen3.8-27B with a scoring head on top.

The reason they matter now is cost. When an agent handles a customer turn, its first model call often only works out what she is asking. In this study, Opus 5.5 working alone spent $0.021 a turn. One Jev call cost $0.000037, about 570 times less. A decision model in front can make that first call unnecessary.

The ones I tested, and who makes them

Three companies shipped decision models between 15 September and 1 October 2026, and OpenAI followed on 6 October. These are the ones I put through the test.

JevTypeSafe AI
closed weightstext only$0.042 per M input tokenslaunched 15 Sep 2026

TypeSafe AI is a new AI lab founded by Diogo Almeida. It calls Jev the first of a class it names System One models, after Daniel Kahneman's fast, intuitive System 1 thinking; the name Jev comes from the economist William Stanley Jevons. TypeSafe says Jev is trained with a method it calls Reinforcement Learning for Calibrated Decisions, so its probabilities mean what they say, and it claims 70 to 500 ms per request. Output is free. TypeSafe opened its own API as early access at launch; Jev has also been on Vercel's AI Gateway since 16 September, which is how I reached it.

In my test, on its own: the right card on 98% of messages, about 150 ms inside TypeSafe's servers, $0.000037 a call.

Source: TypeSafe's launch post, 15 Sep 2026.

Clef and Clef-flashCloudflare
open weights, Apache 2.0reads images too$0.24 and $0.09 per M input tokensreleased 1 Oct 2026

Cloudflare built Clef on top of open Qwen models: Clef is Qwen3.8-27B, Clef-flash is Qwen3.5-9B, each frozen, with a small adapter and a scoring head that rates every option of every question at once. The weights are on Hugging Face under Apache 2.0, so you can run them yourself; Cloudflare serves both on Workers AI. Its launch post says it already uses Clef to classify websites for its threat-intelligence team, and that its API is fully compatible with Jev's. The name is the musical symbol, with "CF" for Cloudflare.

In my test, on its own: Clef chose the right card on 100% of messages, Clef-flash on 97%. Because the weights are open, I also ran Clef-flash on my laptop, free: the right card on 97% of messages there too, at about 3 s a call.

Sources: Cloudflare's launch post, model card, Workers AI pricing.

d1Liquid AI
API accessfree tier at Liquid$0.04 per M input tokens on Vercelannounced 29 Sep 2026

Liquid AI also makes the LFM family of language models. It describes d1 as "purpose-built for structured decisions" that returns probabilities "in a single call with zero generated tokens", using the same three question types as Jev. Liquid offers a free tier on its own API, and d1 is on Vercel's AI Gateway and OpenRouter. Liquid says d1 is the first model to outperform Jev on the community Decision Index; I could not find d1 in that index's published data.

In my test, on its own: the right card on 95% of messages, $0.000039 a call. I tested d1 on its own only, not inside the full process.

Source: Liquid's d1 docs.

Luna DecisionsOpenAI
closed weightstext and images$0.10 per M input tokenspublic beta 6 Oct 2026

OpenAI's Decisions API runs on gpt-6-luna, one of the two reasoning models it released on 22 September. You send text, images or both, with questions of three types: a predicate (the probability that a condition is true), a choice (one of your options) and a score (ordered levels). OpenAI says it answers "about 10x faster than the Responses API" and charges for input tokens only. It is in public beta, with general availability expected "in the coming weeks". Vercel's AI Gateway serves it on the same endpoint as Jev, which is how I reached it, and it took Jev's request unchanged.

In my test, on its own (added 8 October): the right card on all 95 calls it answered, $0.000099 a call, against Jev's 93 of 96 calls and $0.000037 in a same-day rerun. I tested Luna on its own only, not inside the full process.

Sources: OpenAI's Decisions guide, API changelog, 6 Oct 2026.

Two more exist: Laya from Convai Innovations, open and free on Vercel, which picked the right card on only 31% of my messages, and Kev 9B from Jared Palmer, open and self-host only, which I did not test. The benchmark tables on these models are mostly run by the vendors themselves; Cloudflare's own leaderboard marks Clef "self-reported".

What I tried to do, and why

When an AI agent answers a customer, a large model does two jobs in one go. It works out what she is asking, and it writes the answer. Getting the first job wrong sends her the wrong answer, and it does not need a large model. Decision models claim to do it in about a second for a fraction of a cent. I wanted to know whether that holds up in a real workflow, with real prices, behind cheap models as well as expensive ones, and which one to pick.

So I ran an eval: a repeatable test with fixed inputs, a written rule for what counts as a right answer, and the same measurements for every setup. Same 32 customer messages, same agent, same tools, same rule. The only thing I changed was what sat in front of the agent, and I recorded the time, the cost and whether the customer got the right answer, on every run.

The test

One turn in a wealth-onboarding chat. A customer types a message. The turn has to show one of eight prepared cards (her allocation, a growth projection, an advisor call and so on) and write one or two sentences with facts from tools. I wrote 32 messages and the cards each one should get, and ran every message through six agent models, from Amazon Nova Micro at $0.035 per million input tokens to Claude Fable 5.1 at $10.

Each model ran five ways: alone, deciding everything itself; and with Jev, Clef, Clef-flash or Claude Haiku 4.5 deciding the card first. Behind each classifier sat the same rule. Show the card when the classifier is at least 50% sure, ask the customer to choose between 30 and 50%, hand her to a person when it is 80% sure she wants one. The agent then only writes the message.

ONE TURN, AS A PROCESS message Classifier Jev, Clef, Clef-flash or Haiku 4.5 Rule a DMN table 50% sure or more show the card, the agent writes 30 to 50% sure ask her which card she wants wants a person, 80%+ hand to an advisor, no model call Below 30%, or no card fits, the agent decides on its own. Without a classifier, it always does.
Swapping one classifier for another meant changing one task in this diagram. The rule, the agent, its ten tools and the 32 messages stayed the same files.

Every combination, in one grid

Six models, five setups each, the same 32 messages in every cell. Pick what to compare.

Darker means more right answers. Each cell is 64 or 96 turns. The Clef-flash column ran in the first window and is compared with that window's agent-alone numbers, shown on hover.

First, each classifier on its own

Before the full runs, I sent every message straight to each classifier, three times, from my laptop. "Right card" is whether its top pick was one the message should get. "Right after the rule" applies the thresholds above, so a pick it was unsure of becomes a question back to the customer.

Times are a laptop's round trip, not the model's own speed. The first six rows are 672 calls from 2 October. The last two ran on 8 October for the OpenAI update, Jev again so Luna is compared with a same-day Jev; Luna's 95 of 96 calls is one timeout.

Update, 8 October: Jev against Luna

OpenAI's Decisions API went to public beta on 6 October, so I ran it next to Jev on a second job: a bank's lending intake queue. 40 items (uploads, broker emails, chats and system events), each routed to one of seven teams, three times each, through the same gateway. Both models got the same request, byte for byte; only the model name changed.

JevLuna DecisionsLuna ÷ Jev
List price, per million input tokens$0.042$0.102.4×
Lending: tokens it counted, per call5924900.83×
Lending: billed per call$0.000025$0.0000492.0×
Lending: right route117 of 120114 of 120
Lending: median round trip244 ms214 ms
Wealth: tokens it counted, per call8859921.12×
Wealth: billed per call$0.000037$0.0000992.7×
Both sets: billed per call$0.000030$0.0000712.3×

Accuracy was close and changed with the job: Jev led on lending, Luna on the wealth messages. Cost was the difference. Across both sets Luna billed 2.3× Jev per call. Part of that is price, $0.10 per million input tokens against $0.042. The rest is counting: for the same request, Luna counted fewer tokens than Jev on the two-question lending request and more on the five-question wealth one, so the gap depends on your questions. At 10 million lending calls a month that is $249 on Jev and $490 on Luna. Luna was a little faster.

Costs are what Vercel's AI Gateway billed for each call on 8 October 2026. The lending items and their labels are mine. Luna ran on its own only, so the grid and the charts in this guide do not include it.

Which classifier made the turn cheaper

This is the cost of a whole turn with each classifier in front, against the same model working alone. Left of the line is cheaper.

Cost per turn against the agent alone

Table view
Each dot is 64 or 96 turns on the same 32 messages. Clef-flash ran in a separate window with its own agent-alone baseline. Hover for the numbers.

On the Claude models, each of the three decision models cut the cost of a turn by about half. Jev cut 45 to 47% and kept its answers close to the agent alone; Clef cut 43 to 44%. On the two cheapest models the picture flips. One Clef call costs $0.00023, more than Nova Micro's whole turn of $0.00013, so putting Clef in front made that turn 127% more expensive. Jev, at a sixth of Clef's price, stayed cheaper than the model alone on all six.

On frontier models

This is where most teams feel the bill. On Opus 5.5 and Fable 5.1, every classifier in front made the turn cheaper and faster. Jev took about 45% off the cost and 1.5 to 2.8 seconds off the median turn, for two or three points of right answers. The Haiku router took 35 to 39% off and lost nothing on Opus, five points on Fable. Clef-flash was the cheapest and asked the customer most.

Opus 5.5, cost a turn

Fable 5.1, cost a turn

Bars are cost per turn; the labels give right answers and median time. Clef-flash's bar comes from the first window, where the model alone cost the same to within a tenth of a cent.

Try it for your own stack

Measured, not modelled: each number is this study's mean over 32 messages. Your prompts and messages will move them.

Right answers, questions back, and misses

A turn counts as right when the customer sees an accepted card and the message carries the fact that card needs, for example the 0.85% fee or her advisor's name. A question back to her ("would you like the comparison or the risk profile?") never counts as right, because she still has to answer it.

What the customer got, averaged over six models

right card and factasked her to choosewrong card or missing fact
The Haiku router, Jev and Clef ran 384 turns each, Clef-flash 576. The agent alone answered right on 91 to 100% of turns per model.

Clef-flash picked the right card 97% of the time when I tested it on its own. In the process it was right on only 76% of turns, because it asked the customer on 22%. Its answers are fine; its confidence runs lower than Jev's. When it was right it was about 57% sure, where Jev was about 92% sure, and a rule written around Jev's numbers reads 57% as "not sure enough to show". Clef, the larger model, asked on 6% and landed within a few points of Jev.

Benchmark speed and speed in a process

Cloudflare's own benchmark table, run by Cloudflare, puts Clef-flash at 38.8 ms and Jev at 524.1 ms. Inside the process, each classifier step took about a second, start to finish.

Classifier time, milliseconds

Cloudflare's model card, model sidemeasured in the process, median
Measured from the step's start to its end, 384 to 576 calls each. Clef's steps ran across three windows, so its time carries an hour-of-day caveat.

Most of that second is the trip out and back: the call out of the process, the network, the provider's queue. The model's own share is a rounding error. That is fine. Saving one frontier-model call saves several seconds, and that is where the speed came from: with Jev in front, Fable 5.1 answered 3.1 s faster than alone.

Three things I'd tell a team choosing one

Compare the classifier's price with the turn behind it.

A classifier earns its place by replacing a model call that costs much more than it does. Behind Fable or Opus, every one I tried paid. Behind Nova Micro, Clef and Haiku cost more than the turn they were meant to shorten.

Set your thresholds per model.

Probabilities are only useful when your rule reads them on the right scale. Jev's 0.5 is not Clef-flash's 0.5. Tune the rule on a held-out set for each classifier you try, and test a swap the way you would test a new model.

Read "compatible" as "similar".

Clef's API matches Jev's own. It does not match Jev as Vercel serves it, where the yes-or-no type is called boolean and the answer field probability. On Clef the field is noul. Swap one for the other with only a type change, and the needs-a-person probability reads as zero: the handoff never fires, and nothing errors.

The work behind the numbers

Jev came through Vercel's AI Gateway, because TypeSafe's own API was in early access when I started; I bought a dollar of gateway credits when the free tier closed. Clef ran on Cloudflare's free plan, which allows 10,000 neurons a day, about 480 Clef calls. That ran out twice, mid-run. 236 turns fell back to the agent; I voided them and reran them when the allowance came back. Before any of that, every message went to each classifier on its own: 672 calls, including Clef-flash on my laptop. All in, 3,264 scored turns over one weekend, for $36.87 in model spend.

One thing made the volume manageable. Each setup is a BPMN process I drew in Camunda rather than code I wrote. Moving the classifier, adding the rule or adding the handoff meant changing a diagram, and I could replay any of the 3,264 runs step by step, with its timings, straight from Camunda, instead of building tracing over code first. I'm grateful for that.

How I ran it, and its limits

Each turn ran as a BPMN process in Camunda, and every time and cost here comes from the engine's own records. I wrote down five predictions before any Clef call. Three held, the one on speed was wrong, and the one on right answers held for Clef but not for Clef-flash. 32 messages from one demo is a small, single-task test. I hand-labelled the cards on 26 September and reviewed the four checks that looked wrong on 5 October; three changed, and I rescored every saved answer against the review.

The one-paragraph version

A decision model returns probabilities for typed questions in one pass, which lets a rule decide before an agent spends a frontier-model call. On 32 messages and six models, Jev in front cut the cost of a turn by 9 to 47% with answers close to the agent alone. Clef matched it on big models and more than doubled the cost on the cheapest. Clef-flash was cheaper still but asked the customer too often under thresholds tuned for Jev. The classifier mattered less than its price next to the agent, the thresholds around it, and the process that holds them.

References

TypeSafe, Introducing System One models and Jev · Cloudflare, Clef decision models · Clef model card · Workers AI pricing · Vercel AI Gateway, evaluation · Decision Index

Get the next one

New field guides and episodes, straight to your inbox. No noise, unsubscribe anytime.