TL;DR
- Cloudflare released Clef and the smaller Clef-flash on 1 October 2026: open decision models that answer the same kind of request as TypeSafe's Jev. I ran both through the exact tests from my earlier case triage and deployment failure posts, with no change to the questions.
- Accuracy: a three-way tie at the top. Clef, Jev and Einstein were within a point or two of each other on every main question. Clef was the only option that flagged every urgent support case.
- Clef's confidence is the one worth trusting. Sending its unsure answers to a person caught all 15 of its mistakes. The same approach with Jev caught 4 of 17.
- Jev is still faster and cheaper. From my desk Clef typically took 0.52 seconds against Jev's 0.26, and cost about six times as much: still under five cents for 300 cases.
- Clef-flash is quicker but not safe for security checks. It missed a third of the deployments that touched security-sensitive files.
- The question-wording lesson holds for Clef too. Asking "was this caused by a temporary problem?" instead of "would a rerun succeed?" took both Clef models to 28 out of 28, with no false alarms.
What You'll Learn
- How Clef and Clef-flash compare with Jev, Laya and an LLM on the same tests
- Why "how sure the model is" matters as much as how often it is right
- What Clef costs, and how fast it is from outside Cloudflare
- Which model I would choose for which job
The Problem
In the last two weeks, a new kind of AI model has gone from one product to a small market. TypeSafe launched Jev in September. Open alternatives followed, including Laya. Then Cloudflare released Clef, with open weights, a 27 billion parameter main model and a 9 billion parameter "flash" model, hosted on Workers AI and downloadable from Hugging Face.
These are decision models: you send some text and a list of questions with allowed answers, and they return the answers with probabilities. They cannot write a sentence. The pitch is LLM-quality decisions at a fraction of the time and cost.
Most of the comparisons published so far repeat each vendor's own benchmark. Those benchmarks disagree with each other, and none of them runs on the kind of work I care about: sorting support cases and reading failed deployments in Salesforce. I already had a test harness, test sets and results for Jev, Laya and Einstein, so adding Clef was a matter of pointing the harness at one more address.
Common Questions This Article Answers:
- Is Cloudflare's Clef better than Jev?
- Is Clef-flash good enough, or do you need the full Clef?
- How does Clef compare with an LLM such as Salesforce Einstein?
- Do the same question-writing rules apply across decision models?
Quick Answer
Clef is as accurate as Jev and Einstein, and its confidence scores are much more useful. If you plan to let the model act alone on confident answers and send the rest to a person, Clef is the better choice in my tests. Jev is twice as fast and a sixth of the price, and if a quarter of a second matters more than knowing when the model is unsure, it still wins. Avoid Clef-flash for anything where a miss is costly, such as security review. Clef's weights are open, so unlike Jev, you can run it on your own hardware if your data must not leave your systems. I used Cloudflare's hosted version for these tests.
A few words you will see
| Term | What it means here |
|---|---|
| Decision model | AI that only answers fixed questions with fixed answers. It cannot write. |
| LLM | AI that reads and writes text, like ChatGPT or Salesforce's Einstein. |
| Open-weight | The model can be downloaded and run on your own computers. Clef and Laya are open-weight; Jev is not. |
| Confidence | How sure the model says it is, from 0 to 1. Useful only if low numbers really do mean "probably wrong". |
| Caught | Of the cases that really were urgent (or temporary, or security-sensitive), how many the model flagged. |
How Clef differs from Jev and Laya
All three take the same request: some text, plus questions that each have a type (pick one, score on a scale, or yes/no) and their allowed answers. All three return a probability for every allowed answer, and none of them writes text. Behind that shared request, they are very different products.
The difference from an LLM is in how the answer is produced. An LLM writes its answer out piece by piece, and your code has to check what it wrote. A decision model reads the case once and puts a probability on every allowed answer. Here is one real case from the test set, answered both ways:
Kia ora, Can I swap the jacket for the larger one? Best, Alex
An LLM writes an answer
Einstein, given the four questions as a prompt
{"category": "returns", "urgency": 2, "mentions_health": false, "escalate": false}
Written one piece at a time. Then your code checks it:
- Is it valid JSON?
- Is "returns" one of the eight teams?
- Is urgency a whole number from 1 to 5?
- Are the last two true or false?
And there is no measure of how sure it is.
1.35 s for this case
A decision model scores every answer
Clef, given the same four questions
Which team?
How urgent, 1 to 5?
- 138%
- 255%
- 35.0%
- 4<1%
- 5<1%
- Mentions health?
- <1% yes
- Escalate now?
- <1% yes
0.48 s for this case, nothing to check
Real answers to case T000 from the test set. Both got it right. Only the decision model says how sure it is: 95% on returns, and a closer call between urgency 1 and 2. Jev and Laya answer in the same shape.
| Jev | Clef and Clef-flash | Laya | |
|---|---|---|---|
| Made by | TypeSafe AI | Cloudflare | Convai Innovations |
| Released | 15 September 2026 | 1 October 2026 | September 2026 |
| Can you download it? | No, hosted only, with a waitlist | Yes, Apache 2.0 licence | Yes, Apache 2.0 licence |
| Size | Not published | 27 billion and 9 billion parameters | 0.3 to 0.4 billion parameters, built on ModernBERT |
| Where it runs | TypeSafe's servers | Cloudflare Workers AI, or your own GPU server | Your own machine, even a laptop |
| Reads images | No | Yes, images and video | No |
| Useful without training | Yes | Yes | No, train it on your own examples first |
| Hosted price per million input tokens | $0.042 | $0.24 (Clef), $0.09 (Clef-flash) | No hosted service |
The biggest practical difference is where your case text goes, which decides whether you can use a model at all for health or other personal data:
Same request, four places to send it
Switch between the models and point at each hop to see where the case text travels, and what you have to run yourself.
Red hops leave your own systems. Only the open-weight models can avoid one.
What that means in practice:
- Jev is the simplest to start with and the cheapest per call, but your text always goes to TypeSafe, and you cannot run it anywhere else.
- Clef is the one to choose if you need open weights without training. It is large enough to be accurate out of the box, which also means running it yourself needs a proper GPU server. It is also the only one of the three that can look at a screenshot or a photo, which I did not test here. Cloudflare also offers a service for fine-tuning it on your own decisions.
- Laya is small enough to run anywhere, including next to your data on an ordinary computer, but out of the box it is not accurate enough to trust. In my earlier tests, a short training run on a few hundred past examples made the difference.
- Each one means something different by "confidence". Jev and Clef both return a confidence figure next to each answer, but they are calibrated differently, as the results below show. After my quick training run, Laya said it was 99.9% sure of almost everything. Never carry a cut-off over from one model to another.
What I tested
Nothing changed from the earlier posts except the model. Same cases, same questions word for word, one run each.
- Support cases (300): which team should handle it, how urgent it is on a scale of 1 to 5, whether it mentions health, and whether a senior person should act now.
- Failed deployments (300): what kind of failure it is, who should fix it, whether it was a temporary problem, and whether the change touched security-sensitive files.
- Real failures (7): genuine errors from deployments to a Salesforce Developer Edition org.
Clef accepts the same request format as Jev. The only code I wrote was the address and the key. That is a good sign for anyone worried about lock-in: you can build against one and switch to another.
Results: support cases
| Option | Right team | Urgency exactly right | Health mentions caught | Urgent cases caught | Typical time | Slowest 5% |
|---|---|---|---|---|---|---|
| Clef | 95% | 61% | 100% | 100% | 0.52 s | 1.07 s |
| Clef-flash | 93% | 62% | 100% | 96% | 0.35 s | 0.52 s |
| Jev | 94% | 61% | 100% | 96% | 0.26 s | 0.32 s |
| Einstein (LLM) | 94% | 55% | 100% | 97% | 1.16 s | 1.87 s |
| Laya, trained on 500 cases | 88% | 70% | 100% | 83% | 0.30 s | 0.43 s |
On "right team", the top four are within two points of each other. That gap is four or five cases out of 300, and I would not choose a model on it. Clef catching all 76 urgent cases, where the others missed two or three, matters more: a missed urgent case is the expensive mistake.
Results: failed deployments
| Option | Type of failure | Who should act | Temporary failures caught | Security-sensitive changes caught | Typical time |
|---|---|---|---|---|---|
| Clef | 92% | 90% | 68% | 100% | 0.53 s |
| Clef-flash | 88% | 91% | 46% | 66% | 0.34 s |
| Jev | 93% | 86% | 0% | 100% | 0.26 s |
| Einstein (LLM) | 93% | 97% | 100% | 99% | 1.20 s |
| Laya, trained on 500 failures | 90% | 96% | 50% | 96% | 0.39 s |
Clef-flash missed 24 of the 70 changes that touched permission sets, profiles, sharing rules or credentials. The full Clef caught all 70. If the whole point of the check is not to let a risky change slip past, the small model is the wrong one.
On the seven real failures, Clef and Clef-flash both got six right, the same as Jev and Einstein. All of them were fooled by the same misleading Salesforce error, No such column 'Triage_Urgency__c' on entity 'Case', which is really a permissions problem but reads exactly like a missing field. The difference was how sure they were. Jev was 99% sure of its wrong answer. Clef was 85% sure.
Clef knows when it is wrong
This was the result I did not expect, and it is the one that would change how I build.
In a real process you do not let the AI act on everything. You let it act when it is sure, and send the rest to a person. That only works if "not sure" actually lines up with "wrong".
| Option | Review if less sure than | Mistakes sent to a person | Correct answers also sent | Share of cases reviewed |
|---|---|---|---|---|
| Clef | 0.5 | 15 of 15 | 45 | 20% |
| Clef-flash | 0.5 | 16 of 20 | 57 | 24% |
| Jev | 0.7 | 4 of 17 | 11 | 5% |
With Clef, every wrong team choice had confidence below 0.5, and on average its wrong answers scored 0.33 against 0.76 for right ones. Set the cut-off at 0.5, and a person checks one case in five and sees every mistake. Jev's wrong answers scored 0.87 on average, close to its right ones, so no cut-off separates them cleanly. Raising Jev's cut-off just sends more correct answers to review without catching many more mistakes.
So "accuracy" alone undersells Clef. A model that is right 95% of the time and tells you which 5% to check is more useful than one that is right 94% of the time and equally sure about everything.
As in the earlier posts, set the cut-off separately for each model. The 0.7 I used for Jev would send 116 of Clef's 300 answers to review.
Ask about what happened, not what will happen
In the deployment post, Jev never flagged a failure as temporary when asked "Would running the same deployment again probably succeed?". It answered as if asked for odds, and kept every answer below a half. Rewording it as a fact, "The failure was caused by a temporary platform or environment problem rather than by the change itself", fixed it.
Clef was less affected by the original wording, but still affected:
| Option | "Would a rerun succeed?" | "Was it caused by a temporary problem?" | False alarms (reworded) |
|---|---|---|---|
| Clef | 19 of 28 | 28 of 28 | 0 of 272 |
| Clef-flash | 13 of 28 | 28 of 28 | 0 of 272 |
| Jev | 0 of 15 | 15 of 15 | 0 of 15 |
The Jev figures are from the 30-failure retest in the earlier post. The Clef figures cover all 300 failures.
Three different models, the same fix. This is not a quirk of one product. Ask decision models what is true about the thing in front of them, then decide in your own process what to do about it.
Cost and speed
| Option | Price per million input tokens | Cost of 300 support cases | Per 1,000 cases |
|---|---|---|---|
| Jev | $0.042 | $0.008 | about 3 cents |
| Clef-flash | $0.09 | $0.018 | about 6 cents |
| Clef | $0.24 | $0.048 | about 16 cents |
Each case was about 670 tokens for every model. Output is free for all three.
About speed. Cloudflare's own benchmark put Clef well ahead of Jev. From my machine, over the public internet, it was the other way round: Clef typically took 0.52 seconds and Jev 0.26. Both figures include the trip to each provider and back, and a model called from inside a Cloudflare Worker would avoid most of that trip. Treat my numbers as what a business system calling over the internet will see, not as the model's own speed. Even so, Clef was more than twice as fast as Einstein.
Which one would I choose?
- Support case routing with a person checking the unsure ones: Clef. It caught every urgent case and its confidence tells you which answers to check.
- High-volume decisions where speed and price matter most and a few silent mistakes are acceptable: Jev.
- Security checks in a deployment pipeline: Clef or Jev, both caught every risky change. Not Clef-flash.
- Data that must not leave your own systems: Clef or Laya, run yourself. Laya runs on a laptop but needs training first. Clef needs a proper GPU server, but was accurate without any training. I did not self-host Clef for these tests.
- Anything that has to be written, such as a reply or a summary: none of these. Use an LLM.
Before you rely on these numbers
- Practice cases. The 600 test items were generated from a limited set of sentence patterns. They compare the models fairly but flatter all of them. Test on your own data.
- One run each, from one place. Speeds depend on where you are and where the provider runs the model.
- I did not run Clef from inside Salesforce this time. The earlier posts measured Jev and Einstein inside the org. Clef takes the same request, so the same Named Credential pattern should work, but I have not timed it there.
- Clef sends your text to Cloudflare when you use Workers AI. The open weights are what make self-hosting possible, not the hosted service.
Frequently Asked Questions
Q: Is Clef better than Jev?
A: On accuracy they are level. Clef is better at telling you when it might be wrong, and it caught every urgent case. Jev is about twice as fast and a sixth of the price. Which one is better depends on whether a person reviews the unsure answers.
Q: Should I use Clef or Clef-flash?
A: Clef, unless you have tested Clef-flash on your own data and its misses are acceptable. In my tests Clef-flash missed a third of security-sensitive changes and four of its 20 wrong team choices looked confident.
Q: Can I switch from Jev to Clef without rewriting anything?
A: Nearly. Clef accepted my Jev-format questions unchanged. You change the address, the key, and the model name. Then reset your confidence cut-off, because the two models measure confidence differently.
Q: Do I need to train Clef like Laya?
A: Not for these tasks. Clef was accurate straight away. Laya was unsafe until trained on a few hundred examples.
Q: Is Clef free?
A: The weights are free under the Apache 2.0 licence. Running them yourself means paying for a capable GPU server. Cloudflare's hosted version costs $0.24 per million input tokens for Clef and $0.09 for Clef-flash.
Key Takeaways
- The top decision models are now level on accuracy. Choose on speed, price, confidence and where your data goes.
- Check how well confidence works, not just accuracy. Clef's low scores reliably flagged its mistakes; Jev's did not.
- Smaller is not always good enough. Clef-flash was faster but missed too many security-sensitive changes.
- Write questions as facts, not predictions. The fix worked on all three models I tested it on.
- Keep the provider a setting. Adding Clef took an address and a key.
What's Next?
Recommended Reading:
- Salesforce Case Triage With Decision Models: Jev and Laya Against an LLM, Measured
- Why Did My Salesforce Deployment Fail? Letting Jev, Laya and Einstein Read the Error, Measured
Action Items:
- Take a few hundred past decisions with known right answers from your own system.
- Run them through two decision models, and for each one check how many of its mistakes fall below a confidence cut-off.
- Rewrite any question that asks the model to predict the future as a question about what already happened.
Tried Clef on your own data? Leave a comment with what you found.
Responses
Checking your session.
Loading responses.