TL;DR
- Most AI in a Salesforce process makes small decisions, like which queue a case goes to. A new kind of model, the decision model, is built for exactly that. It answers fixed questions and cannot write anything.
- I built one support-case process in a Salesforce org three ways and gave each the same 300 test cases: Einstein, the AI built into Salesforce; Jev, a hosted decision model; and Laya, a free decision model I ran on my own laptop.
- Jev was as accurate as Einstein and about four times faster: a quarter of a second per case against more than a second. All 300 cases cost less than one cent.
- Laya was poor until I trained it. Out of the box it missed most cases mentioning health. After 18 minutes of training on 500 example cases it caught all of them, but it still missed one urgent case in six.
- All three ran from the same Salesforce code. Switching between them was a settings change, not a rebuild. One setting, how sure the AI must be before acting alone, did need retuning for each.
What You'll Learn
- What a decision model is, in plain terms, and how it differs from the AI chat tools you already know
- How accurate, fast and cheap each option was on the same cases
- How to build the process so you can change AI provider later without rewriting it
- Four problems I hit, so you do not have to
The Problem
Picture the support inbox. A message comes in, and somebody has to decide four things: which team should handle it, how urgent it is, whether it mentions something sensitive like a health condition, and whether a senior person needs to see it right now. Doing that by hand is slow. Doing it with keyword rules ("if the email says refund, send it to billing") breaks the first time a customer says "you took my money twice".
The usual AI answer is a large language model, or LLM: the kind of AI behind ChatGPT, Claude, and Einstein in Salesforce. You describe the four questions in plain English, send the message, and ask for an answer. It works, but an LLM is built to write, and you are only asking it to pick from a list. It takes a second or more, its reply is free text that your code has to check, and you pay for every word it reads and writes.
In September 2026 a company called TypeSafe AI released Jev, a model designed only for decisions. You give it the message and the questions, each with its allowed answers, and it returns those answers with a measure of how sure it is. It cannot write a sentence, and that is the point. Within days, free alternatives appeared that answer the same kind of request, including Laya and Cloudflare's Clef.
The promise is "as good as an LLM for decisions, much faster and cheaper". I wanted to know whether that holds in a real Salesforce process, and what it takes to build.
Common Questions This Article Answers:
- Can a decision model replace an LLM for sorting and routing cases in Salesforce?
- How do Jev and Laya compare with Einstein on accuracy, speed and cost?
- How do I build this so I can switch provider later?
Quick Answer
For sorting and routing decisions, Jev matched Einstein's accuracy at about a quarter of the time and a tiny fraction of the cost in my test. If your case data is not allowed to leave your own systems, Laya works, but only after you train it on your own example cases, and it was still weaker at spotting urgent cases. Keep an LLM for anything that has to be written, such as a draft reply to the customer.
Build it so the AI provider is a setting, not part of the code. Then you can start with one and move to another without starting again.
A few words you will see
You do not need any AI background to follow this post. These are the terms that matter:
| Term | What it means here |
|---|---|
| LLM (large language model) | AI that reads and writes text, like ChatGPT or Einstein. Good at almost anything, slower and dearer. |
| Decision model | AI that only answers fixed questions with fixed answers. It cannot write, but it answers quickly and cheaply. |
| Hosted | Someone else runs the model on their servers, and your data is sent to them. Jev is hosted. |
| Open-weight | The model itself is free to download, so you can run it on your own computer or servers and your data never leaves. Laya is open-weight. |
| Training | Showing a model hundreds of examples with the right answers so it learns your situation. Salesforce already has those examples: closed cases record where they really went. |
| Confidence | How sure the model is about an answer, from 0 to 1. A low number means "I might be wrong, let a person check". |
| Latency | How long one answer takes. I report the typical time, and the time for the slowest 5% of cases. |
| Token | Roughly three quarters of a word. AI services charge per token. |
The process I built
A customer emails in, and Salesforce creates a case. The AI answers four questions:
| Question | Kind of answer | Allowed answers |
|---|---|---|
| Which team? | pick one | billing, technical, account access, delivery, returns, privacy request, complaint, general |
| How urgent? | score | 1 (no rush) to 5 (act now) |
| Does it mention health? | yes or no | so the case is treated as sensitive |
| Escalate now? | yes or no | a legal or regulator threat, a safety risk, or an outage affecting many people |
Then Salesforce acts on the answers by itself:
- If the AI is sure enough (confidence 0.7 or more), the case goes straight to the right team's queue.
- If it is unsure, the case goes to a review queue for a person to check.
- If the AI fails or gives a nonsense answer, the case is marked Failed and stays where it was.
To test fairly I made up 300 practice cases with known right answers, written in different styles, plus 2,000 more for training. No real customer messages were used. The training cases were deliberately worded differently from the test cases, so a trained model could not score well by recognising sentences it had already seen.
For developers: a record-triggered Flow calls an invocable Apex action on its asynchronous path, so the AI call never slows down saving the case and outbound calls are allowed. I used Flow rather than a second Apex trigger because the org already had a Case trigger, and Flow lets an admin see the automation and switch it off.
How I made the AI provider a setting
The idea is simple. Salesforce asks every provider the same four questions in the same words, through one standard doorway. Which provider answers is decided by a Named Credential, Salesforce's stored connection details for an outside service: an address and a key. Change the address and key, and a different AI answers. The code does not change.
For developers: every engine implements one Apex interface:
public interface TriageEngine {
TriageDecision decide(String text);
}
The four questions are stored once, as a static resource in Jev's request format, and every engine reads the same file. The Jev engine sends one request carrying all four questions:
req.setEndpoint('callout:Jev/v1/systemone');
req.setBody(JSON.serialize(new Map<String, Object>{
'state' => text, 'model' => 'jev-latest', 'questions' => TriageQuestions.asMap()
}));
Laya's server and Cloudflare's Clef accept that same request, so this code works with any of them.
The Einstein engine turns the same questions into an instruction for the LLM, calls aiplatform.ModelsAPI.createGenerations, and then has to check the reply, because an LLM can say anything: a team that does not exist, an urgency of 7, or a friendly paragraph instead of an answer. The two engines came out about the same size, 65 and 71 lines. The difference is what they guard against, not how much code they need.
Results
Same 300 cases, same questions, one run each. Percentages are the share of cases each option got right. "Caught" means: of the cases that really mentioned health, or really needed escalating, how many the AI flagged. Missing those is the costly mistake.
| Option | Right team | Urgency exactly right | Urgency within 1 | Health mentions caught | Urgent cases caught | Typical time | Slowest 5% |
|---|---|---|---|---|---|---|---|
| Einstein (Salesforce's LLM) | 94% | 55% | 98% | 100% | 97% | 1.16 s | 1.87 s |
| Jev (hosted) | 94% | 61% | 98% | 100% | 96% | 0.26 s | 0.32 s |
| Laya, untrained | 75% | 37% | 86% | 31% | 50% | 0.26 s | 0.40 s |
| Laya, trained on 500 cases | 88% | 70% | 97% | 100% | 83% | 0.30 s | 0.43 s |
These times were measured from my machine, so they include the trip across the internet. Measured inside Salesforce, Einstein typically took 0.75 seconds and Jev about 0.26 to 0.36 seconds.
Cost. Jev read about 630 tokens per case. All 300 cases cost $0.0079, or about 2.6 cents per thousand cases. Einstein ran in a free Developer Edition org; in a production org it uses your AI allowance, so check your own contract for the real price. Laya cost nothing to use, but you pay for whatever computer runs it.
Reliability. Not one call failed, for any option. Einstein never once gave an answer my checking code had to reject. I would keep that check anyway, for the day the model behind it changes.
What the numbers mean
Jev is as good as the LLM at these decisions, and much faster. It matched Einstein within a point everywhere and was better at exact urgency. Even its slowest answers came back quicker than Einstein's typical ones. When the AI runs on every single case, that steadiness matters as much as the average.
Laya straight out of the box is not safe to use. Catching under a third of health mentions and half of the urgent cases would be a liability. Its makers say so themselves: it is meant to be trained on your own examples first.
Training fixes most of that, and Salesforce already has the examples. Every closed case records the team and priority a person finally chose. I trained Laya on 500 examples for 18 minutes on an ordinary laptop (an Apple M4). Health detection went from 31% to 100%, and urgent-case detection from 50% to 83%. That still leaves one urgent case in six unflagged. I would want that fixed before trusting it.
"How sure am I" helps, but less than you might hope. When Jev was wrong, it was still fairly sure of itself: 0.87 on average, against 0.96 when right. Sending everything below 0.7 to a person caught 4 of its 17 mistakes, and also sent 11 correct cases for checking. Worth having, but not a reason to stop spot-checking.
Switching from Jev to Laya, live
To test the "provider is a setting" idea, I ran the trained Laya on my laptop, made it reachable from Salesforce for a short while, and changed only the Named Credential's address and key. Brand new cases were then sorted and routed by Laya, with no change to code, Flow or settings. Then I switched back.
Two lessons came out of it:
- The "how sure" threshold does not carry over. Laya said it was about 99.9% sure of everything, so nothing would ever go to review. Different models measure confidence differently, and my quick training had not tuned Laya's. Same code, but the threshold has to be set again for each provider. I keep it in a setting for that reason.
- Ignore the speed of that test. About 1.6 seconds was a laptop at the end of a temporary internet link, not real hosting. On the laptop itself, Laya took about 0.3 seconds per case.
Four things that broke, so they do not break for you
These are mostly for whoever builds it.
- Jev wants a model name. Most examples show only the message and the questions; Jev rejects a request without
"model": "jev-latest". Laya accepts it too, so always send it. - Salesforce quietly dropped the API key, twice. Both mistakes showed up as the same "forbidden" error. First, the key setting has to be written as one formula:
{!'Bearer ' & $Credential.Jev.ApiKey}. Second, the Named Credential needs Allow Formulas in HTTP Header switched on, or the formula is ignored. Jev answers 403 when no key arrives at all and 401 when the key is wrong. That difference is how I found it. - Tests could not see the new fields. Since API version 67, Salesforce applies field permissions to queries by default, and brand new fields have no permissions yet. The production code deliberately runs with system access, and says why in a comment. The tests' checking queries needed
WITH SYSTEM_MODEtoo. - Laya counts urgency from zero. Urgency 3 is level 2 in its training data and its answers. Mix the two up and every urgency shifts by one without any error.
Before you rely on these numbers
- Practice cases flatter everyone, and trained Laya most. My test cases came from about 38 sentence patterns. That is fine for comparing the options fairly, and nothing like a real inbox. Treat the numbers as a comparison, and test on your own closed cases before deciding.
- One run each. I did not fine-tune the Einstein instructions, and I trained Laya once, briefly. Longer training might close its gap on urgent cases.
- Jev sends your case text to a service in the United States. If your cases contain health or other personal information that must stay in your country, that is a privacy and contract question before it is a technical one. It is the best reason to consider running a model like Laya yourself.
Frequently Asked Questions
Q: Can Jev replace Einstein or Agentforce?
A: For decisions like sorting, scoring and flagging, often yes. For anything that must be written, such as a reply, a summary or an explanation, no. Use a decision model for the sorting and an LLM for the writing.
Q: Do I need to understand AI to use this?
A: No. You need to write clear questions with clear allowed answers, much as you would explain the job to a new colleague, and a set of past cases to test against.
Q: How do I call Jev from Salesforce?
A: Through a Named Credential that stores Jev's address and your key, using the formula and setting in the list above. Call it from a background step, such as a Flow's asynchronous path, so saving a case never waits for it.
Q: Is Laya free?
A: The model is free under the Apache-2.0 licence. You still pay for the computer that runs it, the time to train it, and someone to retrain it as your cases change.
Q: Which should I choose?
A: If sending case text to Jev is acceptable, start there: it is the quickest way to prove the idea. If your data must stay in your own systems, plan for a self-run model like Laya or Clef, with the training and hosting that comes with it. Either way, build it so the provider is a setting.
Key Takeaways
- Deciding is not writing. A model that only makes decisions matched Salesforce's LLM at a quarter of the time and a fraction of the cost.
- Free models need your examples. Untrained Laya was not safe; 18 minutes of training on 500 past cases fixed most of it, but not all.
- Make the provider a setting. One set of questions and a Named Credential let me switch provider live, with no rebuild.
- Reset the "how sure" threshold for each provider. The same number means different things to different models.
What's Next?
Recommended Reading:
- Salesforce Security Audit Tools Compared: Health Check, Code Analyzer, AuraInspector, Org Check and sf-audit
- Salesforce Retires the OAuth Username-Password Flow on 20 February 2027
Action Items:
- Export a few hundred closed cases with the team and priority they ended up with. That is your test set, and your training set if you run a model yourself.
- Ask your four questions of Jev and of your current approach on those cases, and compare the results before building anything.
- If you build it, write the questions down once, and keep the provider in a Named Credential.
Tried a decision model on your own Salesforce data? Leave a comment with what you found.
Responses
Checking your session.
Loading responses.