TL;DR
- AI coding tools get the attention, but a Salesforce team also spends hours on small judgement calls around the code. The most familiar: a deployment failed, so what broke, and whose job is it?
- I asked three AI options four questions about 300 test failures and 7 real failures I produced in a Salesforce org: Einstein (Salesforce's built-in AI), Jev (a hosted decision model) and Laya (a free decision model I ran myself).
- Jev and Einstein were equally good at naming the type of failure (93%), and both caught every change that touched security-sensitive files. Jev was more than four times faster.
- Laya, trained for under an hour on 500 example failures, came close: 90% on the type of failure, and it never sends a log outside your own machine.
- All three were fooled by the same misleading Salesforce error, and Jev was 99% sure of its wrong answer. Some errors only make sense if you know the org.
- How you ask matters more than which AI you use. Asked "would a rerun succeed?", Jev never once said yes. Asked "was this caused by a temporary problem?", it caught all 15 temporary failures I tested, with no false alarms.
What You'll Learn
- Where AI can help in the Salesforce development process without writing a line of code
- How accurately three AI options read real Salesforce deployment errors
- A ready-to-use script that explains a failed deployment in your pipeline's summary
- How to word questions so a decision model answers them well
The Problem
Every Salesforce team knows this moment. The pipeline goes red. Somebody opens the log, scrolls past several hundred lines, finds the error, and works out what it means. Is it our code? A field that does not exist in that org yet? A test that broke? Coverage? A permission? Or the platform having a bad moment, so we run it again?
That reading takes a few minutes each time, and it is usually done by the most experienced person on the team, because they are the one who recognises the error. Multiply it by every failed run on every branch, and it is a real cost.
When people talk about AI in Salesforce development, they usually mean tools that write code. In my previous post I tested a different kind of AI, the decision model, which cannot write anything but answers fixed questions quickly and cheaply. TypeSafe, the company behind Jev, is clear that Jev is not a coding assistant. The interesting question is whether it can take over the small decisions around the code, like reading a failed deployment.
Common Questions This Article Answers:
- Can AI tell me why a Salesforce deployment failed, and who should fix it?
- Can it spot a flaky failure, so the team reruns instead of investigating?
- How do I add this to a GitHub Actions or other CI pipeline?
Quick Answer
Yes, mostly. On these tests, Jev and Einstein named the type of failure correctly 93% of the time and flagged every change touching permissions, profiles, sharing or credentials. Jev answered in about a quarter of a second. Neither can see your org, so an error message that points the wrong way fools them both.
Use it to sort and explain, not to decide alone: put its answer at the top of the pipeline summary, with the errors underneath, so a person can confirm in seconds instead of reading for minutes.
A few words you will see
If you read the previous post, the AI terms are the same. The Salesforce development ones:
| Term | What it means here |
|---|---|
| Deployment | Moving changes (code, fields, settings) into a Salesforce org. It runs tests on the way and fails if anything breaks. |
| Validate-only | A rehearsal deployment: Salesforce compiles and tests everything, then saves nothing. I used it to produce real errors without changing the org. |
| Pipeline or CI | The automated process that runs deployments and tests for every change, such as GitHub Actions. |
| Coverage | How much of your code the tests exercise. Salesforce requires 75% before it lets code into production. |
| Governor limit | A hard cap Salesforce enforces per transaction, such as 100 database queries. |
| Flaky failure | A failure caused by a temporary problem, such as two jobs updating the same record at once. Running again usually works. |
| Decision model | AI that answers fixed questions with fixed answers, quickly and cheaply, and cannot write. Jev and Laya are decision models. |
What I asked
For each failure, four questions:
| Question | Allowed answers |
|---|---|
| What kind of failure is it? | missing dependency, code does not compile, a test failed, coverage below 75%, permissions, org data or configuration, governor limit, temporary environment problem |
| Who should act on it? | developer (code change), admin (configuration, permission or data change), or nobody (run it again) |
| Would running it again probably succeed? | yes or no |
| Do the changed files touch security-sensitive areas? | yes or no: permission sets, profiles, sharing rules, classes without sharing, sites or guest access, credentials |
The last question is worth a moment. A failed deployment is a natural point to notice that a change also edits a permission set or a sharing rule, and to ask for a security review before anyone merges it.
The test failures. I generated 300 failures using Salesforce's own error wording, such as System.LimitException: Too many SOQL queries: 101, wrapped in the kind of log a pipeline produces, with the list of changed files. Training examples used different class names, object names and log formats from the test set, so a trained model could not score well by recognising names.
The real failures. Generated text only proves so much, so I also wrote seven small pieces of deliberately broken Apex and ran each as a validate-only deployment against a Salesforce Developer Edition org: a missing field, a syntax error, a failing test, low coverage, too many queries, a picklist value that does not exist, and a field the test user had no permission to see. Salesforce returned seven genuine errors, and nothing in the org changed.
Results on the 300 test failures
"Caught" means: of the failures that really were temporary, or really did touch security-sensitive files, how many the AI flagged.
| Option | Type of failure | Who should act | Temporary failures caught | Security-sensitive changes caught | Typical time | Slowest 5% |
|---|---|---|---|---|---|---|
| Einstein (Salesforce's AI) | 93% | 97% | 100% | 99% | 1.20 s | 1.86 s |
| Jev | 93% | 86% | 0% (see below) | 100% | 0.26 s | 0.44 s |
| Laya, untrained | 53% | 63% | 7% | 3% | 0.34 s | 0.40 s |
| Laya, trained on 500 failures | 90% | 96% | 50% | 96% | 0.39 s | 1.22 s* |
* Laya's slowest times reflect my laptop running short of memory during that run, not the model.
Measured inside Salesforce, Einstein typically took 0.78 seconds. Jev read about 700 tokens per failure, so all 300 cost $0.0088.
Results on the seven real failures
| Real Salesforce error | Correct answer | Einstein | Jev | Laya, trained |
|---|---|---|---|---|
Variable does not exist: Delivery_Region__c |
missing dependency | ✓ | ✓ | ✗ said compile error |
Missing return statement required return type: Integer |
does not compile | ✓ | ✓ | ✓ |
Assertion Failed: Expected: Critical, Actual: High |
a test failed | ✓ | ✓ | ✓ |
Test coverage of selected Apex Class is 25% |
coverage | ✓ | ✓ | ✓ |
Too many SOQL queries: 101 |
governor limit | ✓ | ✓ | ✓ |
INVALID_OR_NULL_FOR_RESTRICTED_PICKLIST |
data or configuration | ✓ | ✓ (said developer, not admin) | ✓ |
No such column 'Triage_Urgency__c' on entity 'Case' |
permissions | ✗ missing dependency | ✗ missing dependency, 99% sure | ✗ missing dependency |
What the numbers mean
Reading the error type: a tie, at a fifth of the time. Jev and Einstein both named the failure correctly 93% of the time, and both got six of the seven real ones. For a step that runs on every failed build, Jev's quarter-second answer matters.
"Who should act" is partly opinion. Jev agreed with my answer 86% of the time, Einstein 97%. Looking at where Jev differed, most cases are arguable: a validation rule blocking test data (I said admin, Jev said developer), or a field missing from the target org (I said developer, Jev said admin, and an admin can create the field). Einstein happened to share more of my opinions. If you use this, write down your team's rules for who owns what, and put them in the question.
All three were fooled by the same error, and confidence did not help. The seventh real failure was a test user without permission to see a field. Salesforce reports that as No such column 'Triage_Urgency__c'... If you are attempting to use a custom field, be sure to append the '__c', which reads exactly like a missing field. All three called it a missing dependency, and Jev was 99% sure. A person who knows the org can see the field exists, so it must be access. An AI reading only the message cannot. This is the strongest argument for showing the AI's answer next to the original error, not instead of it.
Security-sensitive changes: caught almost every time. Jev flagged all 70 test failures whose change touched permission sets, profiles, sharing rules, classes without sharing, sites or credentials. Einstein caught 69. That alone may justify the step in your pipeline.
Untrained Laya is not usable for this either. At 53% on failure type and 3% on security flags, it needs training first, as it did in the previous post.
Trained Laya gets close. After training on 500 example failures, Laya named the failure type correctly 90% of the time, agreed with me on who should act 96% of the time, and caught 96% of the security-sensitive changes. That is near Einstein, from a model that runs entirely on your own machine. It caught only half of the temporary failures, and on the real failures it got five of seven, mistaking a missing field for a compile error as well as falling for the misleading permissions message. Training took 48 minutes on a laptop that was short of memory; the earlier case-triage run took 18.
The biggest lesson: ask about what happened, not what will happen
Jev's 0% on temporary failures looked like a blind spot. It was not.
When I looked at the probabilities behind the yes/no answers, Jev had ranked the cases correctly: genuinely temporary failures scored 0.23 to 0.31, and everything else 0.07 to 0.20. It never went above 0.5, which is where yes starts. And that is a fair reading of the question I asked. "Would running it again probably succeed?" is a prediction, and even a record-locking error might fail twice. Jev's probabilities are meant to reflect real odds, so it said "about a one in four chance".
So I changed the question to a fact about what already happened: "Was the failure caused by a temporary platform or environment problem rather than by the change itself?" On the same 30 failures, Jev caught all 15 temporary ones, with no false alarms.
That is the practical takeaway for anyone writing questions for a decision model, whatever their background. Ask what is true about the thing in front of it, not what will happen next, and decide what to do about it in your own process.
Putting it in your pipeline
I wrote a small script that does this after a failed deployment. It reads the JSON the Salesforce CLI prints, asks the four questions (with the reworded temporary-problem question), and writes a short verdict for the pipeline summary. On the real failing test above, it produced:
A test failed (confidence 1.00). Next step: developer.
and on the governor limit failure, with a permission set in the same change:
Governor limit exceeded (confidence 1.00). Next step: developer. The change touches security-sensitive files. Ask for a security review before merging.
Every summary ends with a reminder that the answer can be confidently wrong, and lists the original errors underneath.
For developers: in GitHub Actions it is two steps. The script uses only Python's standard library, so nothing to install:
- name: Validate deployment
id: deploy
run: sf project deploy validate --source-dir force-app --target-org ci --json > deploy.json
continue-on-error: true
- name: Explain the failure
if: steps.deploy.outcome == 'failure'
env:
JEV_API_KEY: ${{ secrets.JEV_API_KEY }}
run: |
python3 ci/triage_deploy_failure.py deploy.json \
--files "$(git diff --name-only origin/main...HEAD)" >> "$GITHUB_STEP_SUMMARY"
exit 1
The checkout step needs fetch-depth: 0 for the git diff. The final exit 1 keeps the job red: the explanation is for people, not a reason to let a failed deployment through.
Where Laya fits. Deployment logs contain class names, field names and sometimes data from test records. If your organisation does not allow those to leave your network, run Laya on your own build machine instead. Its server accepts the same request as Jev, so the script works unchanged: point JEV_URL at it. Laya is free and open-weight, but it has to be trained on your own failures first.
Other places in the development process
These are ideas I have not measured. Each is the same pattern: a fixed question, a fixed set of answers, and a person who acts on it.
- Story readiness: does this user story have acceptance criteria? Does it mention data that needs a privacy review?
- Pull request review: does this change touch sharing, guest access or credentials, so it needs a second reviewer?
- Release notes: does this Salesforce release-note item affect a feature our org uses?
- AI coding assistants: does this task need the expensive model, or will the cheap one do? Decision models are fast and cheap enough to make that call before every task.
Each is worth testing on a few hundred of your own past examples before you trust it, exactly as here.
Before you rely on these numbers
- Error messages are very regular, which flatters every option. Salesforce words each error the same way every time. That makes this task easier than most, especially for a model trained on examples.
- The "correct" answers are partly my opinion, particularly for who should act. Your team's rules will differ.
- Seven real failures is a small sample. They are there to show the approach works on genuine output, not to give a precise score.
- Jev sends the log to a US-hosted service. If your logs must stay in your network, use a self-hosted model.
Frequently Asked Questions
Q: Can Jev fix the failure for me?
A: No. A decision model cannot write code or text. It tells you what kind of failure it is and who should look at it. An AI coding assistant is still the tool for fixing it.
Q: Does this replace reading the log?
A: It replaces most of the hunting. Show the AI's answer above the original errors, so a person confirms it in seconds. Do not let it decide alone, because some Salesforce errors point the wrong way.
Q: Which should I use, Jev, Einstein or Laya?
A: Jev if sending deployment logs to it is acceptable: it was as accurate as Einstein and much faster. Einstein if you would rather stay inside Salesforce, though your pipeline would have to call into an org. Laya if logs must stay in your network: trained on 500 of my example failures it reached 90% on failure type, close to the other two, but it was weaker at spotting temporary failures.
Q: How much does it cost?
A: For Jev, a fraction of a cent per hundred failures. For Laya, the cost of the machine that runs it. For Einstein, check your Salesforce AI allowance.
Q: Why did the wording of one question make such a difference?
A: Decision models give honest odds. "Would this succeed next time?" is genuinely uncertain, so the odds stayed low. "Was this caused by a temporary problem?" is a fact about the error, so the model could answer it clearly.
Key Takeaways
- AI can read failed Salesforce deployments well. Jev and Einstein both named the failure type correctly 93% of the time, and Jev did it in a quarter of a second.
- It flags security-sensitive changes almost perfectly, a cheap safety net in any pipeline.
- Some errors need knowledge of the org. All three models were fooled by the same misleading message, so show their answer beside the error, not instead of it.
- Ask about facts, not predictions. One reworded question took temporary-failure detection from 0 to 15 out of 15.
What's Next?
Recommended Reading:
- Salesforce Case Triage With Decision Models: Jev and Laya Against an LLM, Measured
- Salesforce Security Audit Tools Compared: Health Check, Code Analyzer, AuraInspector, Org Check and sf-audit
Action Items:
- Collect the error output from your last 50 failed deployments, and note what each one turned out to be.
- Write down your team's rule for who owns each kind of failure.
- Run those failures through a decision model with your four questions, and compare its answers with what really happened before adding it to the pipeline.
Tried this on your own pipeline? Leave a comment with what it got right and what fooled it.
Responses
Checking your session.
Loading responses.