There are two decision models worth knowing about, and the comparison between them is more interesting than the usual "open versus closed" framing suggests.
Jev is the closed, hosted model that popularized the category. Laya is the open, local model that followed three days later. Both do the same job: they answer typed questions about a piece of text or a JSON state, returning a probability rather than generating prose.
The honest summary is this. Jev is more accurate out of the box. Laya is free, local, and trainable. Neither of those facts cancels the other out, and which one matters depends entirely on what you are doing.
This article lays out the comparison with the baseline named for every number, including the places where the published benchmarks are not actually comparable.
The short answer
Choose Jev if you want the best zero-shot accuracy, you are not going to fine-tune, your questions have wide option sets, or you do not want to own the infrastructure.
Choose Laya if your data cannot leave your machine, you need to fine-tune on your own labels, you need to run offline, or your volume is high enough that per-call pricing matters.
Choose both if you want the best of each: run Laya locally as a cheap first pass and escalate the uncertain cases to Jev. That is the subject of the next article.
Zero-shot: Jev wins clearly
This is the most important section, because it is the one most often glossed over.
Out of the box, with no training on your task, Jev is substantially more accurate. The published comparisons are consistent on this point.
Test | Laya | Jev |
|---|---|---|
Zero-shot RAG routing, multi-turn (322M base) | 0% | 61.4% |
Head-to-head, same labels, 600 held-out items | 72.2% | 87% |
Japanese business benchmark | 36.9 | 97.6 |
Phishing detection accuracy | 0.505 | 0.626 |
Phishing detection recall | 1.2% | 43.2% |
Wide option set (77 labels) | 0.425 | 0.870 |
The phishing result is the starkest, and it is worth dwelling on. Laya's base model scored 0.505 accuracy with 1.2% recall, which means it predicted almost everything negative. That is near-random behavior with a catastrophic failure to find the thing it was looking for. Applying calibration lifted accuracy to 0.611, still below Jev's raw 0.626.
The zero-shot RAG routing result is even more extreme: 0% for the base model against 61.4% for Jev. When a task is outside the model's training distribution, the base checkpoint can fail completely rather than degrade gracefully.
What this means for you: if you plan to use Laya without fine-tuning, test it on your own data before you commit. The base model is a starting point, not a finished classifier.
Fine-tuned: Laya is competitive and sometimes ahead
Here is where the picture changes.
Task | Laya fine-tuned | Jev |
|---|---|---|
Typed-decisions, hard-label accuracy | 0.766 | 0.727 |
AG News classification | 0.950 | 0.910 |
Enron spam filtering | 0.993 | — |
Phishing detection | 0.980 | — |
Jailbreak / guardrail holdout | 0.755 to 0.762 | — |
After fine-tuning on a specific task, Laya matches or beats Jev's published numbers on several of them. One developer reported fine-tuning Laya on an application-specific dataset and cutting inference latency sixfold while matching results, entirely on a laptop.
But read the caveats, because they matter more than the numbers.
The 0.766 figure came from a checkpoint fine-tuned on the benchmark's own training split. That measures how well the model learned that specific distribution, not whether it generalizes.
On the fine-tuned tasks, Laya exceeded the teacher model's self-consistency ceiling. The teacher's consistency is a ceiling on what its labels can teach, so exceeding it suggests some memorization of training noise rather than deeper understanding.
And the typed-decisions checkpoint is a specialist, fine-tuned on four specific synthetic workflows. It may perform poorly on anything else.
What this means for you: fine-tuning is where Laya's advantage lives, but the published fine-tuning numbers are not a fair comparison to Jev's zero-shot numbers. They are measuring different things.

Zero-shot: Jev leads on every published test.
Wide option sets: Jev wins clearly
This is a structural difference, not a training difference, and it is the one that will surprise people.
Laya's choice options share a fixed 256-token budget in the model's output head. Past roughly twenty options, each label gets too few tokens and accuracy falls off sharply. The model card says so directly.
On a 77-label classification task, Laya scored 0.425 against Jev's 0.870. That is not a gap you close with a better prompt. It is an architectural limit.
What this means for you: if your task has a wide answer set, you have three options. Redesign it as a cascade of narrow questions, fine-tune Laya on your specific label set, or use Jev. Do not ask Laya a forty-option question and expect it to work.
Calibration: Jev is better, and it matters
Calibration is how well the confidence number reflects reality. It is easy to overlook and it decides whether you can trust a threshold.
Metric | Laya | Jev |
|---|---|---|
Expected calibration error, typed-decisions | 0.213 | 0.144 |
Laya's probabilities are more overconfident. The temperature parameter used to tune them was fitted on training data, which means the confidence values are not well calibrated for your inputs.
There is also a documented failure that shows how bad this can get. On Bengali script, Laya's English checkpoint scored 0.080 accuracy while reporting 0.945 confidence. The model could not read the input and was extremely sure of its answer.
What this means for you: never take a confidence threshold from documentation. Measure it on your own labeled data, for both models. And if you handle multiple languages, route by script before you decide anything.
Cost and privacy: Laya wins
Laya | Jev | |
|---|---|---|
Per-call cost | zero when self-hosted | per token |
Hardware | yours | theirs |
Data leaves your machine | no | yes |
Works offline | yes | no |
Fine-tunable | yes | no |
The cost advantage is real but needs context. Self-hosting converts a per-call cost into hardware, electricity, and your own time. At low volume, the hosted model is often cheaper once you value your time honestly. At high volume, or when the privacy requirement is absolute, local wins.
The privacy advantage is the one that settles the question for many teams. If your data cannot leave your machine, there is no comparison to make.
The unfair-comparison problem
Every published benchmark between these two models has at least one of these problems, and you should assume all of them are present until proven otherwise.
Different baselines. Vendor benchmarks compare against different reference points. The model card for Laya concedes this directly.
Local inference versus a hosted API call. Speed comparisons put local GPU inference against a hosted call that includes network time. That is not a like-for-like measurement, and Jev's own speed claims have the same issue in reverse.
Benchmark-split training. A fine-tuned number measured on the benchmark's own test split is not comparable to a zero-shot number.
Self-reported results. Almost every number in this article comes from the party with an interest in the outcome. Treat them as credible order-of-magnitude evidence, not as settled facts.
Two reasonable readings exist. Independent reviewers have found metrics favoring Jev (soft accuracy 0.580 vs 0.471, calibration 0.144 vs 0.213) and metrics favoring Laya (hard-label accuracy after fine-tuning 0.766 vs 0.727). Both readings are defensible because they are measuring different things.
The only benchmark that matters is yours. Run both models on a few hundred examples you have labeled yourself, and compare them on your task.
The decision table
Your situation | Recommendation |
|---|---|
Prototyping, low volume | Jev. Do not build infrastructure to test an idea. |
Data cannot leave your machine | Laya. The comparison is over. |
Wide option sets, no fine-tuning planned | Jev. Laya's option budget is a hard limit. |
Narrow task, you have labeled data | Laya, fine-tuned. This is where it wins. |
Need offline operation | Laya. Jev requires a network. |
High volume, cost-sensitive | Laya. Per-call pricing scales linearly; local does not. |
Need the best zero-shot accuracy | Jev. |
Want to own the model | Laya. Apache 2.0. |
Unsure | Run both on your own labeled data. |
The recommendation
Run both, and route by difficulty.
Laya handles the high-volume, narrow decisions locally and for free. Jev handles the cases Laya is unsure about, where its better zero-shot accuracy and calibration are worth paying for. You get the cost profile of the local model and the accuracy of the hosted one, and you only pay for the hard cases.
That hybrid pattern is the subject of the next article, and it is the practical answer for most teams that have read this far.
If you must pick one: pick Jev if you are not going to fine-tune, and pick Laya if you are.
What to do next
Take three hundred examples from your own work, label them, and run both models on them. Compare accuracy, calibration, and the cost per correct decision.
That last metric is the one people forget. A model that is 5% more accurate but costs a hundred times more per call may still be the right choice, or it may not, and you cannot know without the number.
Read the rest of the series
This article is part of a thirteen-part series on using Laya for SEO and GEO work.
- Start here: how to use Laya for SEO and GEO
- What Laya is: the open decision model, explained
- Running Laya locally: hardware, latency, and the real cost model
- Laya for search intent classification at scale
- Fine-tuning Laya on your own SEO labels
- Laya as a local reranker for internal search and RAG
- Laya for GEO answer scoring, offline
- Laya for content audits: keep, update, merge, remove
- Laya for internal linking, and where it breaks
- Guardrails: using Laya to check your own agents
- Building a hybrid stack: Laya local, Jev cloud
- The open-model trade: what you own when you self-host
Author: Marcus Ellery, Growth Experimenter Behind 150+ SEO Tests at Auspia. Marcus writes about experiments, benchmarks, learning loops, and evidence-led growth content.




