Laya and Jev, side by side
Two engines that take the same input — a piece of text, a question, a fixed set of options — and return one probability per option. Jev is a closed hosted API. Laya is a local engine with open weights. Their trade-offs run in opposite directions, and in three places Jev is simply the better choice.
What each one is
Jev
Jev is a closed hosted API — a service you call over the network, where the model itself is not published. You send the text and the question to someone else's servers and get probabilities back.
You cannot inspect the model, run it yourself, or fine-tune it, and the text you send has left your machine. In exchange there is nothing to install, nothing to size, and nothing to keep running.
Laya
Laya is a local, open-weight engine — the trained parameters are published, so you can download them and run them on your own hardware. The weights are Apache-2.0, a licence that allows commercial use and modification. There is no API key and no network call.
The text never leaves the machine, and the machine is yours to run.
Both are decision engines rather than chatbots. Neither writes prose, and both return one probability per option that sums to 1. Same job, opposite trade-offs.
The numbers, side by side
Each row names where its figures come from, and that matters here: not all of them were produced under the same conditions.
| What | Laya | Jev |
|---|---|---|
| Latency for one question | 32.8 ms (upstream's published figure, Tesla T4) | 236–276 ms p50 (measured by a third party) |
| Price | Self-hosted, no per-token charge | About $0.042 per 1M tokens (published price) |
| Options in one question | Smaller budget, so descriptions get trimmed at high counts | Up to 255 out of the box (published figure) |
| Calibration as shipped | ECE 0.213 (we measured) | ECE 0.144 (we measured) |
| Calibration after fitting a temperature | ECE 0.081 (we measured, one temperature per question type) | Not published |
| Weights | Apache-2.0, run locally | Closed, hosted |
| What you have to run | The model, on your hardware | Nothing |
One caveat before the latency row is used for anything. The two figures were not produced under the same conditions. Laya's 32.8 ms is upstream's published measurement on a Tesla T4, a modest older GPU, and the English laya checkpoint is published at 39.5 ms. Jev's 236–276 ms p50 comes from third-party measurement of the hosted service, where the hardware is unknown and a network round trip is included. Treat the gap as a real difference in order of magnitude, not as a controlled head-to-head.
Where Laya wins
Latency: about 32.8 ms against 236–276
An p50 figure is a median: half the requests came back faster than the number and half slower.
Upstream's published figure for one question on laya-multilingual is 32.8 ms on a Tesla T4. A third party measuring Jev's hosted service reports 236–276 ms at the median. On those figures, Laya answers a single question roughly 6–7× faster.
The gap matters most when the questions arrive in bulk. A pipeline scoring a large set of documents at tens of milliseconds per document is a different piece of software from the same pipeline at hundreds of milliseconds. For one request sitting behind a human's click, both are fast enough that the wait does not register.
Cost: the marginal question is close to free
Jev's published price is about $0.042 per 1M tokens. Laya has no per-token price at all: the weights are free to download and they run on hardware you already have or buy.
Self-hosting is not free, and it would be dishonest to pretend otherwise. A GPU, its power draw, and someone to keep the service up are real costs, and at low volume the hosted price can easily be the cheaper answer. The advantage shows up at volume and in predictability: once the machine exists, the price of the next question is electricity.
Openness: the text does not leave the machine
Laya's weights are published under Apache-2.0 and run locally. There is no API key, no account, and no network round trip, so the text you are judging stays on the machine that judges it. For health records, legal documents, or anything under a data-residency rule, that is often the deciding factor on its own.
Open weights also mean you can fine-tune (keep training the model on your own labelled examples). We measured +9.0 to +16.5 points from fine-tuning the existing decision head with the encoder frozen, reproducible across runs. With a closed hosted API, that option does not exist at all.
Calibration, after you fit a temperature
As shipped, Jev is the better calibrated of the two: ECE 0.144 against Laya's 0.213 in our measurements. ECE, or expected calibration error, is the average gap between the confidence the engine stated and how often it was right, so lower is better.
Fit one temperature — a single number that rescales the raw scores before they become probabilities — per question type, and Laya reaches 0.081, better than Jev's raw figure. That is the win, and it is worth being precise about what it costs: the fit needs labelled examples, it is done per question type, and it is a step you own rather than one the vendor does for you.
Where Jev wins
More than a few dozen options in one question
Jev supports up to 255 options in a single question out of the box, and taking a long list of labels costs you nothing extra. Laya's option budget is smaller, and every option's description takes room in the fixed head of the sequence described on the how it works page. Push the count up and the descriptions have to be trimmed until they stop carrying much meaning.
The published guidance is that Jev is the better choice from 50 options up without tuning. In our reading the strain begins earlier, near 20 labels. That 20 is our reading of the published figures, not a measured threshold, and the diagram below is drawn that way.
Raw calibration, before anyone fits anything
Take both engines exactly as they ship. Jev's ECE is 0.144 and Laya's is 0.213, in our measurements. If you have no labelled data to fit on, or no appetite to own that step, Jev's probabilities mean more on day one. Laya only moves ahead after the fit.
Nothing to run
Jev is a hosted product. There is no checkpoint to download, no GPU to size, no serving stack to keep alive, no version to pin, and nobody on your team has to learn any of it. For a team without machine-learning infrastructure and with a modest volume of questions, that is a substantial advantage. Self-hosting Laya is a project before it is a decision engine.
One failure worth showing, fairly
On the DAIR Emotion benchmark, Jev assigned zero probability to the correct answer on 16% of examples (we measured). Laya did not assign zero to the correct answer in that run.
Zero is a particular kind of wrong. It is not “unsure”. It is a claim that the option is impossible.
If your code branches on confidence, the two cases you care about are “the engine knows” and “the engine does not know”. A zero erases the difference. A correct option that scored zero looks exactly like an option the engine considered and ruled out, so you cannot route the hard cases to a person, because the engine never says it is unsure.
Zeros also break the arithmetic downstream. Combine several questions by multiplying probabilities, or score a whole evaluation set by the product of the probability given to the true label, and one zero takes the entire product to zero — and its logarithm to infinity. An engine that does this on 16% of examples cannot be used in that pattern at all.
Fairness, because it matters here: this is one benchmark and one measured behaviour. It does not tell you how either engine will do on your task, and it does not undo Jev's better raw calibration. It says something narrower and more useful, which is that if you plan to act on confidence, you should test for zeros before you ship.
How to choose
Five questions you can answer about your own situation. None of them need a benchmark.
How many options will one question really have?
A handful, and Laya is comfortable and fast. Dozens or hundreds, and Jev takes the list as it is while Laya needs the descriptions trimmed. That single fact decides a lot of cases on its own.
What is your latency budget?
One question behind a human's click: either engine is fast enough, and the difference is not worth an argument. A pipeline scoring thousands of documents: Laya is roughly 6–7× faster per question on the published and third-party figures, and that difference is the shape of the job.
Can the text leave the machine?
If it cannot, the comparison is over. Laya runs locally and Jev cannot. If it can, Laya's biggest structural advantage disappears and the other rows decide.
Can you label data and fine-tune?
Laya's best numbers assume work: a fitted temperature per question type for calibration, and a fine-tuned decision head for accuracy, which we measured at +9.0 to +16.5 points with the encoder frozen. If neither step is available to you, Jev's out-of-the-box calibration of 0.144 against Laya's 0.213 is the better starting point.
Will you branch on confidence?
Then calibration is the row that matters, and zeros are the trap in it. Jev assigned zero probability to the correct answer on 16% of DAIR Emotion examples in our run, and Laya did not. If your logic multiplies probabilities or takes logarithms, check that before you build on top of either engine.
There is no overall winner, and we are not going to manufacture one. Laya wins on latency, cost at volume, openness, and calibration after fitting. Jev wins on option count, calibration before fitting, and being a product rather than a project. Which list matters is a question about your constraints, not about the models.