The sun: one glance, no deliberation

Laya and Jev, side by side

Two engines that take the same input — a piece of text, a question, a fixed set of options — and return one probability per option. Jev is a closed hosted API. Laya is a local engine with open weights. Their trade-offs run in opposite directions, and in three places Jev is simply the better choice.

A low band of cloud along the base of the sky

What each one is

Jev

Jev is a closed hosted API — a service you call over the network, where the model itself is not published. You send the text and the question to someone else's servers and get probabilities back.

You cannot inspect the model, run it yourself, or fine-tune it, and the text you send has left your machine. In exchange there is nothing to install, nothing to size, and nothing to keep running.

Laya

Laya is a local, open-weight engine — the trained parameters are published, so you can download them and run them on your own hardware. The weights are Apache-2.0, a licence that allows commercial use and modification. There is no API key and no network call.

The text never leaves the machine, and the machine is yours to run.

Both are decision engines rather than chatbots. Neither writes prose, and both return one probability per option that sums to 1. Same job, opposite trade-offs.

The numbers, side by side

Each row names where its figures come from, and that matters here: not all of them were produced under the same conditions.

Laya figures are for the laya-multilingual checkpoint unless stated. Jev is the hosted service.
WhatLayaJev
Latency for one question 32.8 ms (upstream's published figure, Tesla T4) 236–276 ms p50 (measured by a third party)
Price Self-hosted, no per-token charge About $0.042 per 1M tokens (published price)
Options in one question Smaller budget, so descriptions get trimmed at high counts Up to 255 out of the box (published figure)
Calibration as shipped ECE 0.213 (we measured) ECE 0.144 (we measured)
Calibration after fitting a temperature ECE 0.081 (we measured, one temperature per question type) Not published
Weights Apache-2.0, run locally Closed, hosted
What you have to run The model, on your hardware Nothing

One caveat before the latency row is used for anything. The two figures were not produced under the same conditions. Laya's 32.8 ms is upstream's published measurement on a Tesla T4, a modest older GPU, and the English laya checkpoint is published at 39.5 ms. Jev's 236–276 ms p50 comes from third-party measurement of the hosted service, where the hardware is unknown and a network round trip is included. Treat the gap as a real difference in order of magnitude, not as a controlled head-to-head.

Where Laya wins

Latency: about 32.8 ms against 236–276

An p50 figure is a median: half the requests came back faster than the number and half slower.

Upstream's published figure for one question on laya-multilingual is 32.8 ms on a Tesla T4. A third party measuring Jev's hosted service reports 236–276 ms at the median. On those figures, Laya answers a single question roughly 6–7× faster.

The gap matters most when the questions arrive in bulk. A pipeline scoring a large set of documents at tens of milliseconds per document is a different piece of software from the same pipeline at hundreds of milliseconds. For one request sitting behind a human's click, both are fast enough that the wait does not register.

Bar chart on one shared millisecond axis: Laya ends just past 30 ms, Jev ends at 250 ms one question, one shared millisecond scale Laya 32.8 ms upstream's published figure, on a Tesla T4 roughly 6–7× faster on a single question Jev 236–276 ms p50, measured by a third party p50 0 50 100 150 200 250 300 milliseconds for one question, both bars on the same axis
Diagram H. Both bars are drawn to scale on one millisecond axis, so the gap is visible rather than asserted. Jev's bar is drawn at 250 ms, the midpoint of the published 236–276 ms p50 range, and that range is marked as a whisker inside it. The two measurements come from different hardware, as the table above notes.

Cost: the marginal question is close to free

Jev's published price is about $0.042 per 1M tokens. Laya has no per-token price at all: the weights are free to download and they run on hardware you already have or buy.

Self-hosting is not free, and it would be dishonest to pretend otherwise. A GPU, its power draw, and someone to keep the service up are real costs, and at low volume the hosted price can easily be the cheaper answer. The advantage shows up at volume and in predictability: once the machine exists, the price of the next question is electricity.

Openness: the text does not leave the machine

Laya's weights are published under Apache-2.0 and run locally. There is no API key, no account, and no network round trip, so the text you are judging stays on the machine that judges it. For health records, legal documents, or anything under a data-residency rule, that is often the deciding factor on its own.

Open weights also mean you can fine-tune (keep training the model on your own labelled examples). We measured +9.0 to +16.5 points from fine-tuning the existing decision head with the encoder frozen, reproducible across runs. With a closed hosted API, that option does not exist at all.

Calibration, after you fit a temperature

As shipped, Jev is the better calibrated of the two: ECE 0.144 against Laya's 0.213 in our measurements. ECE, or expected calibration error, is the average gap between the confidence the engine stated and how often it was right, so lower is better.

Fit one temperature — a single number that rescales the raw scores before they become probabilities — per question type, and Laya reaches 0.081, better than Jev's raw figure. That is the win, and it is worth being precise about what it costs: the fit needs labelled examples, it is done per question type, and it is a step you own rather than one the vendor does for you.

Where Jev wins

More than a few dozen options in one question

Jev supports up to 255 options in a single question out of the box, and taking a long list of labels costs you nothing extra. Laya's option budget is smaller, and every option's description takes room in the fixed head of the sequence described on the how it works page. Push the count up and the descriptions have to be trimmed until they stop carrying much meaning.

The published guidance is that Jev is the better choice from 50 options up without tuning. In our reading the strain begins earlier, near 20 labels. That 20 is our reading of the published figures, not a measured threshold, and the diagram below is drawn that way.

Option count from 1 to 255 on a log axis: Laya comfortable at low counts and trimming descriptions from about 20, Jev comfortable across the range to 255 about 20 the flip zone published guidance: Jev is better from 50 options up Laya comfortable descriptions get trimmed our reading puts the strain near 20 labels, not a measured threshold Jev comfortable across the range: up to 255 options 1 10 20 50 100 255 options in one question, on a log scale: equal distances are equal ratios a reading of the published figures, not a measured curve
Diagram I. This is not a measured curve. It shows the shape the published figures describe: Laya is the comfortable choice at low option counts, its option descriptions have to be trimmed as the count rises, and the recommendation flips somewhere in the shaded band — our reading puts the strain near 20 labels, while the published guidance puts Jev ahead from 50. Jev is drawn as comfortable to its published ceiling of 255 options. The horizontal axis is logarithmic, so equal distances are equal ratios rather than equal numbers of labels.

Raw calibration, before anyone fits anything

Take both engines exactly as they ship. Jev's ECE is 0.144 and Laya's is 0.213, in our measurements. If you have no labelled data to fit on, or no appetite to own that step, Jev's probabilities mean more on day one. Laya only moves ahead after the fit.

Nothing to run

Jev is a hosted product. There is no checkpoint to download, no GPU to size, no serving stack to keep alive, no version to pin, and nobody on your team has to learn any of it. For a team without machine-learning infrastructure and with a modest volume of questions, that is a substantial advantage. Self-hosting Laya is a project before it is a decision engine.

One failure worth showing, fairly

On the DAIR Emotion benchmark, Jev assigned zero probability to the correct answer on 16% of examples (we measured). Laya did not assign zero to the correct answer in that run.

Zero is a particular kind of wrong. It is not “unsure”. It is a claim that the option is impossible.

If your code branches on confidence, the two cases you care about are “the engine knows” and “the engine does not know”. A zero erases the difference. A correct option that scored zero looks exactly like an option the engine considered and ruled out, so you cannot route the hard cases to a person, because the engine never says it is unsure.

Zeros also break the arithmetic downstream. Combine several questions by multiplying probabilities, or score a whole evaluation set by the product of the probability given to the true label, and one zero takes the entire product to zero — and its logarithm to infinity. An engine that does this on 16% of examples cannot be used in that pattern at all.

Fairness, because it matters here: this is one benchmark and one measured behaviour. It does not tell you how either engine will do on your task, and it does not undo Jev's better raw calibration. It says something narrower and more useful, which is that if you plan to act on confidence, you should test for zeros before you ship.

How to choose

Five questions you can answer about your own situation. None of them need a benchmark.

How many options will one question really have?

A handful, and Laya is comfortable and fast. Dozens or hundreds, and Jev takes the list as it is while Laya needs the descriptions trimmed. That single fact decides a lot of cases on its own.

What is your latency budget?

One question behind a human's click: either engine is fast enough, and the difference is not worth an argument. A pipeline scoring thousands of documents: Laya is roughly 6–7× faster per question on the published and third-party figures, and that difference is the shape of the job.

Can the text leave the machine?

If it cannot, the comparison is over. Laya runs locally and Jev cannot. If it can, Laya's biggest structural advantage disappears and the other rows decide.

Can you label data and fine-tune?

Laya's best numbers assume work: a fitted temperature per question type for calibration, and a fine-tuned decision head for accuracy, which we measured at +9.0 to +16.5 points with the encoder frozen. If neither step is available to you, Jev's out-of-the-box calibration of 0.144 against Laya's 0.213 is the better starting point.

Will you branch on confidence?

Then calibration is the row that matters, and zeros are the trap in it. Jev assigned zero probability to the correct answer on 16% of DAIR Emotion examples in our run, and Laya did not. If your logic multiplies probabilities or takes logarithms, check that before you build on top of either engine.

There is no overall winner, and we are not going to manufacture one. Laya wins on latency, cost at volume, openness, and calibration after fitting. Jev wins on option count, calibration before fitting, and being a product rather than a project. Which list matters is a question about your constraints, not about the models.