The sun: one glance, no deliberation

How Laya answers a question without writing a word

Laya takes a piece of text and a question with a fixed set of options. It gives back one probability per option. Nothing is generated, nothing is written, and there is no loop. This page is the whole mechanism.

A low band of cloud along the base of the sky

The shape of the input

Laya does not take a prompt. A prompt is open-ended: you write a sentence and hope the model works out the task. Laya takes two things instead, and both of them are structured.

The first is the state: the text to be judged. A support ticket, a product review, one paragraph of a contract. It is data, not a request.

The second is the question, which arrives with a fixed set of options, each carrying a short description. Those options are part of the input. A question without options is not a question Laya can answer.

An example request

State — the text to judge:

The tenant reports the radiator in the north bedroom has not worked since Tuesday. Third report this month.

Question — What should happen to this ticket?

The four options, each with the short description that travels with it. The description is part of the input: it tells the engine what the label means.
OptionDescription
send a repair crewBook a visit and tell the tenant when to expect it.
ask for more informationThe report is too vague to act on yet.
close as duplicateThe same fault is already being tracked.
escalate to the building managerThe fault is a safety matter or a lease matter.

The word “fixed” is doing work there. Laya never invents a label. Its answer is always one of the options you supplied, and every option is described in the request so the engine can tell them apart.

Laya's questions come in three types: choice (pick among your options), bool (true or false), and score (a value on a scale). Everything below uses choice, because it is the type that makes the machinery visible.

One more word before we build anything. A token is a chunk of text — usually a word, sometimes part of a word — that the model reads as a single unit.

How the question and the state become one sequence

Laya has no separate slot for the question and no separate slot for the document. Both are written into a single sequence, in a fixed order, and the model reads that one sequence as a single thing.

The order looks like this:

[CLS] · type · instructions · [SEP] · [MASK] opt0 · [MASK] opt1 · [MASK] opt2 · [MASK] opt3 · [SEP] · <the document> · [SEP]

Every part of that line is worth naming.

[CLS]
A marker token meaning “the sequence starts here”. It carries no word of its own. Laya reads its decision from the [MASK] positions rather than from here, so this marker is mostly a starting boundary.
type
One short marker naming the question type: choice, bool, or score. The model has to know which kind of answer it is being asked for.
instructions
A fixed sentence, the same for every request built from the same template, that tells the model how to read the options and the document. It is part of the request and it costs tokens, like everything else here.
[SEP]
A separator token. It marks where one block ends and the next begins. There are three in this sequence: after the instructions, after the last option marker, and at the very end of the document.
[MASK] opt0 … opt3
One placeholder token per option. A [MASK] is a token the model was trained to fill in; here nothing is ever filled in. Each one is an address: a fixed position whose internal state the model will read as its judgement of that one option.
<the document>
The state, in full, followed by a final separator. On laya-multilingual the whole sequence can run to 8,192 tokens; the checkpoint's default setting is 1,024.

One [MASK] per option

The obvious design would be a single answer slot and a list of labels for the model to choose between in text. Laya does not do that. It lays down one placeholder per option before it has seen the document, and each placeholder ends up holding a separate opinion about one option.

That is why the markers are the decision. The model does not weigh the options against each other in a sentence and then commit to one. It produces one judgement per position, and the positions never see each other's judgements.

It also means the number of markers is the number of options. Four options, four markers. A question with more options would have more markers; the count is set by your question, not by the model.

Why the options come before the document

Two reasons, and the second one is why this site exists.

The first is that the head stays put. The question and the markers are short and fixed, so building them first means the markers always sit at the same offsets in the sequence, whatever the document's length happens to be. A reading step that has learned to look at position N keeps working when the document grows. If the options came last, their positions would move with every document.

The second is a cost. Because the markers sit at the start and the document comes after them, a marker that needs a sentence 4,000 tokens into the document has to reach back across 4,000 tokens to get it. The layout that keeps the head stable is the same layout that makes distance matter.

In our four-option template, the whole head — [CLS], type marker, instructions, separators and the four [MASK] positions — comes to about 45–48 tokens. Hold on to that figure; it is the fixed part of every request.

The assembled sequence: a short head of instructions and option markers, then the document the head: about 45–48 tokens everything else [CLS] type instructions [SEP] [MASK] opt0 [MASK] opt1 [MASK] opt2 [MASK] opt3 [SEP] the document up to 8,192 tokens [SEP] the question type and instructions one [MASK] per option a placeholder position, not a word the state the text being judged assembled left to right, once per request
Diagram E. The sequence is one strip. The head is a fixed, short block: a start marker, the question type, the instructions, and one placeholder per option. The document is everything after it. The strip is not drawn to scale in length — the head is about 45–48 tokens, while the document can run to thousands — so the document block carries a break mark.

One pass, then four numbers

Once the sequence is assembled, the model does exactly one thing with it. Here is that one thing in four stages.

  1. The whole sequence goes in

    All of it at once, not one token after another. Nothing waits for the left-hand side to finish before the right-hand side is touched.

  2. Every position sees every other position

    Laya is an encoder — a model that reads a whole sequence and hands back an internal state for every position, rather than writing new text. Its work happens in layers (the repeated processing blocks the sequence passes through; laya-multilingual has 22 of them). Inside a layer, each position may look at every other position, in both directions at the same time. That is what bidirectional means here. It is the opposite of a chatbot model, where each position may look only backwards, at text that has already been written.

  3. The four markers are read

    After the last layer, the model does not read the sequence as text. It reads only the internal state sitting at each [MASK] position. A small learned part — the decision head, the piece that turns an internal state into an answer — holds one set of weights per option, compares the state at marker opt0 against the weights for opt0, and produces a single score for that option: a raw number, where higher means more likely.

  4. The scores become probabilities

    The four scores go through a softmax, a function that turns any set of raw scores into positive numbers adding up to 1. Four scores in, four probabilities out, one per option. Higher is more likely, and softmax gives every option a positive number, so the engine never states outright that an option is impossible.

No token is generated. Nothing is written. There is no loop, no draft, and no second guess. The model reads the sequence once and returns four numbers.

Four placeholder positions fanning out across the whole document, then converging on one probability per option the four decision points one marker per option [MASK] opt0 [MASK] opt1 [MASK] opt2 [MASK] opt3 the document the state: the text being judged, up to 8,192 tokens one probability per option opt0 opt1 opt2 opt3 these four numbers sum to 1 the answer can sit thousands of tokens to the right
Diagram F. Each placeholder has to reach the whole document to judge its option, and the document begins far to the right of it. The twelve thin arrows are the reach; the three heavier arrows converge on the single vector of four probabilities. The bar lengths in that vector only show the shape of a distribution — no particular answer is being reported.

Why this is called non-autoregressive

Autoregressive describes a model that builds its answer one piece at a time, where each piece is produced from the pieces produced before it. A chatbot writing a sentence is autoregressive: it cannot write the fifth word until it has written the first four.

Laya is non-autoregressive, and the word is almost too generous, because Laya does not produce a sequence at all. The four probabilities are computed in the same pass, and none of them is an input to any of the others. There is no first number and no last number, only four numbers that appear together.

That is why there is no loop anywhere in the machinery above. Autoregression is what makes a writing model slow: every token is a step it has to take before it can take the next one. Laya takes no steps. It looks once.

If you want the contrast in the sky language of this site: a cloud builds. A sun is simply there.

Where the long document goes, and why that is the problem

Look back at the strip. The four markers are at the far left, and the document is everything to their right.

For a marker to judge its option using a sentence deep in the document, information has to travel from that sentence back to the marker's position. In this kind of model that travel happens inside the layers, and whether one position can reach another is decided layer by layer.

On laya-multilingual, the checkpoint we ran the long-context work on, the published configuration gives only 8 of the 22 layers the whole document to look at. The other 14 layers use a sliding window: each position can see only itself and the 64 tokens on either side of it, written as ±64 tokens.

So in 14 of the 22 layers, a marker at the start and a sentence thousands of tokens into the document cannot touch each other at all. Any link between them has to be carried by the 8 layers that can see both, and has to survive the 14 that cannot.

There is a blunter version of the same problem at the front door. At the checkpoint's shipped default of 1,024 tokens, a document longer than that is cut off: the tail is not in the sequence at all, so no layer can reach it. We measured a document of 1,000+ tokens at that setting and found the request itself was no longer in the model's input (request_tokens_kept = 0.00).

Twenty-two layers drawn twice: against a long document, 8 bars cover it fully and 14 cover a sliver; against a short document, every bar covers it a layer that sees the whole document a layer that sees ±64 tokens only a long document the shipped default context of 1,024 tokens 8 of 22 layers see the whole document 14 layers see ±64 tokens only the start 1,024 tokens a short document short enough to fit inside the ±64-token window every layer sees the whole document the start the whole document
Diagram G. Each bar is one layer of laya-multilingual. On the left, a document at the shipped default length: 8 of the 22 layers see all of it, and the other 14 see only the ±64-token window around each position. On the right, the same 22 layers against a document short enough to fit inside that window, where every layer sees everything. The short bars are drawn at the window's share of a 1,024-token document, worked out from the published window size; the 8-and-14 split is the published configuration.

Set the limit to 8,192 tokens so that the whole document fits, and the problem does not go away. We measured accuracy falling from 0.840 at length 0 to 0.420 at 7,000 filler tokens, with the entire document present. And at a fixed 4,000-token document, moving the request from one position to another changed accuracy by 0.40–0.45. The mechanism on this page is the reason. The measurements are on the findings page.

What “calibrated probability” means

A probability is only useful if you can act on it. If Laya gives an option 0.82, and you collect every case where it gave an option 0.82, then that option should turn out to be right in about 82% of those cases. When that holds, the probabilities are calibrated: the stated confidence matches the observed frequency.

The measure is expected calibration error, or ECE — the average gap between how confident the engine claimed to be and how often it was right. Lower is better. A zero would mean the probabilities mean exactly what they say.

Raw scores coming out of a network are rarely calibrated on their own. The usual fix is temperature scaling: fit a single number, the temperature, that divides the raw scores before the softmax turns them into probabilities. The number is chosen by comparing stated confidence against observed accuracy on data where you already know the right answer.

Fitting a temperature is a step, not a default. The laya-multilingual checkpoint we tested ships without one, so out of the box you get its raw calibration: ECE 0.213 in our measurements. Fit one temperature per question type, and it improves to 0.081.

That is why the number matters. The moment your code branches on confidence — auto-approve the confident cases, hand the rest to a person — you are choosing a threshold on a number that either does or does not mean what it says. On the comparison page we put that number beside Jev's, including the case where Jev's raw calibration is the better one.

One pass, four numbers, no text. Everything after this page is about how far those four numbers can be trusted.