A sun with rays: the one bold element on every page of this site

Glossary

Every term this site uses, in plain words, with an example or a number beside it. It is written to be read on its own, without the other pages.

A low band of cloud along the horizon

Read an entry's first line for the plain explanation and its second for the concrete case — a number, an example, or why the term matters here. The map below puts the whole vocabulary of the site on one picture, so the terms can be met in relation to each other rather than one at a time. Terms that appear on another page link to it.

The layout of one Laya request as a strip of token boxes, labelled with the head, the markers, the state, the context window and the boundary where truncation cuts the document's tail One request, and the words for its parts context window — everything the model reads; its size is the request's max_len head — the question type, the instructions and the option markers truncation everything past here is cut marker — a [MASK] token, one per option [CLS] [SEP] [MASK] [MASK] [SEP] choice question: which dept.? billing technical document tokens … one token — each box is one state — the document the question is about The head comes first, so when a request is too long it is the document's tail that gets cut.
The reading map: one request laid out from left to right. A token is one box. The head carries the question type, the instructions and a marker for each option, and the state — the document being asked about — comes last, inside the same context window. Everything past the dashed line is truncation. The decision head reads the model's internal state at the marker positions, which is what the entries below describe.
ablation
Switching one part of a model off to see whether it was doing anything, and comparing the two runs.
Our cross-attention head's own ablation scored 0.4717 with the new part on and 0.4717 with it switched off (p=1.0), so the part was doing nothing. See it used in What we found.
attention
The mechanism by which a token pulls in information from other tokens: its new state is a weighted blend of the states it is allowed to look at.
Attention is the step that lets a marker in the head take on information from the document. This site meets two versions of it: global attention and sliding (local) attention.
autoregressive
A model is autoregressive if it writes its answer one token at a time, with every new token conditioned on the ones it has already written.
A chatbot writing a sentence is autoregressive. Laya is not — see non-autoregressive, and How it works for what it does instead.
benchmark
A fixed set of questions with known correct answers, run the same way every time, so that two models or two versions of one model can be compared.
The long-context benchmark we inherited pads a support request with filler and always puts the request at the end. Ours moves the request through a document of fixed length, which is the axis the inherited one could not see. See What we found.
bidirectional encoder
An encoder that reads the whole sequence at once, so every token can use the tokens on both sides of it, not only the ones that came before.
This is why a marker in the head can carry the meaning of an option while the document that follows it is still being read. Contrast autoregressive models, where a token sees only its own past.
calibration (and ECE)
Calibration is whether a stated probability is true of the world: among answers given at 0.80 confidence, about 80 per cent should be right. Expected calibration error (ECE) is the average gap between stated confidence and observed accuracy, so a lower number is better.
In the published comparison, Jev is better calibrated out of the box (ECE 0.144 against Laya's 0.213), and fitting one temperature per question type brings Laya to 0.081. See Laya vs Jev.
checkpoint
A saved set of trained weights — the model at one moment in training — which you load in order to run it.
Two checkpoints appear in this research: laya (English; ModernBERT, 28 layers; default context 512 tokens) and laya-multilingual (mmBERT-base, 22 layers; default context 1024 tokens). Both support up to 8192 tokens.
choice / bool / score question
The three shapes of question Laya answers about a piece of text: a choice question picks one option from a list, a bool question answers yes or no, and a score question puts the answer on a scale.
choice: which department should handle this? bool: does the customer threaten to cancel? score: how urgent is this? Each answer comes back as a label with a probability. See How it works.
context window
The stretch of tokens a model can read at once. Anything longer has to be cut before the model sees it.
laya ships with a 512-token context and laya-multilingual with 1024; both encoders support up to 8192, which is the setting max_len controls. The failure in What we found begins with tokens that fall outside the window.
control
A second measurement run beside the real one, where the change should make no difference. If the control moves, what you measured was not the change.
Widening the sliding window moved accuracy by −0.158 (p=0.0094); narrowing it moved the same cell by the same amount, which is how we knew the hint we were chasing was noise. See p-value.
cross-attention
Attention that reads from one set of positions into another — here, a new path for the markers in the head to read the document directly, instead of only through the encoder's own layers.
We built ours to start bit-identical to the shipped model, so that any gain would belong to the design. It trained, and its ablation then showed no gain at all: 0.4717 against 0.4717, p=1.0. See What we found.
decision engine
A model that answers structured questions about a piece of text — a label, with a probability for it — rather than writing prose.
Laya is one: local, Apache-2.0, runs on your own machine with no API key, and answers choice, bool and score questions in a single forward pass. Start at the beginning.
decision head
The small network that sits on top of the encoder and turns the states at the markers into one score per option. The encoder reads; the decision head decides.
It is the part we fine-tuned: with the encoder frozen, training this head on the task gained 9.0 to 16.5 points. Do not confuse it with head, which on this site means the block of instructions and markers at the front of the input.
dilution (attention dilution)
The idea that as a document grows, the attention paid to any one sentence is spread thinner, so that sentence influences the answer less.
It is a tempting explanation for our long-document failure, and our measurement points away from it: with the answer 4000 tokens away, a linear probe recovered the label 0.82 of the time from the same states the shipped head reads. The information arrived.
encoder
The part of the model that reads the text and turns it into internal signals, one per token. It does not decide anything by itself.
laya uses a ModernBERT encoder and laya-multilingual an mmBERT-base encoder. Fine-tuning the decision head while leaving the encoder frozen was the one intervention that worked.
filler
Unrelated text used to pad a document to a target length in a benchmark.
In the benchmark we inherited, the filler is meeting notes. With max_len=8192 and the whole document present, accuracy still fell from 0.840 with no filler to 0.420 at 7000 filler tokens. See What we found.
fine-tuning
Continuing to train a model, or part of one, on your own labelled examples, starting from the shipped weights.
Fine-tuning the existing decision head on the task, with the encoder frozen, gained 9.0 to 16.5 points and reproduced across runs. That is the result behind the conclusion that the fix is adaptation, not design.
forward pass
One trip through the model, from input tokens to output. Laya answers every question in a request in a single forward pass.
Upstream's published latency for one question is 39.5 ms on laya and 32.8 ms on laya-multilingual, measured on a Tesla T4. See non-autoregressive for why one pass is enough.
global attention
A layer with global attention lets every token look at every other token inside the context window.
In laya-multilingual, 8 of the 22 layers are global and the other 14 see only ±64 tokens. Making every layer global collapsed accuracy to 0.000 at one document length, so the alternation is load-bearing. See What we found.
head
In the layout of a request, the head is the block at the front of the input: the question type, the instructions, and one marker in front of each option. The document comes after it.
A four-option head is small — we measured 45 tokens — but it is not free: at max_len=1024 it leaves about 978 tokens for the document. See marker and max_len and head_max_len, and the reading map above.
Jev
A closed, hosted decision API that answers the same kind of structured questions as Laya, over the network, at about $0.042 per 1M tokens.
It supports up to 255 options per question, where Laya's option budget is smaller, so Jev is the better choice for 50+ options without tuning. Third parties measured Jev at 236–276 ms p50, against upstream's published 39.5 ms and 32.8 ms for Laya on a Tesla T4. See Laya vs Jev.
Laya
The subject of this site: a local decision engine, Apache-2.0, that runs on your own machine with no API key and answers structured questions about a piece of text in one forward pass.
Upstream's published latency is 39.5 ms for one question on laya (English) and 32.8 ms on laya-multilingual, measured on a Tesla T4 — roughly 6–7× faster than Jev, which a third party measured at 236–276 ms p50.
layers
A transformer is a stack of blocks, each of which passes the sequence through attention and then through a small network. The number of blocks is the model's depth.
laya has 28 layers and laya-multilingual has 22, of which 8 read the whole document and 14 read only ±64 tokens around each position.
linear probe
A deliberately simple model — one linear layer — trained on a network's internal states to test whether a piece of information is present in them. It is far too weak to do the network's work for it, which is the point.
With the answer 4000 tokens away, and on templates it had not seen, our probe recovered the label 0.82 of the time from the states the shipped head reads; the shipped head scored 0.30 on the same states. See What we found.
logit
The raw, unbounded score the model gives an option before anything turns it into a probability. Higher means more likely; the number itself is not yet a probability.
The decision head produces one logit per marker, and softmax turns that set of logits into a probability vector.
majority-class rate
The accuracy you get by always answering the commonest label in the data — the score to beat before a model can be said to be answering at all.
In our long-document cells it is 0.450, and the interior of a 4000-token document scores close to it. A score at the majority-class rate is not a weak answer; it is close to no answer.
marker
A [MASK] token placed immediately in front of each option in the head. The model's internal state at that position is what the decision head reads, and one option means one marker.
In the layout [MASK] billing [MASK] technical, the answer is the marker whose state scores highest. See the reading map above.
max_len and head_max_len
max_len is the cap on the total tokens in one request — head, document and separators together. head_max_len is the slice of that budget reserved for the head.
laya-multilingual ships with max_len=1024 and both checkpoints reach 8192. head_max_len is a cap, not the head's real length: a four-option head comes out at about 45 tokens, so about 978 of the 1024 are left for the document. See What we found.
ModernBERT / mmBERT
The two encoder families the checkpoints are built on: laya uses ModernBERT, and laya-multilingual uses mmBERT-base, its multilingual counterpart.
ModernBERT is 28 layers in this checkpoint; mmBERT-base is 22, of which 8 see the whole document and 14 see only ±64 tokens. See encoder and layers.
non-autoregressive
Producing the whole answer in one pass, instead of one token at a time.
Laya is non-autoregressive: it reads the request and returns a label with a probability, and there is no generated sentence to parse. See How it works.
option budget
How many tokens a question's options may take up inside the head. When the options overflow it, each one is trimmed to fit, and labels that look similar start to look the same to the model.
Jev supports up to 255 options per question; Laya's budget is smaller, so Jev is better for 50+ options without tuning. See Laya vs Jev.
p-value
The probability of seeing an effect at least as large as the one you saw, if the change did nothing. A small p-value means the effect is unlikely to be noise; it says nothing about how large or how useful the effect is.
The sliding-window change moved accuracy by −0.158 with p=0.0094: a real effect, in the wrong direction. See control for the second measurement that made sense of it.
position sensitivity
How much a model's answer depends on where in the document the evidence sits, with the document's length held fixed.
At a fixed 4000-token document, moving the request through it changed accuracy by 0.40–0.45; at 7000 tokens the best position flipped from the end to the start. See What we found.
probability vector
The model's answer in full: one probability per option, all of them positive, adding up to 1.
Laya reports the largest entry as the answer, and that vector is what a confidence threshold is applied to. See softmax for how the vector is made from the logits, and calibration for whether its numbers can be trusted.
RLCD
Upstream's name for the training recipe behind Laya: reinforcement learning against strictly proper scoring rules, so that the probabilities the model produces are meaningful rather than only ranked.
It is why a Laya probability can be compared against a threshold at all, and why fitting one temperature is enough to improve calibration. The four letters are upstream's shorthand; this page describes the recipe rather than expanding them.
RoPE
Rotary position embedding: the way a transformer tells its tokens where they are, by rotating their vectors by an amount that depends on the token's position in the sequence.
Stretching RoPE is the usual trick for reading beyond a trained context. We tested it and it could not help: inside the window the model was trained on, the change is mathematically a no-op. See What we found.
sliding (local) attention
Attention that lets a token see only a small neighbourhood around its own position, instead of the whole context window.
In laya-multilingual, 14 of the 22 layers are local with a reach of ±64 tokens, and the other 8 are global. Widening the local window cost 0.158 (p=0.0094) rather than gaining anything.
softmax
The function that turns a list of raw scores into probabilities: every value becomes positive, and the whole list adds up to 1.
The decision head applies softmax to the logits of the markers, and what comes out is the probability vector over the options. See temperature for how that distribution can be reshaped.
state
The internal vector a token carries after the model's layers have read the sequence. The state at a marker is what the decision head turns into an answer.
The word is also used for the text itself: the document Laya is asked about is called the state in the API, which is why request_tokens_kept = 0.00 means that none of it survived into the input.
System 1 and System 2
System 1 is an answer that arrives in one glance, without deliberation. System 2 is the slow kind that builds up over time, the way a reasoning model works through a problem.
Laya is the sun in this site's sky: System 1, one forward pass. A reasoning LLM is a cloud: System 2, slower and layered. See System 1.
temperature
A number that rescales a model's scores before softmax: above 1 it flattens the distribution, below 1 it sharpens it. It changes how confident the model claims to be, not what it answers.
Fitting one temperature per question type took Laya's calibration error to 0.081 in the published comparison — an adjustment to the output, not a change to the model.
token
The unit of text a model reads: a short common word is usually one token, and a long or unusual word is often several. Every budget on this site is counted in tokens.
The context window, max_len and the head are all measured in tokens, which is why the same document can be readable at 8192 and cut off at 1024. See the reading map above.
truncation
Cutting the input down to the context window by dropping tokens from one end. In Laya's layout the document comes last, so it is the document's tail that goes.
At the shipped max_len=1024, a document of 1000+ tokens loses its tail, and in the inherited benchmark the tail is the request: request_tokens_kept = 0.00, so the request never reached the model. See What we found.
[CLS] and [SEP]
Two special tokens from the model's vocabulary. [CLS] opens the sequence, and [SEP] marks where one part of it ends and the next begins.
A request is laid out as [CLS] head [SEP] options [SEP] document [SEP], which is what makes the head come first and the document last.
[MASK]
The token the model's vocabulary reserves as a blank slot. Laya puts one in front of every option and reads the model's state at that position.
In [MASK] billing [MASK] technical, the model scores the states of the two blank slots rather than generating either word. That is what marker means.

Numbers at a glance

Every figure used anywhere on this site, with where it came from. Nothing here is an estimate: a figure is either something upstream published, something a third party measured, or something we measured ourselves — in which case it comes from a run on an RTX 4070 Laptop with 200 to 600 items per cell on balanced label sets.

Figures used on this site, with their source. Ranges are quoted as published or measured.
Figure Value Where it comes from
laya: encoder, layers, context ModernBERT · 28 layers · 512 tokens default · up to 8192 upstream's published figure
laya-multilingual: encoder, layers, context mmBERT-base · 22 layers · 1024 tokens default · up to 8192 upstream's published figure
Layers that read the whole document (laya-multilingual) 8 of 22; the other 14 see ±64 tokens upstream's published figure
Latency, one question on laya, Tesla T4 39.5 ms upstream's published figure
Latency, one question on laya-multilingual, Tesla T4 32.8 ms upstream's published figure
Latency, Jev, p50 236–276 ms a third party measured
Laya against Jev, one question roughly 6–7× faster arithmetic on the three latency rows above
Jev price about $0.042 per 1M tokens upstream's published figure
Options per question Jev up to 255; Laya's budget is smaller, so Jev leads at 50+ options without tuning upstream's published figure
Calibration error (ECE), raw Jev 0.144 · Laya 0.213 published comparison (Jev's figures are a third party's)
Calibration error after one temperature per question type Laya 0.081 published comparison
Jev assigning zero probability to the correct answer, DAIR Emotion 16% of examples; Laya did not published comparison
Head length, four-option question about 45–48 tokens we measured
Document room at max_len=1024 about 978 tokens we measured
request_tokens_kept for a document of 1000+ tokens at max_len=1024 0.00 we measured
Accuracy, no filler → 7000 filler tokens, whole document present 0.840 → 0.420 we measured
Accuracy by position in a fixed 4000-token document 0.850 · 0.550 · 0.650 · 0.600 · 0.950 we measured
Change caused by position alone in that sweep 0.40–0.45 we measured
Accuracy at 7000 tokens best position (the start) 0.850; worst (the end) 0.400 we measured
Majority-class rate in the long-document cells 0.450 we measured
Linear probe against the shipped head, answer 4000 tokens away 0.82 against 0.30 we measured
Widening the sliding window −0.158 (p=0.0094) we measured
Making all layers global 0.000 at one document length we measured
Cross-attention head against its own ablation 0.4717 against 0.4717 (p=1.0) we measured
Fine-tuning the decision head, encoder frozen +9.0 to +16.5 points we measured
Conditions for every figure we measured 200–600 items per cell · balanced label sets · RTX 4070 Laptop we measured