Glossary
Every term this site uses, in plain words, with an example or a number beside it. It is written to be read on its own, without the other pages.
Read an entry's first line for the plain explanation and its second for the concrete case — a number, an example, or why the term matters here. The map below puts the whole vocabulary of the site on one picture, so the terms can be met in relation to each other rather than one at a time. Terms that appear on another page link to it.
- ablation
- Switching one part of a model off to see whether it was doing anything, and comparing the two runs.
- Our cross-attention head's own ablation scored 0.4717 with the new part on and 0.4717 with it switched off (p=1.0), so the part was doing nothing. See it used in What we found.
- attention
- The mechanism by which a token pulls in information from other tokens: its new state is a weighted blend of the states it is allowed to look at.
- Attention is the step that lets a marker in the head take on information from the document. This site meets two versions of it: global attention and sliding (local) attention.
- autoregressive
- A model is autoregressive if it writes its answer one token at a time, with every new token conditioned on the ones it has already written.
- A chatbot writing a sentence is autoregressive. Laya is not — see non-autoregressive, and How it works for what it does instead.
- benchmark
- A fixed set of questions with known correct answers, run the same way every time, so that two models or two versions of one model can be compared.
- The long-context benchmark we inherited pads a support request with filler and always puts the request at the end. Ours moves the request through a document of fixed length, which is the axis the inherited one could not see. See What we found.
- bidirectional encoder
- An encoder that reads the whole sequence at once, so every token can use the tokens on both sides of it, not only the ones that came before.
- This is why a marker in the head can carry the meaning of an option while the document that follows it is still being read. Contrast autoregressive models, where a token sees only its own past.
- calibration (and ECE)
- Calibration is whether a stated probability is true of the world: among answers given at 0.80 confidence, about 80 per cent should be right. Expected calibration error (ECE) is the average gap between stated confidence and observed accuracy, so a lower number is better.
- In the published comparison, Jev is better calibrated out of the box (ECE 0.144 against Laya's 0.213), and fitting one temperature per question type brings Laya to 0.081. See Laya vs Jev.
- checkpoint
- A saved set of trained weights — the model at one moment in training — which you load in order to run it.
- Two checkpoints appear in this research:
laya(English; ModernBERT, 28 layers; default context 512 tokens) andlaya-multilingual(mmBERT-base, 22 layers; default context 1024 tokens). Both support up to 8192 tokens. - choice / bool / score question
- The three shapes of question Laya answers about a piece of text: a
choicequestion picks one option from a list, aboolquestion answers yes or no, and ascorequestion puts the answer on a scale. choice: which department should handle this?bool: does the customer threaten to cancel?score: how urgent is this? Each answer comes back as a label with a probability. See How it works.- context window
- The stretch of tokens a model can read at once. Anything longer has to be cut before the model sees it.
layaships with a 512-token context andlaya-multilingualwith 1024; both encoders support up to 8192, which is the settingmax_lencontrols. The failure in What we found begins with tokens that fall outside the window.- control
- A second measurement run beside the real one, where the change should make no difference. If the control moves, what you measured was not the change.
- Widening the sliding window moved accuracy by −0.158 (p=0.0094); narrowing it moved the same cell by the same amount, which is how we knew the hint we were chasing was noise. See p-value.
- cross-attention
- Attention that reads from one set of positions into another — here, a new path for the markers in the head to read the document directly, instead of only through the encoder's own layers.
- We built ours to start bit-identical to the shipped model, so that any gain would belong to the design. It trained, and its ablation then showed no gain at all: 0.4717 against 0.4717, p=1.0. See What we found.
- decision engine
- A model that answers structured questions about a piece of text — a label, with a probability for it — rather than writing prose.
- Laya is one: local, Apache-2.0, runs on your own machine with no API key, and answers
choice,boolandscorequestions in a single forward pass. Start at the beginning. - decision head
- The small network that sits on top of the encoder and turns the states at the markers into one score per option. The encoder reads; the decision head decides.
- It is the part we fine-tuned: with the encoder frozen, training this head on the task gained 9.0 to 16.5 points. Do not confuse it with head, which on this site means the block of instructions and markers at the front of the input.
- dilution (attention dilution)
- The idea that as a document grows, the attention paid to any one sentence is spread thinner, so that sentence influences the answer less.
- It is a tempting explanation for our long-document failure, and our measurement points away from it: with the answer 4000 tokens away, a linear probe recovered the label 0.82 of the time from the same states the shipped head reads. The information arrived.
- encoder
- The part of the model that reads the text and turns it into internal signals, one per token. It does not decide anything by itself.
layauses a ModernBERT encoder andlaya-multilingualan mmBERT-base encoder. Fine-tuning the decision head while leaving the encoder frozen was the one intervention that worked.- filler
- Unrelated text used to pad a document to a target length in a benchmark.
- In the benchmark we inherited, the filler is meeting notes. With
max_len=8192and the whole document present, accuracy still fell from 0.840 with no filler to 0.420 at 7000 filler tokens. See What we found. - fine-tuning
- Continuing to train a model, or part of one, on your own labelled examples, starting from the shipped weights.
- Fine-tuning the existing decision head on the task, with the encoder frozen, gained 9.0 to 16.5 points and reproduced across runs. That is the result behind the conclusion that the fix is adaptation, not design.
- forward pass
- One trip through the model, from input tokens to output. Laya answers every question in a request in a single forward pass.
- Upstream's published latency for one question is 39.5 ms on
layaand 32.8 ms onlaya-multilingual, measured on a Tesla T4. See non-autoregressive for why one pass is enough. - global attention
- A layer with global attention lets every token look at every other token inside the context window.
- In
laya-multilingual, 8 of the 22 layers are global and the other 14 see only ±64 tokens. Making every layer global collapsed accuracy to 0.000 at one document length, so the alternation is load-bearing. See What we found. - head
- In the layout of a request, the head is the block at the front of the input: the question type, the instructions, and one marker in front of each option. The document comes after it.
- A four-option head is small — we measured 45 tokens — but it is not free: at
max_len=1024it leaves about 978 tokens for the document. See marker and max_len and head_max_len, and the reading map above. - Jev
- A closed, hosted decision API that answers the same kind of structured questions as Laya, over the network, at about $0.042 per 1M tokens.
- It supports up to 255 options per question, where Laya's option budget is smaller, so Jev is the better choice for 50+ options without tuning. Third parties measured Jev at 236–276 ms p50, against upstream's published 39.5 ms and 32.8 ms for Laya on a Tesla T4. See Laya vs Jev.
- Laya
- The subject of this site: a local decision engine, Apache-2.0, that runs on your own machine with no API key and answers structured questions about a piece of text in one forward pass.
- Upstream's published latency is 39.5 ms for one question on
laya(English) and 32.8 ms onlaya-multilingual, measured on a Tesla T4 — roughly 6–7× faster than Jev, which a third party measured at 236–276 ms p50. - layers
- A transformer is a stack of blocks, each of which passes the sequence through attention and then through a small network. The number of blocks is the model's depth.
layahas 28 layers andlaya-multilingualhas 22, of which 8 read the whole document and 14 read only ±64 tokens around each position.- linear probe
- A deliberately simple model — one linear layer — trained on a network's internal states to test whether a piece of information is present in them. It is far too weak to do the network's work for it, which is the point.
- With the answer 4000 tokens away, and on templates it had not seen, our probe recovered the label 0.82 of the time from the states the shipped head reads; the shipped head scored 0.30 on the same states. See What we found.
- logit
- The raw, unbounded score the model gives an option before anything turns it into a probability. Higher means more likely; the number itself is not yet a probability.
- The decision head produces one logit per marker, and softmax turns that set of logits into a probability vector.
- majority-class rate
- The accuracy you get by always answering the commonest label in the data — the score to beat before a model can be said to be answering at all.
- In our long-document cells it is 0.450, and the interior of a 4000-token document scores close to it. A score at the majority-class rate is not a weak answer; it is close to no answer.
- marker
- A
[MASK]token placed immediately in front of each option in the head. The model's internal state at that position is what the decision head reads, and one option means one marker. - In the layout
[MASK] billing [MASK] technical, the answer is the marker whose state scores highest. See the reading map above. - max_len and head_max_len
max_lenis the cap on the total tokens in one request — head, document and separators together.head_max_lenis the slice of that budget reserved for the head.laya-multilingualships withmax_len=1024and both checkpoints reach 8192.head_max_lenis a cap, not the head's real length: a four-option head comes out at about 45 tokens, so about 978 of the 1024 are left for the document. See What we found.- ModernBERT / mmBERT
- The two encoder families the checkpoints are built on:
layauses ModernBERT, andlaya-multilingualuses mmBERT-base, its multilingual counterpart. - ModernBERT is 28 layers in this checkpoint; mmBERT-base is 22, of which 8 see the whole document and 14 see only ±64 tokens. See encoder and layers.
- non-autoregressive
- Producing the whole answer in one pass, instead of one token at a time.
- Laya is non-autoregressive: it reads the request and returns a label with a probability, and there is no generated sentence to parse. See How it works.
- option budget
- How many tokens a question's options may take up inside the head. When the options overflow it, each one is trimmed to fit, and labels that look similar start to look the same to the model.
- Jev supports up to 255 options per question; Laya's budget is smaller, so Jev is better for 50+ options without tuning. See Laya vs Jev.
- p-value
- The probability of seeing an effect at least as large as the one you saw, if the change did nothing. A small p-value means the effect is unlikely to be noise; it says nothing about how large or how useful the effect is.
- The sliding-window change moved accuracy by −0.158 with p=0.0094: a real effect, in the wrong direction. See control for the second measurement that made sense of it.
- position sensitivity
- How much a model's answer depends on where in the document the evidence sits, with the document's length held fixed.
- At a fixed 4000-token document, moving the request through it changed accuracy by 0.40–0.45; at 7000 tokens the best position flipped from the end to the start. See What we found.
- probability vector
- The model's answer in full: one probability per option, all of them positive, adding up to 1.
- Laya reports the largest entry as the answer, and that vector is what a confidence threshold is applied to. See softmax for how the vector is made from the logits, and calibration for whether its numbers can be trusted.
- RLCD
- Upstream's name for the training recipe behind Laya: reinforcement learning against strictly proper scoring rules, so that the probabilities the model produces are meaningful rather than only ranked.
- It is why a Laya probability can be compared against a threshold at all, and why fitting one temperature is enough to improve calibration. The four letters are upstream's shorthand; this page describes the recipe rather than expanding them.
- RoPE
- Rotary position embedding: the way a transformer tells its tokens where they are, by rotating their vectors by an amount that depends on the token's position in the sequence.
- Stretching RoPE is the usual trick for reading beyond a trained context. We tested it and it could not help: inside the window the model was trained on, the change is mathematically a no-op. See What we found.
- sliding (local) attention
- Attention that lets a token see only a small neighbourhood around its own position, instead of the whole context window.
- In
laya-multilingual, 14 of the 22 layers are local with a reach of ±64 tokens, and the other 8 are global. Widening the local window cost 0.158 (p=0.0094) rather than gaining anything. - softmax
- The function that turns a list of raw scores into probabilities: every value becomes positive, and the whole list adds up to 1.
- The decision head applies softmax to the logits of the markers, and what comes out is the probability vector over the options. See temperature for how that distribution can be reshaped.
- state
- The internal vector a token carries after the model's layers have read the sequence. The state at a marker is what the decision head turns into an answer.
- The word is also used for the text itself: the document Laya is asked about is called the state in the API, which is why
request_tokens_kept = 0.00means that none of it survived into the input. - System 1 and System 2
- System 1 is an answer that arrives in one glance, without deliberation. System 2 is the slow kind that builds up over time, the way a reasoning model works through a problem.
- Laya is the sun in this site's sky: System 1, one forward pass. A reasoning LLM is a cloud: System 2, slower and layered. See System 1.
- temperature
- A number that rescales a model's scores before softmax: above 1 it flattens the distribution, below 1 it sharpens it. It changes how confident the model claims to be, not what it answers.
- Fitting one temperature per question type took Laya's calibration error to 0.081 in the published comparison — an adjustment to the output, not a change to the model.
- token
- The unit of text a model reads: a short common word is usually one token, and a long or unusual word is often several. Every budget on this site is counted in tokens.
- The context window,
max_lenand the head are all measured in tokens, which is why the same document can be readable at 8192 and cut off at 1024. See the reading map above. - truncation
- Cutting the input down to the context window by dropping tokens from one end. In Laya's layout the document comes last, so it is the document's tail that goes.
- At the shipped
max_len=1024, a document of 1000+ tokens loses its tail, and in the inherited benchmark the tail is the request:request_tokens_kept = 0.00, so the request never reached the model. See What we found. - [CLS] and [SEP]
- Two special tokens from the model's vocabulary.
[CLS]opens the sequence, and[SEP]marks where one part of it ends and the next begins. - A request is laid out as
[CLS]head[SEP]options[SEP]document[SEP], which is what makes the head come first and the document last. - [MASK]
- The token the model's vocabulary reserves as a blank slot. Laya puts one in front of every option and reads the model's state at that position.
- In
[MASK] billing [MASK] technical, the model scores the states of the two blank slots rather than generating either word. That is what marker means.
Numbers at a glance
Every figure used anywhere on this site, with where it came from. Nothing here is an estimate: a figure is either something upstream published, something a third party measured, or something we measured ourselves — in which case it comes from a run on an RTX 4070 Laptop with 200 to 600 items per cell on balanced label sets.
| Figure | Value | Where it comes from |
|---|---|---|
laya: encoder, layers, context |
ModernBERT · 28 layers · 512 tokens default · up to 8192 | upstream's published figure |
laya-multilingual: encoder, layers, context |
mmBERT-base · 22 layers · 1024 tokens default · up to 8192 | upstream's published figure |
Layers that read the whole document (laya-multilingual) |
8 of 22; the other 14 see ±64 tokens | upstream's published figure |
Latency, one question on laya, Tesla T4 |
39.5 ms | upstream's published figure |
Latency, one question on laya-multilingual, Tesla T4 |
32.8 ms | upstream's published figure |
| Latency, Jev, p50 | 236–276 ms | a third party measured |
| Laya against Jev, one question | roughly 6–7× faster | arithmetic on the three latency rows above |
| Jev price | about $0.042 per 1M tokens | upstream's published figure |
| Options per question | Jev up to 255; Laya's budget is smaller, so Jev leads at 50+ options without tuning | upstream's published figure |
| Calibration error (ECE), raw | Jev 0.144 · Laya 0.213 | published comparison (Jev's figures are a third party's) |
| Calibration error after one temperature per question type | Laya 0.081 | published comparison |
| Jev assigning zero probability to the correct answer, DAIR Emotion | 16% of examples; Laya did not | published comparison |
| Head length, four-option question | about 45–48 tokens | we measured |
Document room at max_len=1024 |
about 978 tokens | we measured |
request_tokens_kept for a document of 1000+ tokens at max_len=1024 |
0.00 | we measured |
| Accuracy, no filler → 7000 filler tokens, whole document present | 0.840 → 0.420 | we measured |
| Accuracy by position in a fixed 4000-token document | 0.850 · 0.550 · 0.650 · 0.600 · 0.950 | we measured |
| Change caused by position alone in that sweep | 0.40–0.45 | we measured |
| Accuracy at 7000 tokens | best position (the start) 0.850; worst (the end) 0.400 | we measured |
| Majority-class rate in the long-document cells | 0.450 | we measured |
| Linear probe against the shipped head, answer 4000 tokens away | 0.82 against 0.30 | we measured |
| Widening the sliding window | −0.158 (p=0.0094) | we measured |
| Making all layers global | 0.000 at one document length | we measured |
| Cross-attention head against its own ablation | 0.4717 against 0.4717 (p=1.0) | we measured |
| Fine-tuning the decision head, encoder frozen | +9.0 to +16.5 points | we measured |
| Conditions for every figure we measured | 200–600 items per cell · balanced label sets · RTX 4070 Laptop | we measured |