One glance, then the answer
You hand Laya a piece of text and a question with a fixed set of answers. It reads the text once — one pass, nothing written in between — and returns a probability for each answer. No API key, no round trip to somebody else's machine: it runs on yours. This site is about how that works, and about what we measured when we made it read a very long document.
What a decision engine is
A decision engine is a program that answers a question about a piece of text by choosing among answers you give it. It does not chat, and it does not write prose.
Laya takes three things: the text, a question, and the list of answers it is allowed to choose from. It gives back one of those answers with a probability attached — how confident it is, on a scale from 0 to 1.
One worked example
text I was charged twice for invoice 4411
question which team should handle this?
options billing · technical · sales · other
Laya returns billing, at 0.82. That is the whole shape of an answer: a label and a number.
There are three shapes of question. choice picks one label from a list you supply, bool answers yes or no, and score returns a number on a scale. Anything you want from Laya has to fit one of those three.
Laya is Apache-2.0 (an open licence that lets you run it, change it and pass it on), it runs on your own machine, and it needs no API key.
The sun, the cloud, and the cloud cover
Two styles of thinking already have names. Psychologists call the fast, automatic kind System 1 and the slow, deliberate kind System 2; the names come from the work of Daniel Kahneman. Catching a ball is System 1. Checking long division is System 2.
In this site's sky, System 1 is the sun: always there, one glance, no deliberation. Laya is a sun. A reasoning model such as ChatGPT is a cloud — it builds up slowly, one token at a time (a token is a word-piece, the unit these models write in), and each token is produced from the tokens before it.
So far, so tidy. Here is the part we measured. Put the evidence far from the question inside a long document and the answer stops being reachable. The document still contains it; the model's reading of the document does not. In the sky, that is cloud covering the sun.
Two more of our measurements give the shape of it. At the shipped limit of 1024 tokens, a document of 1000+ tokens loses its tail entirely: we measured request_tokens_kept = 0.00, which means the request was not in the model's input at all. And with the whole document present and room for 8192 tokens, accuracy still fell from 0.840 with no filler to 0.420 at 7000 filler tokens.
How a decision engine differs from a chatbot
The left panel is the one that saves work. A chat model writes its way to an answer, one token at a time, because that is what writing is. A decision engine has nothing to write: the answer is one label out of a list you already fixed.
System 1 builds that idea up slowly, from catching a ball to a forward pass. How it works opens the box.
What this site covers
Five more pages, in the order they are worth reading.
- System 1
- The idea, built from zero: two ways to answer anything, what a language model actually does, and how a decision engine skips the writing.
- How it works
- Inside the box: the encoder that reads the text, the decision head that turns one question into one number, and what changes when the document is long.
- Laya vs Jev
- A local decision engine against a hosted reasoning API, on latency, calibration and how many options a question may have — including where the hosted API wins.
- What we found
- The long-context study: what we measured, the four fixes that failed, and the one that worked.
- Glossary
- Every term used on this site, defined in one place.
What we did, and what we found
We took the shipped checkpoint (a checkpoint is a saved copy of a trained model) and gave it documents far longer than its default limit. The whole document was present, one request was hidden inside it, and we moved that request around on purpose. Then we tried to fix what broke.
Four fixes failed, each tested against a control. RoPE scaling (a way of stretching how far a model can look) turned out to be a mathematical no-op inside the window the model was trained on. Widening the sliding window (the short stretch of text each layer is allowed to see) made accuracy worse by 0.158, p=0.0094 — p being the chance of a result that large if the change did nothing. Making every layer see the whole document collapsed accuracy to 0.000 at one length. A cross-attention head scored exactly the same as the model without it: 0.4717 against 0.4717, p=1.0.
One fix worked: fine-tuning the existing decision head on the task with the encoder frozen, worth +9.0 to +16.5 points, reproducible across runs.
The story includes a correction. Pull request #696 was a documentation change; it came back with one review that refuted a claim we had made. We checked the counterexample, the reviewer was right, and we corrected it. The findings page has all of it, and PR #697 holds the position-sensitivity benchmark this page just showed you.