What we found
We set out to make Laya better at long documents. We ended up with a smaller fix than we wanted, a benchmark that had been blind to the real problem, and four experiments that made things worse or changed nothing.
The question we started from
A support ticket arrives with its history attached: a few thousand tokens of meeting notes, then the actual request at the end. You ask Laya one question about it — a choice question, which department should handle this? On short text Laya gets this right most of the time. On a long document it starts answering technical whatever the ticket says.
That is the failure we started from. Laya is a decision engine: you give it a piece of text and a set of structured questions, and it answers in one forward pass, without writing prose. The questions travel in a small block at the front of the input, which this site calls the head. Throughout this page, the request means the piece of text being asked about — the ticket itself, the part the answer depends on.
Long documents turn out to mean more than one thing. What follows is the order we actually learned it in, including the parts where our first explanation was wrong.
Unless a figure is labelled otherwise, every number on this page is our own measurement on an RTX 4070 Laptop: 200 to 600 items per cell, balanced label sets, and the majority-class rate — the score you get by always answering the commonest label — printed beside every accuracy figure.
The first thing we measured, and the trap in it
The first measurement looks like a verdict on long input. At the shipped max_len=1024, accuracy collapses and then stops moving: a document of 1000 tokens and a document of 7000 tokens get the same score. It reads like a model that has given up somewhere past a thousand tokens.
It is not that. Nothing has been given up, because nothing long was ever attempted. request_tokens_kept = 0.00: none of the request reached the model's input at all.
max_len is a cap on the whole input, not on the document. The input is laid out as the head first, then the document, then a final separator. The head is small, but it is not free: on a four-option question we measured it at 45 tokens, so 1024 − 45 − 1 leaves about 978 tokens for the document, and the cut falls at the end. The benchmark we inherited places the request at the very end of the padded document. Past roughly a thousand tokens, that is exactly where the cut lands.
Strip the request out and what is left is filler — pages of unrelated meeting notes. The model reads the filler and answers from habit, returning the commonest label almost every time. That is why the score is so flat: there is nothing in the input for a longer document to make worse.
The first measurement was not about reading. It was about the input being cut before the model saw it, and no amount of work on attention, position encoding or window size can recover tokens that were never in the input.
The second measurement, which is the real finding
So we raised max_len to 8192, the encoder's maximum, and checked that nothing was being cut. The collapse did not go away. With the whole document present, accuracy still falls from 0.840 with no filler to 0.420 at 7000 filler tokens (we measured). Different failure, same symptom.
Growing the document had already told us what it could. So we held the document still and moved the request through it: a fixed 4000-token document, the request placed at five positions, everything else identical.
| Position in the document | 0.00 right after the question | 0.25 | 0.50 | 0.75 | 1.00 very end |
|---|---|---|---|---|---|
| Accuracy | 0.850 | 0.550 | 0.650 | 0.600 | 0.950 |
The document never changed length. Only the position of the request moved, and that alone changes accuracy by 0.40–0.45.
At 7000 tokens the picture inverts. The best position flips to the start (0.850) and the end becomes the worst (0.400). Nothing about the request changed between those two runs except how much text sat around it.
The interior of the document sits near the majority-class rate of 0.450 — the score you get by always answering the commonest label. That is what the middle of a long document does to this model: not a weak answer, but close to no answer at all. At the two ends it reads.
The benchmark we inherited samples this axis at exactly one point. It always appends the request to the end of the filler, so every number it can produce is a length number and never a position number. At 4000 tokens the end is the best position we measured (0.950); at 7000 tokens it is the worst (0.400). A one-point sample cannot tell those apart, which is why this axis had been invisible.
Two failures now look identical from the outside — the same flat, wrong answer — and they are not the same thing at all.
max_len=1024 the input ends partway through the document, and the request — which the inherited benchmark places at the end — sits outside it, ghosted. Right: at max_len=8192 the whole document is inside the window, but the answer's position is covered. Same symptom, different disease, and only the second one can be fixed inside the model.Where the failure actually lives
The easy explanation for the second failure is that the information is gone by the time it reaches the head — diluted across a longer document, or overwritten on the way. We tested that directly instead of assuming it.
We trained a linear probe on exactly the internal signals the shipped head reads: one linear layer, a far weaker reader than the head it stands in for. With the answer 4000 tokens away, on templates it had not seen, it recovered the label 0.82 of the time. The shipped head, reading the same signals, scored 0.30 (we measured).
The information arrives. The reading step is what fails.
That is the turning point of the investigation. Before it, the target looked like the encoder: make it carry the evidence better across a long document. After it, the target is the reading step, and every later experiment can be judged by a simple question — does it help the reader, or does it only rearrange what the reader is given?
Four things we tried that did not work
Each of these was a reasonable idea, and each came with a control — a second measurement that should not move, run beside the one that should. Negative results are results, so they are written down here rather than deleted.
RoPE scaling: stretching the position encoding
RoPE is how a transformer tells its tokens where they are, and stretching it is the usual trick for making a model read further than it was trained to. We tested it and it could not have helped: inside the window the model was trained on, the change is mathematically a no-op. The rotations the model sees are the ones it always saw, so there was nothing to gain.
Widening the sliding window
Of the 22 layers in laya-multilingual, 14 see only ±64 tokens around each position and 8 see the whole document. Giving the local layers a wider reach is the obvious move if long-range reading is the problem. Widening it moved accuracy by −0.158 (p=0.0094). The control is what settled it: narrowing the same window moved the same cell by the same amount, which is how we knew the original hint was noise rather than a lever.
Making all layers see the whole document
If some layers are local, remove the restriction and let every layer read everything. At one document length accuracy collapsed to 0.000. The alternating pattern is load-bearing: the local layers are not a limitation to be removed, and when they were removed the model stopped answering at all.
A cross-attention head
We built a new mechanism for the markers to read the document directly, and built it so that it started bit-identical to the existing model: at step zero the two produced the same outputs, so any later gain would belong to the design rather than to a luckier start. It trained. Then we ran its own ablation — the same trained model with the new part switched off — and the two scored 0.4717 against 0.4717, p=1.0 (we measured). The new part was doing nothing.
The one thing that worked
Fine-tuning the existing decision head on the task, with the encoder frozen, gained +9.0 to +16.5 points, and the gain reproduced across runs (we measured). Frozen means the encoder's weights never move; only the small network that reads the markers is trained.
The fix is adaptation, not design. The encoder was already carrying the information — the probe result in the previous section says so — and what it needed was a reader trained on this task. Every architectural change we tested either did nothing, made things worse, or had to be switched off again.
Corrections we had to make to our own work
Four claims on this project were wrong, and three of them were ours. They are listed here because a result you cannot correct is not a measurement, and because each one was caught by a check that was designed to be able to fail.
- A floor that was one measurement repeated four times
- We published a value as Laya's long-document floor. It was not a floor. The four rows behind it differed only in document length, and the retained state was identical in all four: the request had been truncated away in every one of them. It was a truncation result wearing the clothes of a long-context result. What replaced it is the token-budget account above, checked with
request_tokens_keptrather than assumed. - A comparison that was our own test setup
- Our first position sweep left out a delimiter in the document construction that the published benchmark uses. It read 0.600 where upstream publishes 0.950 for the same nominal cell, and we were about to report that our numbers were worse than the published ones. The difference was our construction, not the model. What replaced it is the published construction, reproduced exactly, at which point the 0.950 cell reproduced too.
- A curve described as monotone from two points
- We sampled the two ends of the document and described the result as a decline with distance. A fuller sweep showed it is not a decline: the middle is worse than either end, and at 7000 tokens the best position is the start. Two endpoints cannot tell a decay from an edge effect, and that is precisely what our two-point version had been reporting. What replaced it is the five-position sweep in this page's second section.
- A sentence in our own pull request that a reviewer refuted
- In PR #696 we wrote that the head could exceed its token cap by only a small fixed amount. The reviewer said that was false in the direction that matters, and gave the mechanism: past a certain number of options, the floor on how much text each option keeps makes the head grow well beyond its cap — and those requests are accepted, not rejected, so a document sized from the cap gets truncated. We reproduced the counterexample with both checkpoints' own tokenisers before changing anything. What replaced it is a measured table covering both directions, where the cap understates the head for a small question and overstates it for a large one.
None of the four was caught by being careful. Each was caught by something that could have gone the other way: a hash comparison, a rerun of the published construction, a sweep that sampled the middle, and a reviewer reading a sentence as a claim.
The two pull requests
#696: a documentation correction
It fixes a statement about how much of the input is left for the document. The state budget is max_len minus the head that was actually built, not max_len - head_max_len, because head_max_len is a cap rather than the head's real length. At max_len=1024 we measured about 978 tokens of document room where the documentation implied 768. The pull request received one review. The reviewer confirmed the finding and the measured figures, and refuted a claim we had made in the same branch — the bound on head growth described above — which we verified and corrected.
#697: the position-sensitivity benchmark
It adds the measurement this page is built on: the same requests, the same construction and the same filler as the published long-context benchmark, with the request moved through a document of fixed length instead of always appended to the end. At position 1.00 it reproduces the published cell (0.950 on a 4000-token document), which is the check that the new benchmark measures the same thing as the old one.
Both pull requests were open at the time of writing. Neither changes the model: one corrects documentation, the other adds a benchmark. So the practical finding for anyone running Laya on long documents today is the one from the section above — fine-tune the decision head on your own task, and do not assume the answer at the end of a long document is being read.