A sun with rays: System 1, the answer in one glance

What we found

We set out to make Laya better at long documents. We ended up with a smaller fix than we wanted, a benchmark that had been blind to the real problem, and four experiments that made things worse or changed nothing.

A low band of cloud along the horizon

The question we started from

A support ticket arrives with its history attached: a few thousand tokens of meeting notes, then the actual request at the end. You ask Laya one question about it — a choice question, which department should handle this? On short text Laya gets this right most of the time. On a long document it starts answering technical whatever the ticket says.

That is the failure we started from. Laya is a decision engine: you give it a piece of text and a set of structured questions, and it answers in one forward pass, without writing prose. The questions travel in a small block at the front of the input, which this site calls the head. Throughout this page, the request means the piece of text being asked about — the ticket itself, the part the answer depends on.

Long documents turn out to mean more than one thing. What follows is the order we actually learned it in, including the parts where our first explanation was wrong.

Unless a figure is labelled otherwise, every number on this page is our own measurement on an RTX 4070 Laptop: 200 to 600 items per cell, balanced label sets, and the majority-class rate — the score you get by always answering the commonest label — printed beside every accuracy figure.

The first thing we measured, and the trap in it

The first measurement looks like a verdict on long input. At the shipped max_len=1024, accuracy collapses and then stops moving: a document of 1000 tokens and a document of 7000 tokens get the same score. It reads like a model that has given up somewhere past a thousand tokens.

It is not that. Nothing has been given up, because nothing long was ever attempted. request_tokens_kept = 0.00: none of the request reached the model's input at all.

max_len is a cap on the whole input, not on the document. The input is laid out as the head first, then the document, then a final separator. The head is small, but it is not free: on a four-option question we measured it at 45 tokens, so 1024 − 45 − 1 leaves about 978 tokens for the document, and the cut falls at the end. The benchmark we inherited places the request at the very end of the padded document. Past roughly a thousand tokens, that is exactly where the cut lands.

Strip the request out and what is left is filler — pages of unrelated meeting notes. The model reads the filler and answers from habit, returning the commonest label almost every time. That is why the score is so flat: there is nothing in the input for a longer document to make worse.

The first measurement was not about reading. It was about the input being cut before the model saw it, and no amount of work on attention, position encoding or window size can recover tokens that were never in the input.

The second measurement, which is the real finding

So we raised max_len to 8192, the encoder's maximum, and checked that nothing was being cut. The collapse did not go away. With the whole document present, accuracy still falls from 0.840 with no filler to 0.420 at 7000 filler tokens (we measured). Different failure, same symptom.

Growing the document had already told us what it could. So we held the document still and moved the request through it: a fixed 4000-token document, the request placed at five positions, everything else identical.

Accuracy by position of the request in a fixed 4000-token document, max_len=8192, whole document present (we measured; majority-class rate 0.450).
Position in the document 0.00 right after the question 0.25 0.50 0.75 1.00 very end
Accuracy 0.850 0.550 0.650 0.600 0.950

The document never changed length. Only the position of the request moved, and that alone changes accuracy by 0.40–0.45.

Line chart of accuracy against the position of the request in the document, for a 4000-token document across five positions and a 7000-token document at its two measured endpoints, with the majority-class rate of 0.450 drawn as a dashed line and the region below it shaded 1.00 0.90 0.80 0.70 0.60 0.50 0.40 0.30 majority-class rate 0.450 — the score from always answering the commonest label below this line: no better than a constant answer 0.850 0.550 0.650 0.600 0.950 0.400 4000-token document, five positions measured 7000-token document, two endpoints measured in the interior the score falls to 0.550–0.650 and at 7000 tokens the end drops to 0.400 accuracy (share answered correctly) position of the request in the document (0.00 = right after the question, 1.00 = very end) 0.00 0.25 0.50 0.75 1.00 the vertical axis starts at 0.30, not at zero
Accuracy against the position of the request, with the document length held fixed. The solid line is a 4000-token document sampled at five positions; the dotted line joins the two measured endpoints of a 7000-token document, and the space between them is not a measurement. The dashed line is the majority-class rate of 0.450 — what you score by always answering the commonest label — and the shaded band is everything at or below it. At 4000 tokens the interior stays above that band, between 0.550 and 0.650; at 7000 tokens the very end of the document falls into it, at 0.400.

At 7000 tokens the picture inverts. The best position flips to the start (0.850) and the end becomes the worst (0.400). Nothing about the request changed between those two runs except how much text sat around it.

The interior of the document sits near the majority-class rate of 0.450 — the score you get by always answering the commonest label. That is what the middle of a long document does to this model: not a weak answer, but close to no answer at all. At the two ends it reads.

The benchmark we inherited samples this axis at exactly one point. It always appends the request to the end of the filler, so every number it can produce is a length number and never a position number. At 4000 tokens the end is the best position we measured (0.950); at 7000 tokens it is the worst (0.400). A one-point sample cannot tell those apart, which is why this axis had been invisible.

Two failures now look identical from the outside — the same flat, wrong answer — and they are not the same thing at all.

Two panels comparing the two failures: on the left, a document cut off by a 1024-token window with the request outside the boundary and absent from the input; on the right, the whole document inside an 8192-token window with cloud covering the position where the answer sits Two failures, one symptom: the same constant answer max_len = 1024 the shipped default for laya-multilingual max_len = 8192 the encoder's maximum everything the model can see question the input ends here the request the answer is not in the input the request was cut off before the model read it question the answer sits here, under the cloud the answer is in the input and cannot be read out of it The left failure is about the input. The right failure is about the reading.
The two long-document failures side by side. Left: at the shipped max_len=1024 the input ends partway through the document, and the request — which the inherited benchmark places at the end — sits outside it, ghosted. Right: at max_len=8192 the whole document is inside the window, but the answer's position is covered. Same symptom, different disease, and only the second one can be fixed inside the model.

Where the failure actually lives

The easy explanation for the second failure is that the information is gone by the time it reaches the head — diluted across a longer document, or overwritten on the way. We tested that directly instead of assuming it.

We trained a linear probe on exactly the internal signals the shipped head reads: one linear layer, a far weaker reader than the head it stands in for. With the answer 4000 tokens away, on templates it had not seen, it recovered the label 0.82 of the time. The shipped head, reading the same signals, scored 0.30 (we measured).

The information arrives. The reading step is what fails.

One document with the answer at the far end, whose internal state at the marker positions feeds two different readers: a linear probe that recovers the label 0.82 of the time and the shipped head that scores 0.30 a linear probe reading it 0.82 recovers the label the shipped head reading it 0.30 returns the wrong label most of the time the state at the marker positions the same signals in both reads question the signal arrives the answer, 4000 tokens from the question the document (the state) the reader differs; the source does not
One document, the answer at the far end, and the same internal state offered to two readers. The probe is trained on these states; the shipped head has never been trained on this task, so 0.82 is not a promise of what the head could score. What the pair establishes is where the limit sits: the signal reaches the markers, and the step that turns it into an answer is what fails.

That is the turning point of the investigation. Before it, the target looked like the encoder: make it carry the evidence better across a long document. After it, the target is the reading step, and every later experiment can be judged by a simple question — does it help the reader, or does it only rearrange what the reader is given?

Four things we tried that did not work

Each of these was a reasonable idea, and each came with a control — a second measurement that should not move, run beside the one that should. Negative results are results, so they are written down here rather than deleted.

RoPE scaling: stretching the position encoding

RoPE is how a transformer tells its tokens where they are, and stretching it is the usual trick for making a model read further than it was trained to. We tested it and it could not have helped: inside the window the model was trained on, the change is mathematically a no-op. The rotations the model sees are the ones it always saw, so there was nothing to gain.

Widening the sliding window

Of the 22 layers in laya-multilingual, 14 see only ±64 tokens around each position and 8 see the whole document. Giving the local layers a wider reach is the obvious move if long-range reading is the problem. Widening it moved accuracy by −0.158 (p=0.0094). The control is what settled it: narrowing the same window moved the same cell by the same amount, which is how we knew the original hint was noise rather than a lever.

Making all layers see the whole document

If some layers are local, remove the restriction and let every layer read everything. At one document length accuracy collapsed to 0.000. The alternating pattern is load-bearing: the local layers are not a limitation to be removed, and when they were removed the model stopped answering at all.

A cross-attention head

We built a new mechanism for the markers to read the document directly, and built it so that it started bit-identical to the existing model: at step zero the two produced the same outputs, so any later gain would belong to the design rather than to a luckier start. It trained. Then we ran its own ablation — the same trained model with the new part switched off — and the two scored 0.4717 against 0.4717, p=1.0 (we measured). The new part was doing nothing.

The one thing that worked

Fine-tuning the existing decision head on the task, with the encoder frozen, gained +9.0 to +16.5 points, and the gain reproduced across runs (we measured). Frozen means the encoder's weights never move; only the small network that reads the markers is trained.

The fix is adaptation, not design. The encoder was already carrying the information — the probe result in the previous section says so — and what it needed was a reader trained on this task. Every architectural change we tested either did nothing, made things worse, or had to be switched off again.

Corrections we had to make to our own work

Four claims on this project were wrong, and three of them were ours. They are listed here because a result you cannot correct is not a measurement, and because each one was caught by a check that was designed to be able to fail.

A floor that was one measurement repeated four times
We published a value as Laya's long-document floor. It was not a floor. The four rows behind it differed only in document length, and the retained state was identical in all four: the request had been truncated away in every one of them. It was a truncation result wearing the clothes of a long-context result. What replaced it is the token-budget account above, checked with request_tokens_kept rather than assumed.
A comparison that was our own test setup
Our first position sweep left out a delimiter in the document construction that the published benchmark uses. It read 0.600 where upstream publishes 0.950 for the same nominal cell, and we were about to report that our numbers were worse than the published ones. The difference was our construction, not the model. What replaced it is the published construction, reproduced exactly, at which point the 0.950 cell reproduced too.
A curve described as monotone from two points
We sampled the two ends of the document and described the result as a decline with distance. A fuller sweep showed it is not a decline: the middle is worse than either end, and at 7000 tokens the best position is the start. Two endpoints cannot tell a decay from an edge effect, and that is precisely what our two-point version had been reporting. What replaced it is the five-position sweep in this page's second section.
A sentence in our own pull request that a reviewer refuted
In PR #696 we wrote that the head could exceed its token cap by only a small fixed amount. The reviewer said that was false in the direction that matters, and gave the mechanism: past a certain number of options, the floor on how much text each option keeps makes the head grow well beyond its cap — and those requests are accepted, not rejected, so a document sized from the cap gets truncated. We reproduced the counterexample with both checkpoints' own tokenisers before changing anything. What replaced it is a measured table covering both directions, where the cap understates the head for a small question and overstates it for a large one.

None of the four was caught by being careful. Each was caught by something that could have gone the other way: a hash comparison, a rerun of the published construction, a sweep that sampled the middle, and a reviewer reading a sentence as a claim.

The two pull requests

#696: a documentation correction

It fixes a statement about how much of the input is left for the document. The state budget is max_len minus the head that was actually built, not max_len - head_max_len, because head_max_len is a cap rather than the head's real length. At max_len=1024 we measured about 978 tokens of document room where the documentation implied 768. The pull request received one review. The reviewer confirmed the finding and the measured figures, and refuted a claim we had made in the same branch — the bound on head growth described above — which we verified and corrected.

#697: the position-sensitivity benchmark

It adds the measurement this page is built on: the same requests, the same construction and the same filler as the published long-context benchmark, with the request moved through a document of fixed length instead of always appended to the end. At position 1.00 it reproduces the published cell (0.950 on a 4000-token document), which is the check that the new benchmark measures the same thing as the old one.

Both pull requests were open at the time of writing. Neither changes the model: one corrects documentation, the other adds a benchmark. So the practical finding for anyone running Laya on long documents today is the one from the section above — fine-tune the decision head on your own task, and do not assume the answer at the end of a long document is being read.