The fast kind of thinking, built up from zero

This page carries you from catching a ball to a decision engine, in six steps. Nothing here assumes you have met a token, a transformer, or a probability distribution before. Each step uses only the step before it.

Six steps, in order

The first two steps have no computers in them at all. They are about two kinds of thinking you already do. The rest of the page moves that distinction onto a machine, one step at a time.

  1. Two ways to answer anything

    Some things you answer without trying. Someone throws a ball at you and your hand is already moving; you could not explain afterwards how you judged the speed. You read a face across a room and know the mood before you have found a word for it.

    Other things will not go that way. Long division takes a page and your full attention. Planning a trip for a group, with one car and one fixed week, takes a list, some arithmetic, and several attempts before the plan holds.

    The two feel different from the inside. The first is instant and seems to cost nothing. The second is slow, effortful, and easy to get wrong — and you can feel the effort while it is happening.

  2. Now call them System 1 and System 2

    Psychologists use two names for those ends of the line. System 1 is the fast, automatic kind: it fires on its own, you cannot switch it off, and it costs you nothing. System 2 is the slow, deliberate kind: it needs attention, it works step by step, and it tires you out. Both names come from the work of Daniel Kahneman.

    They name two styles of thinking, not two organs in the brain. Everyone uses both all day, and neither one is the good one.

    The question this site cares about is narrower. When a computer has to answer something, which of the two does the task actually need? Some tasks look hard and are really one glance. Some look easy and need many steps. Getting that wrong is expensive in both directions.

  3. What an LLM like ChatGPT is

    A large language model (LLM: a program trained on a very large amount of text to predict what comes next) is a System 2 machine in the sense that matters here — it works by writing.

    It writes in tokens (a token is a word-piece, roughly a word or a piece of a word, and it is the unit these models read and write in). Give it a question and it produces one token, then the next, then the next.

    Each token is computed from everything that came before it, including the tokens the model has just written itself. That property has a name: autoregressive, meaning each output is produced from the outputs before it.

    So the loop is literal. Pick a token. Append it to the text. Read the longer text. Pick the next token from that. Repeat.

    The autoregressive loop: a token is picked, appended, and the longer sequence is read again input I was charged twice for invoice 4411 which label? pick a token the sequence so far 1 · The 2 · The charge 3 · The charge looks 4 · The charge looks like append it to the text read the longer text the next pass step n cannot start until step n − 1 has finished
    The loop, once per token. The model picks a token, appends it, reads the longer text, and picks again; the sequence inside the ring grows one token at a time, from one token to two to three to four. Because each step reads the output of the last, step n has to wait for step n − 1, and the answer is only complete when the writing stops.

    The waiting is the part that costs. Step n cannot begin until step n − 1 has finished, because the text the model reads changes every time it writes a token. Nothing about the answer is settled until the writing is done.

  4. Why that is slow, and sometimes wasteful

    Writing one token at a time is exactly the right shape for writing. It is a strange shape for deciding.

    Take a real request: is this email a refund request? The possible answers are yes and no. Nobody asking that wants a paragraph about it. They want the label, and a number they can act on.

    So ask the question back: does this decision need a paragraph? For a refund question in a support queue, no. The email either is a refund request or it is not, and the evidence is already sitting in the text the model was given.

    Every token a chat model writes before it reaches the label is work nobody asked for, and the label waits behind all of it. On a support queue, that waiting is the whole cost of the system.

  5. What a System 1 AI does instead

    Laya does not write. You give it the text, the question, and the list of answers it may choose from; it reads the whole text at once and returns a probability for every answer on the list.

    It does this in one forward pass (a forward pass is one trip through the model, from the input to the answer, with no stopping in between and nothing fed back in). One trip, one set of numbers.

    The word for this is non-autoregressive: not built from its own previous outputs. There are no previous outputs. The text goes in, the probabilities come out together, and billing does not have to be produced before technical can be considered.

    Out comes the answer with its confidence, on the same input the loop used above.

    One forward pass over the same input, returning a probability for each option at once input I was charged twice for invoice 4411 which label? every token enters together one forward pass the whole text at once no loop, no waiting probability per option billing technical sales other 0.82 0.11 0.05 0.02
    One pass over the same input. Every token of the text enters the single block together, the block runs once, and a probability comes back for each option on the list: billing 0.82, technical 0.11, sales 0.05, other 0.02. Nothing is written on the way, and nothing is fed back in.

    The answer is the four numbers in that row. The largest one wins, but the others are not thrown away: they are how you know whether to trust it.

    The English checkpoint that ships as laya (a checkpoint is a saved copy of a trained model) reads up to 512 tokens by default and 8192 at most. The multilingual checkpoint reads 1024 by default. Those limits matter a great deal once documents get long, and that is where the research on this site begins: what we found.

  6. The trade, stated honestly

    System 1 is fast and cheap. Upstream's published figure for Laya is 39.5 ms for a single question on a Tesla T4. Third parties measured a hosted reasoning API at 236–276 ms p50 (p50: the middle of the distribution, so half of the requests were faster than that) for the same job, which makes the one-glance answer roughly 6–7× faster.

    It is also local. Laya is Apache-2.0 (an open licence that lets you run it, change it and pass it on), it runs on your own machine, and there is no API key and no per-token bill.

    It is not a general reasoner, though, and this is the part worth being clear about. It answers the structured question you gave it and nothing else: it will not write the reply, plan the trip, or notice something you did not ask about. If the answer has to be composed, if the steps are not known in advance, or if the question has 50+ options to choose from, System 2 is the right tool. A hosted API such as Jev supports up to 255 options per question, where Laya's option budget is smaller without tuning.

    The honest summary: System 1 answers the question you actually asked, once, in about the time it takes you to read the text. Everything it cannot do is System 2's job.

    Next: how it works opens the box and shows the parts inside, Laya vs Jev makes the comparison in detail, and the glossary keeps every term in one place.