Skip to content

LLMs for People Who Can Count to 4 but not 5

An LLM explained with the simplest math possible.

• Nabeel Siddiqui

Plenty of people write “LLMs explained for kids.” Fine, but kids can count pretty high, so I wanted something simpler. Pretend you can only count to 4. Five and up do not exist for you. That’s all the math you get. Using only that math, can we understand what an LLM is doing? I think so.

The first thing you need to do is build two simple helpers. Neither needs to understand a word. In fact, LLMs do not understand words because they are not conscious (despite what techno-enthusiasts claim), and understanding requires self-awareness which LLMs do not have.

The first helper is a sorter. It reads a massive body of text, such as the Internet, and assigns each word a number from 1 to 4. Each word gets a different number, and no two share the same one. Words with numbers close together mean similar things in that body of text.

WordNumber
Taco1
Cat2
Dog3
Puppy4

As you can see, “Dog” and “puppy” sit at 3 and 4, right next to each other, so they mean similar things. “Cat” is a 2, still in the animal neighborhood. “Taco” gets a 1, the farthest away, because it has nothing to do with any of them. Of course, in reality, there would be a lot more words. Between “taco” at 1 and “cat” at 2, for example, you’d expect words like “burrito,” “lunch,” “dollar,” and everything in between. We only get four spots since we can only count to 4. A real model, of course, has room for as many words as it needs.

The second helper is a guesser. You show it some words, and it tries to predict the next one. For example, if you give it “I fed my hungry ___”, a good guesser might answer “dog” or “cat.” Then you show it the real next word. If it said “taco” instead, you measure how far off it was and nudge it to be a little less wrong. Then you repeat that again and again, across billions of examples, until the guesses stop being terrible.

Using our numbers. Here’s the math. The correct answer, “dog,” is a 3. If the guesser says “cat,” that’s a 2, so it missed by a little:

3−2=13 - 2 = 1

Small miss, small nudge. If it says “taco,” that’s a 1, so it missed by more:

3−1=23 - 1 = 2

Bigger miss, bigger nudge. The size of the gap is the size of the correction. The guesser has no idea what any of these words mean. It just wants that gap to be as close to 0 as possible.

The unsettling part is not that this is clever. It’s that this is enough. Run it at a large enough scale, and the guesser starts finishing your sentences better than the person next to you.

NumberGive each word a 1–4GuessPredict the next wordScoreHow far off?AdjustNudge it to be less wrong

If you want the real words

Obviously, this is a simplification, but knowing this will likely keep you ahead of 90% of people. Here’s what those pieces are actually called, if you ever want to read more.

The words aren’t squeezed into single numbers. Each chunk of text (called a token) gets a long list of numbers instead of a single number. These lists are called embeddings, and embeddings that are close together tend to represent related things. For our four words, a toy version might look like this:

WordEmbedding
Taco[4, 4, 4, 4]
Cat[2, 1, 1, 2]
Dog[3, 1, 1, 1]
Puppy[4, 1, 1, 1]

Notice that “dog” and “puppy” have almost the same list, “cat” is close behind, and “taco” looks nothing like the rest. Since this list has four numbers, this is called a four-dimensional embedding. Most real AI uses far more dimensions, often hundreds or thousands, but the idea is the same. The sorter and the guesser aren’t really separate machines either; the numbers get learned right alongside the guessing, not in a first pass of their own.

The guessing game is next-token prediction. The “how far off” number is the loss. Real training doesn’t subtract one word-number from another, though. It measures how much confidence the guesser put in the actual next word and encourages it to put more there next time. The nudge that shrinks the loss is gradient descent, and backpropagation is the bookkeeping that figures out which knob to turn. Running the loop over and over is training.

And the nudging only happens while it learns. Once it’s chatting with you, it still turns your words into numbers and predicts the next one. However, it stops updating its settings.

Same idea. More than four numbers.