Ask a chatbot to write a poem, debug your code or summarise a 40-page PDF and it answers in seconds. It feels like magic, but under the hood a large language model (LLM) is doing something surprisingly simple — repeated billions of times, very fast.
This guide walks through what actually happens between pressing Enter and seeing a reply. No maths required.
1. It all starts with tokens
Computers don’t read words; they read numbers. So the first thing an LLM does is chop your text into small pieces called tokens. A token is often a whole word (“cat”), sometimes part of a word (“un” + “believ” + “able”), and sometimes punctuation.
As a rough rule of thumb, 100 tokens is about 75 English words. That’s why AI tools talk about “context windows” in tokens rather than pages.
Languages written in non-Latin scripts — like Urdu, Arabic or Hindi — usually need more tokens for the same sentence. That’s one reason some models feel slower or “forget” sooner in those languages.
2. Tokens become vectors (embeddings)
Each token is converted into a long list of numbers called an embedding. Think of it as coordinates on a giant map of meaning, where similar ideas sit close together:
- “king” and “queen” are near each other
- “Python” (the language) lives near “JavaScript”, while “python” (the snake) lives near “cobra”
- “fast” and “quick” are almost neighbours
The model learns these coordinates during training — nobody types them in by hand.
3. Attention: figuring out what matters
Here’s the breakthrough that made modern AI possible. In 2017, Google researchers published a paper called “Attention Is All You Need”, introducing the Transformer architecture.
The key idea is self-attention: for every token, the model looks at every other token in the text and decides how much each one matters. In the sentence:
The trophy didn’t fit in the suitcase because it was too big.
attention helps the model work out that “it” refers to the trophy, not the suitcase. A Transformer stacks dozens of these attention layers, each one building a richer understanding of the text.
4. The only real job: predict the next token
Strip away everything else and an LLM does one thing: given some text, it predicts what token comes next. That’s it.
When you ask a question, the model:
- Reads your whole prompt
- Calculates a probability for every possible next token
- Picks one (usually a likely one, with a little randomness)
- Adds it to the text and repeats — token by token — until it decides to stop
It’s autocomplete on a colossal scale. The intelligence we perceive emerges from having learned patterns across trillions of words of books, code, websites and conversations.
5. Training: from autocomplete to assistant
Building a model like ChatGPT or Claude happens in stages:
| Stage | What happens | Result |
|---|---|---|
| Pre-training | The model reads a huge slice of the internet and learns to predict the next token | A knowledgeable but unruly “base model” |
| Fine-tuning | It’s trained on high-quality example conversations | Follows instructions and answers questions |
| Human feedback (RLHF) | People rank answers; the model learns which ones are more helpful and safe | A polished, helpful assistant |
Pre-training is the expensive part — it can take thousands of specialised chips running for months.
6. Why do LLMs “hallucinate”?
Because the model predicts plausible text, not verified text. If it has never seen the answer, it may still produce something that sounds right — a fake citation, a function that doesn’t exist, a confident but wrong date.
Always double-check facts, numbers, legal or medical information and code that an AI gives you. Treat it like a brilliant intern: fast and helpful, but it needs review.
Modern tools reduce hallucinations by letting the model search the web, read your documents (a technique called retrieval-augmented generation, or RAG) or run code to check its work.
7. What “parameters” and “context window” mean
You’ll often see models described with two numbers:
- Parameters — the internal settings the model adjusts during training. More parameters generally means more capacity to learn patterns, but also more cost to run.
- Context window — how many tokens the model can consider at once. A bigger window means you can paste longer documents or have longer conversations before it loses track.
The takeaway
An LLM is a next-token predictor trained on an enormous amount of text, using attention to understand how words relate. It doesn’t “know” things the way you do, but it has absorbed so many patterns that it can reason, write and code remarkably well.
Understanding this helps you use these tools better: give clear context, ask for step-by-step reasoning, and verify anything important. Do that, and an LLM becomes one of the most powerful tools you’ve ever had.
Finally an explanation that doesn’t assume I have a maths degree. The autocomplete analogy really clicked for me.
Great section on hallucinations. I’m sharing this with my whole team.
Could you do a follow-up on how fine-tuning works?