Almost everyone now uses a language model. Very few people have a working picture of what happens between typing a question and reading an answer — and without that picture, the technology alternates between seeming miraculous and seeming broken, with no way to predict which you are about to get.
This article builds that picture from the ground up. No mathematics beyond arithmetic, no code, and no hand-waving at the parts that are genuinely interesting. By the end you should be able to predict, reasonably well, what one of these systems will be good at and where it will fail.
What you will learn
- What a language model is trained to do, and why that produces competence
- How text becomes numbers, and why that explains several odd behaviours
- What attention does and why it made everything else possible
- Why scale changed the field
- Why models hallucinate, and what genuinely reduces it
- How to phrase requests so you get better answers
- Start with autocomplete
- Tokens: chopping text into pieces
- Embeddings: meaning as position
- The transformer, and what attention does
- Layers, and what depth buys
- Training: learning from being wrong
- Why prediction produces reasoning
- Sampling: choosing the next token
- The context window
- From raw model to assistant
- Why they hallucinate
- Giving them tools
- What they are genuinely good at
- What they are genuinely bad at
- How to get better answers
- Twelve misconceptions
- A worked example: one sentence, start to finish
- Glossary
- Frequently asked questions
1. Start with autocomplete
Your phone's keyboard suggests the next word. It looks at what you have typed and offers what usually follows. Crude, occasionally useful, easily confused.
A large language model does the same job, and the difference in scale is so extreme that it becomes a difference in kind. Your keyboard considers the last word or two. A large model considers everything in front of it — your entire question, the whole conversation, any documents included — and produces, for every possible next token, a probability.
Then it picks one, appends it, and does the whole thing again. An essay is that loop run a few thousand times.
The obvious objection is that this cannot possibly produce reasoning. Working through that objection is what the rest of this article does — and the short answer is that predicting text well, across enough diverse writing, requires reasoning, because reasoning is what produced the text.
2. Tokens: chopping text into pieces
Models do not work with letters or words. Text is first split into tokens — pieces of roughly three to four characters in English.
Common words are one token each. Rarer words split: "unbelievable" might become "un", "believ", "able". Punctuation and spaces are tokens. Numbers often split in ways that ignore their mathematical structure.
Why this design? A vocabulary of whole words would be enormous and would still fail on anything new. Individual letters would make sequences far too long. Subword pieces are the compromise: a manageable vocabulary that can spell out anything.
This detail explains several persistent oddities:
- Counting letters is hard. The model never sees letters, so asking how many of one appear in a word is asking about something below its perception.
- Arithmetic on long numbers is unreliable. The digits are split into chunks that do not align with place value.
- Unusual names behave oddly. They fragment into many pieces and are handled less consistently than common words.
- Cost is measured in tokens. Roughly 750 English words per thousand tokens; considerably fewer for code or other scripts.
3. Embeddings: meaning as position
Each token is converted to a list of numbers — an embedding, typically several thousand numbers long. That list is a position in a very high-dimensional space.
The useful property is that this space is organised by meaning, because training arranged it that way. Tokens that appear in similar contexts end up near each other. "King" sits near "queen", "monarch" and "throne", and far from "spreadsheet".
More striking, directions in the space carry meaning. The direction from "king" to "queen" is roughly the same as from "man" to "woman" — a gender direction. There are directions corresponding to plurality, to tense, to country-and-capital. Nobody designed these; they emerged because they help predict text.
This is the model's raw material: not words, but positions in a space where similarity of meaning is literally proximity, and relationships between concepts are literally directions.
4. The transformer, and what attention does
The architecture behind every current large model is the transformer, and its central idea is attention.
Take the sentence: The trophy did not fit in the suitcase because it was too large. What does "it" refer to? Trophy. Change "large" to "small" and it refers to the suitcase. Resolving this requires connecting "it" to a word several positions back, and choosing which one based on the adjective at the end.
Attention is the mechanism for exactly that. When processing each token, the model computes how relevant every other token is to it, and builds a representation weighted accordingly. Processing "it", attention weights "trophy" and "suitcase" heavily and "the" barely at all.
Three properties made this transformative:
Distance is irrelevant. A connection between the first and thousandth token is as direct as between adjacent ones. Earlier architectures processed sequentially and lost information across distance.
It is parallel. All positions are processed simultaneously rather than one after another, which is what makes training on enormous datasets feasible at all. This is the practical reason transformers won.
It is learned. Nobody specifies what to attend to. Training discovers it, and different attention heads specialise — some track syntax, some resolve references, some follow topics.
5. Layers, and what depth buys
A single attention step is not enough. Models stack many layers — dozens to over a hundred — each refining the representation from the one below.
The rough progression, from what researchers have been able to inspect: early layers handle surface structure — parts of speech, phrase boundaries. Middle layers handle meaning, resolving references and tracking entities. Later layers handle task-level structure — what kind of answer this is, what tone it needs, what comes next.
By the top, the representation of the final position encodes enough to produce a distribution over next tokens. A final step converts it into probabilities across the whole vocabulary.
Depth is what allows abstraction to build. No single layer decides anything complicated; each makes a modest transformation, and the composition of many is where sophistication lives. It is also why these systems are hard to interpret — there is no layer where "the answer" is decided.
6. Training: learning from being wrong
All of this is governed by parameters — numbers determining how each computation transforms its input. A large model has hundreds of billions of them. Training is the process of setting them.
The procedure is conceptually simple. Take real text. Show the model a prefix. Ask what comes next. Compare its prediction to what actually followed. Adjust every parameter slightly in the direction that would have made the correct answer more likely. Repeat trillions of times across an enormous body of text.
Nobody labels anything. The text itself is the answer key, which is why this scales — the internet is full of training data by construction.
The mechanism for the adjustment works backwards through the network, computing for each parameter how much it contributed to the error and nudging accordingly. Any single nudge is tiny. The accumulation across trillions of examples is what produces the result.
The scale is worth stating concretely: training a frontier model means processing trillions of tokens on thousands of specialised processors for months, at a cost in the tens or hundreds of millions. This is why very few organisations train them and very many use them.
7. Why prediction produces reasoning
Here is the crux. Why does an objective this narrow produce something that can debug code, structure an argument or explain a concept?
The argument runs like this. To predict text well across genuinely diverse material, you must model whatever produced that text. Consider what different predictions require:
- Finishing a sentence about a country's capital requires the fact.
- Finishing a mathematical proof requires following the argument.
- Predicting a program's stated output requires tracing its logic.
- Predicting a character's next line requires a model of their intentions.
- Predicting the conclusion of an argument requires following the reasoning.
None of these capabilities is programmed. Each is instrumentally useful for prediction, so training finds it — because a model with them predicts better than one without, and better prediction is exactly what training selects for.
This also explains why capabilities are uneven in ways that seem arbitrary. Ability tracks what appears in text, not human intuitions about difficulty. Explaining a subtle concept in economics may be easy because thousands of people have written good explanations. Counting characters in a string may be hard because almost nobody writes that out, and the tokenisation works against it. Difficulty for the model and difficulty for a person are different quantities.
8. Sampling: choosing the next token
The model outputs probabilities, not a token. Something has to choose.
Always taking the most likely token sounds right and produces flat, repetitive text that gets stuck in loops. Instead, the choice is sampled from the distribution, and a setting called temperature governs how adventurous that sampling is.
| Temperature | Behaviour | Suits |
|---|---|---|
| Near zero | Nearly always the top choice; repeatable | Extraction, classification, structured output |
| Moderate | Balanced; natural but focused | General use, explanation, analysis |
| High | Adventurous; varied and less reliable | Brainstorming, creative variation |
This is why the same question asked twice gives different answers. It is sampling, not indecision, and it is deliberate: without it the output quality drops noticeably.
It also explains something practically important. Asking the same question again and getting the same wrong answer is not confirmation — the same failure produces the same result. Asking a differently framed question is a real check; repeating one is not.
9. The context window
Everything the model can see at once is its context window: your instructions, the conversation, any attached material, and its own output as it goes.
Modern windows are large — the biggest current models handle around a million tokens, enough for a substantial codebase or hundreds of pages. That changes what is practical, because you can supply whole documents rather than extracts.
Three things to understand:
It is not memory. When the conversation ends, it is gone. Anything that persists was stored and re-supplied by the application.
It costs on every turn. Each message reprocesses everything. This is why long conversations get slower and more expensive, and why caching a stable prefix — a long instruction block, a large attached document — is worth doing when the same material appears repeatedly.
Position matters. Material at the start and end is used more reliably than material in the middle. Put what matters at the edges.
10. From raw model to assistant
A model straight out of training is not an assistant. Ask it a question and it may produce more questions — a perfectly good continuation of a document containing one.
Two further stages turn it into something useful.
Instruction tuning: further training on examples of requests followed by good responses. This teaches the shape of being an assistant — that a question should be answered, that formats should be respected, that a task should be completed.
Alignment: training toward being helpful, honest, and appropriately careful. Some approaches use human preference comparisons; some use an explicit written set of principles the model critiques its own outputs against, which has the advantage of being auditable — you can read the principles rather than inferring them.
Neither stage adds much knowledge. Almost everything the model knows came from the first stage; these two shape how it behaves with it.
11. Why they hallucinate
Models sometimes state false things fluently and with apparent confidence. The mechanism follows from everything above.
The model produces plausible continuations. For well-represented facts, plausible and true coincide — the true version is what appeared in the text. For obscure ones, the model holds a compressed approximation, and the most plausible continuation may be a smooth invention: a citation with the right authors and the wrong title, a function name that ought to exist, a date in the right decade.
The model is not lying. Lying requires knowing the truth and choosing otherwise. It is doing precisely what it does, in a region where doing it well and being right have come apart.
Two structural reasons this is hard to eliminate. The training data itself contains errors, so some falsehoods are learned as facts. And there is no internal check comparing output to reality — a false statement costs nothing extra to produce.
What genuinely reduces it:
- Supply the material. A model answering from a document you provided is grounded in it. This is the largest single improvement available.
- Give it a lookup tool. Retrieval beats recall, always.
- Ask for reasoning. Working through a problem surfaces uncertainty that a direct answer conceals.
- Invite uncertainty. Explicitly asking it to say when unsure changes behaviour.
- Verify the checkable. Names, numbers, citations, identifiers — exactly where approximation is invisible.
12. Giving them tools
Alone, a model can only produce text. Tools are what let it do anything else.
The mechanism: the application describes what tools exist. The model, instead of answering, produces a structured request to use one. The application executes it and returns the result, which enters the context. The model continues with that information.
Two consequences worth being clear about. First, tools fix the two biggest weaknesses at once — currency, because a search returns today's information, and accuracy, because a calculator computes rather than approximates. Second, and importantly for security: the model cannot do anything the application does not do for it. It has no independent access to networks, files or systems. It can only ask.
Running that loop repeatedly — request, observe, decide the next step — is what an "agent" is. There is no separate technology involved.
13. What they are genuinely good at
- Transforming text. Summarising, rewriting, translating, changing register, extracting structure. The strongest category by a distance, because the source material is right there.
- Explaining. Taking a concept and pitching it at a stated level, with analogies.
- Drafting. A reasonable first version of almost any document, faster than starting from blank.
- Code. Writing, explaining and debugging — well-represented in training and highly structured.
- Generating options. Twenty angles on a problem, of which three are good, in seconds.
- Working with supplied material. Answering from documents you provide, which is where accuracy is highest.
14. What they are genuinely bad at
- Current events. Knowledge stops at the training cutoff. Without a search tool, anything recent is guesswork.
- Precise arithmetic. Tokenisation fights it. Use a calculator tool.
- Character-level operations. Counting letters, reversing strings — below its perception.
- Obscure specifics. Exactly where compression turns into invention.
- Knowing what it does not know. Poorly calibrated precisely where errors happen.
- Anything requiring real-world state it has not been told — your systems, your data, your organisation.
- Genuine consistency at length. Contradictions creep into very long outputs.
15. How to get better answers
Supply context rather than relying on recall. The single highest-leverage habit. Paste the document.
Be specific about the output. "Summarise this" has a thousand valid readings. "Three bullets, each a risk with its mitigation, for a non-technical reader" has one.
Show an example when format matters. One demonstration beats a paragraph of description, because matching a shown pattern is exactly what the system does well.
Ask for the reasoning on anything you need to check. A visible argument is verifiable; a bare conclusion is not.
Split complex requests. Several focused ones beat one that asks for everything, and each is independently checkable.
Say who it is for. "Explain to a board member" and "explain to a backend engineer" produce genuinely different and appropriately different answers.
Iterate rather than accept. Say what is wrong with the first answer. The second is usually markedly better, and the context makes it cheap.
Put the important thing first or last in a long prompt.
16. Twelve misconceptions
- "It searches the internet." Not unless given a tool.
- "It remembers previous conversations." Only if the application replays them.
- "It learns from what I tell it." Within the conversation, yes; permanently, no.
- "It knows when it is wrong." Least reliable exactly where errors occur.
- "It has a database of facts." It has compressed generalisations, not records.
- "It thinks before answering." Only if it produces reasoning first — otherwise it commits immediately.
- "Asking again confirms an answer." Same failure, same result.
- "Bigger is always better." Smaller models are faster and cheaper and adequate for many tasks.
- "It understands like a person." Rich representations, no experience — expecting either extreme misleads.
- "Longer prompts are better." Clearer prompts are better; padding dilutes.
- "It can explain its own reasoning." It produces a plausible account, not a readout.
- "Telling it to be accurate makes it accurate." Grounding does; instruction alone does not.
17. A worked example: one sentence, start to finish
Trace what actually happens for: "Explain why the sky is blue to a ten-year-old."
Tokenisation. The sentence becomes roughly a dozen tokens. "Explain", "sky", "blue" are single common tokens; "ten-year-old" splits across several because of the hyphens.
Embedding. Each token becomes a list of several thousand numbers — its position in meaning-space. At this point "blue" is just near other colours; nothing yet knows this is a physics question for a child.
Early layers. Surface structure resolves. "Explain" is identified as an imperative opening a request. "Why the sky is blue" is grouped as its object. "To a ten-year-old" attaches as an audience specification rather than as part of the question.
Middle layers. Meaning assembles. Attention connects "blue" to "sky", activating regions of the representation associated with light, scattering and atmosphere. Separately, "ten-year-old" activates simple vocabulary, concrete analogies, short sentences. These two threads combine — the answer must be about scattering and pitched for a child, and neither constraint alone determines the output.
Later layers. Task shape resolves. This is an explanation, so it should open with a direct answer rather than a preamble. It should probably use an analogy. It should stay under a few hundred words. None of this was requested; it comes from having seen a great many explanations pitched at children.
The first token. The top layer produces a probability over the whole vocabulary. High-probability candidates include "The", "Sunlight", "Light", "Great". Sampling picks one — say "Sunlight". Note that this choice constrains everything after: starting with "Sunlight" commits to leading with the physics rather than with a framing sentence.
The loop. "Sunlight" is appended and the whole process runs again with it included. Now the model is predicting what follows "Explain why the sky is blue to a ten-year-old. Sunlight" — and "looks" and "might" and "is" are the plausible continuations. A few hundred repetitions later, a complete explanation exists.
What emerges. Something like: sunlight looks white but is really all colours mixed; blue light bounces around more than red light when it hits the tiny bits of air; so blue arrives at your eyes from every direction at once. Accurate at that level, correctly pitched, using an analogy nobody supplied.
What did not happen. No lookup of an article about scattering. No plan drafted before writing. No internal check that the physics is right. Each token was chosen because it was a likely continuation, and the result is correct because correct explanations are what the training text contained. Ask the same system about an obscure phenomenon with little written about it and the identical process produces something equally fluent and possibly wrong — the process cannot tell the difference, which is precisely the thing worth remembering.
18. Glossary
| Term | Meaning |
|---|---|
| Token | A chunk of text, roughly 3–4 characters in English; the unit models process. |
| Embedding | A list of numbers representing a token as a position in meaning-space. |
| Transformer | The architecture behind all current large models. |
| Attention | The mechanism letting each position draw on any other, regardless of distance. |
| Layer | One stage of refinement; models stack dozens to hundreds. |
| Parameter | One of the billions of learned numbers governing the computation. |
| Pre-training | The large, expensive stage of learning next-token prediction from vast text. |
| Instruction tuning | Training that turns a text predictor into something that answers requests. |
| Alignment | Training toward helpfulness, honesty and appropriate caution. |
| Temperature | How adventurous sampling is; low is repeatable, high is varied. |
| Context window | Everything the model can see at once. |
| Training cutoff | The date after which built-in knowledge stops. |
| Hallucination | Fluent, confident, incorrect output. |
| Grounding | Supplying source material so answers derive from it rather than recall. |
| Tool use | The model requesting an action the application performs on its behalf. |
19. Frequently asked questions
Is it just a very large autocomplete?
Mechanically, yes. But learning to autocomplete well across all of written language requires internalising facts, grammar, reasoning patterns and argument structure — because those are what determined the text. The objective is simple; what it forces the system to learn is not. Judge it by what it does rather than by how modest the training objective sounds.
Does it understand what it is saying?
It has internal representations rich enough to generalise to problems it never saw, which is more than "no" allows. It has no experience, no body and no persistent goals, which is less than "yes" implies. The honest position is that the question is poorly posed and the practical answer is to test it on your task.
Why does it get simple things wrong?
Because difficulty for the model is not difficulty for a person. Counting letters is hard because it never sees letters. Long arithmetic is hard because digits are chunked awkwardly. Explaining a subtle idea is easy because thousands of good explanations exist in the training text. The mismatch is systematic once you know where it comes from.
Why do I get a different answer each time?
Sampling. The model produces probabilities and one token is drawn from them; different draws give different text. Lowering the temperature setting makes outputs more repeatable, at the cost of flatter and more repetitive writing.
How do I stop it making things up?
Give it the material. A model answering from a document you supplied is far more accurate than one recalling from training, because the answer is present rather than approximated. Add a lookup tool for anything current, ask for reasoning on anything complex, and verify names, numbers and citations — those are exactly where invention is invisible.
Is a bigger model always better?
No. Bigger models are more capable and also slower and more expensive. For classification, extraction, routing and simple transformation, a smaller model is frequently as accurate and dramatically cheaper. Match the model to the task rather than defaulting to the largest.
Can it access my files or the internet?
Only through tools the application provides and executes. The model produces a request; the surrounding code decides whether to act on it. It has no independent access to anything, which is why permission decisions belong in that code rather than in instructions written to the model.
What is the single most useful habit?
Supplying the source material instead of relying on memory. Almost every complaint about accuracy traces back to asking the model to recall something it only approximately knows, when the authoritative version could simply have been pasted in.
Key takeaways
- Next-token prediction is the objective, and doing it well across diverse text is what produced the capability.
- Attention — letting any position draw on any other — is the idea that made this possible.
- Knowledge is compressed and time-bounded. Ground anything factual in supplied material.
- Hallucination is approximation, not deception, and grounding is the fix.
- Context is not memory. Persistence is always something the application does.
- Capability is uneven in ways that track written text, not human intuitions about difficulty.
Understanding the mechanism does not make it less impressive. It makes it predictable — and predictability is what lets you rely on something. Once you know why it invents citations and why it explains concepts well, you stop being surprised by either, and you start using it for what it is actually good at.
Enjoyed this article?
Get more engineering insights from ELIVTECH — or talk to us about your project.
Get in touch