Skip to main content
Blog

Under the Hood of Claude: How Anthropic's Model Thinks

Last updated AI

People describe using a capable language model as "talking to something that thinks", and then immediately worry that the phrase is misleading. Both reactions are reasonable. Something genuinely useful is happening between a question and an answer, and it is neither magic nor a lookup table.

This article opens the box on how a modern assistant like Claude actually produces a response — the training that shapes it, the reasoning it does before answering, how it is steered toward being helpful and honest, and where its limits sit. No mathematics beyond arithmetic, and no code.


What you will learn
  • What the underlying model is and what it was trained to do
  • How pre-training, fine-tuning and alignment differ
  • What "extended thinking" actually is and when it helps
  • How context, memory and tools work together
  • Why models hallucinate, and what reduces it
  • Practical implications for how you should prompt
In this article
  1. The one-sentence version
  2. Tokens: the unit of everything
  3. Pre-training: absorbing the pattern of language
  4. Why prediction produces competence
  5. Fine-tuning: from text engine to assistant
  6. Alignment and the constitutional approach
  7. The context window
  8. Extended thinking
  9. Effort, and the cost of thinking
  10. Tool use and agency
  11. Memory, and why it is not what you assume
  12. Hallucination: the honest account
  13. What the model does not know about itself
  14. Practical implications for prompting
  15. Twelve misconceptions
  16. A worked example: watching a hard question get answered
  17. Glossary
  18. Frequently asked questions

1. The one-sentence version

A large language model is a very large statistical function that, given a sequence of text, produces a probability distribution over what comes next — and everything else is a consequence of how that function was trained and how it is wrapped.

That sentence sounds deflating, and it is worth resisting the deflation. "Predicting the next word" is a training objective, not a description of the resulting capability. Learning to predict text well enough, across a large enough and diverse enough body of writing, requires internalising an enormous amount about grammar, facts, reasoning patterns, code, and the structure of arguments — because all of those things influence what word comes next. The objective is simple; what it forces the system to learn is not.

2. Tokens: the unit of everything

Models do not see letters or words. They see tokens — chunks of text, typically three to four characters in English. Common words are single tokens; rare words split into several; whitespace and punctuation are tokens too.

This is not a trivial implementation detail. It explains several persistent oddities:

  • Counting letters is genuinely hard for a model, because it never sees letters. Asking how many times a letter appears in a word is asking it to reason about something below its perceptual resolution.
  • Rare names and identifiers fragment into many tokens, which makes them harder to handle consistently than common words.
  • Cost and limits are measured in tokens, not words. A rough conversion: about 750 words per thousand tokens in English; considerably fewer for code or non-Latin scripts.

Every input is converted to tokens, and the output is produced one token at a time — each one chosen based on everything that came before it, including the tokens the model itself just produced.

3. Pre-training: absorbing the pattern of language

The first and by far the most expensive stage. The model is shown vast quantities of text and repeatedly asked to predict the next token; where it is wrong, its internal parameters are nudged. Repeat trillions of times.

Nobody labels this data. Nobody tells the model what is true. The training signal is entirely "what actually came next in this text". What emerges is a system that has, in a compressed and lossy way, absorbed the statistical structure of an enormous amount of written material.

Two consequences follow directly and explain a great deal of behaviour:

Knowledge is baked in at a point in time. The model's parameters reflect the text it saw during training, which ended on some date. It does not know what happened afterwards unless you tell it, or unless it can look it up. This is the training cutoff, and it is why current information should always come from a search or a tool rather than from the model's memory.

Knowledge is compressed, not stored. The model is far smaller than its training data. It cannot have memorised it; it has learned generalisations. Well-represented facts — repeated across many sources — are reliably retained. Obscure ones are approximated, and approximation is where confident errors come from.

4. Why prediction produces competence

The interesting question is why an objective this simple yields something that can debug code or structure an argument.

The intuition: to predict the next token well across genuinely diverse text, you have to model whatever generated that text. Predicting the last word of a proof requires something proof-like. Predicting the output of a described function requires something computation-like. Predicting how a character responds requires something like a model of intentions. None of this is programmed; all of it is instrumentally useful for the prediction task, so training finds it.

Scale matters here in a way that surprised the field. Beyond certain sizes, capabilities appear that were absent at smaller scales — following instructions in a format never explicitly trained on, doing arithmetic on unseen numbers, translating between languages that never appeared paired. The capability was not designed; it fell out of scale and data.

This is also why capabilities are uneven. There is no guarantee that a system which handles a hard problem well handles an easy-looking one — competence follows the structure of what appears in text, not human intuitions about difficulty.

5. Fine-tuning: from text engine to assistant

A purely pre-trained model is not an assistant. Given a question, it might continue with more questions — because that is a plausible continuation of a document containing a question.

Turning it into something that answers takes further training on demonstrations of the desired behaviour: conversations where a helpful, accurate, appropriately careful response follows a request. This is far smaller in scale than pre-training but does the shaping. It is where the assistant's characteristic register comes from — direct answers, acknowledged uncertainty, a consistent format.

An important framing: fine-tuning does not add much knowledge. It shapes behaviour. Almost everything the model knows came from pre-training; fine-tuning determines how it uses it.

6. Alignment and the constitutional approach

Beyond being helpful, a deployed assistant needs to be honest and to decline genuinely harmful requests without becoming useless or preachy about ordinary ones. This is the alignment stage.

The conventional approach uses human feedback: people compare responses, preferences train a reward model, the model is optimised against it. It works, and it has limits — it is expensive, slow, inconsistent between raters, and it is difficult for a human comparing two answers to explain why one is better in a way that generalises.

Anthropic's constitutional AI approach adds an explicit written set of principles — a constitution. The model critiques and revises its own outputs against those principles, and the revisions train it. The principles are visible and debatable rather than implicit in thousands of individual judgements.

The practical advantages are worth naming. Behaviour becomes auditable: you can read the principles rather than inferring them from behaviour. It becomes adjustable: changing a principle is a legible edit. And it scales better, because generating critiques does not require a human for every example.

It is not a solved problem. Principles conflict; being maximally helpful and maximally cautious pull in opposite directions on genuinely ambiguous requests, and where the line falls is a judgement call that reasonable people disagree about. The visible symptoms — occasionally declining something harmless, occasionally hedging where a direct answer would serve better — are the cost of a system tuned to avoid worse failures in the other direction.

7. The context window

The context window is everything the model can see at once: your instructions, the conversation so far, any documents included, tool results, and its own output as it generates.

Current Claude models support very large windows — up to a million tokens on the largest, which is roughly a substantial codebase or several hundred pages of documents. That size changes what is practical: you can include an entire specification rather than summarising it.

Three things are commonly misunderstood about it.

It is not memory. When a conversation ends, the context is gone. A new conversation starts blank. Anything that persists does so because an application stored it and re-supplied it.

Everything in it costs, on every turn. Each message reprocesses the whole context. A conversation with a large document attached pays for that document repeatedly. This is what prompt caching addresses — a stable prefix can be cached so it is not reprocessed at full cost each time, which for a long system prompt or a large attached document is a substantial saving. Caching has a minimum useful size, so it pays off for large stable prefixes rather than short ones.

Position matters. Material at the beginning and end of a long context is attended to more reliably than material buried in the middle. If something is critical, place it deliberately rather than trusting it will be found.

8. Extended thinking

Left to answer immediately, a model commits to its first token before working anything out — and everything after that is a continuation of a possibly wrong start.

Extended thinking changes this. The model produces reasoning before its answer: considering approaches, checking steps, noticing errors. Only then does it respond. The improvement on multi-step problems is large and reliable, and the reason is mechanical rather than mysterious. Generating reasoning tokens is literally more computation applied to the problem, and each reasoning step is visible to the model when producing the next one. It is the difference between answering instantly and being allowed a sheet of paper.

Where it helps most: mathematics and logic, multi-step planning, debugging, analysis with several interacting constraints, and anything where a wrong early commitment is hard to recover from.

Where it helps least: recall, simple rewriting, formatting, classification, and short factual answers. Thinking about "what is the capital of France" adds latency and nothing else.

On the current generation, thinking is adaptive — the model decides how much reasoning a given problem warrants, rather than being handed a fixed budget to spend. A trivial request gets little; a hard one gets more. That removes a tuning burden that used to fall on the developer.

One point that surprises people: on the most capable current models, the raw reasoning is not returned. You can request a summary of it, and that summary is genuinely useful for understanding how a conclusion was reached, but the unfiltered chain is not exposed. Raw reasoning is exploratory — it contains discarded branches and half-formed ideas — and presenting it as if it were considered output is misleading.

9. Effort, and the cost of thinking

Thinking costs tokens and time, which is why current models expose an effort control — a dial from low through high to the maximum, governing how much work the model puts into a response.

The practical framing is a trade: low effort for high-volume, low-stakes work where speed and cost dominate; high effort for genuinely hard problems where being right matters more than being fast. The default sits in the middle for a reason, and most workloads never need to touch it.

What is worth internalising is that this is a real trade-off rather than a quality setting to be maxed out. Running everything at maximum effort makes a classification pipeline expensive and slow without making it more accurate — the problem was never hard enough to benefit.

10. Tool use and agency

A model alone can only produce text. Tool use is what lets it affect anything.

The mechanism is straightforward. The application describes the available tools. The model, instead of answering, produces a structured request to use one. The application executes it — the model does not — and returns the result, which enters the context. The model continues with that information available.

The security consequence is worth stating plainly, because it is frequently misunderstood: the model cannot do anything your application does not do on its behalf. It cannot reach the network, read a file or call an API. It can only ask, and your code decides. Every permission decision belongs there, not in instructions to the model.

Repeating this loop — request a tool, see the result, decide the next step — is what an agent is. There is no separate agent technology; it is this loop plus a goal and a stopping condition. Agents are useful precisely where the sequence of steps cannot be known in advance, because each step is chosen after seeing the previous result. Where the sequence is known, a script is cheaper and more predictable.

11. Memory, and why it is not what you assume

The default state is complete amnesia between conversations. Anything that looks like memory is an application feature.

Common implementations: replaying prior messages into the context; summarising older exchanges to fit; storing facts in a database and retrieving relevant ones; or maintaining explicit notes the model reads and writes.

Each has costs. Full replay is faithful and expensive. Summarisation is cheap and lossy — details are dropped, and which details is not always predictable. Retrieval is scalable and only surfaces what a query matches. Explicit notes are auditable and require deciding what is worth keeping.

Knowing which mechanism you are using explains otherwise puzzling behaviour. An assistant that "forgot" something from earlier in a long conversation has almost certainly had that portion summarised away, not chosen to ignore it.

12. Hallucination: the honest account

Models sometimes state false things fluently and confidently. The mechanism follows directly from everything above.

The model produces plausible continuations. For well-represented facts, plausible and true coincide, because the true version is what appeared in the training text. For obscure ones, the model has a compressed approximation, and the most plausible continuation may be a smooth invention: a citation with the right shape, an API function that ought to exist, a date in the right decade.

Crucially, the model is not lying — that would require knowing the truth and choosing otherwise. It is doing exactly what it does, in a region where doing that well is not the same as being right.

What actually reduces it, in order of effectiveness:

  • Supply the source material. A model answering from a document you provided is grounded in that document. This is the single largest reduction available and the reason retrieval-augmented approaches dominate factual applications.
  • Let it use tools. Looking something up beats recalling it, always.
  • Enable thinking on hard questions. Reasoning surfaces the model's own uncertainty more often than immediate answering does.
  • Invite uncertainty explicitly. Asking it to say when it is unsure genuinely changes behaviour, because the training rewards honesty over confident filling.
  • Verify anything checkable — names, numbers, citations, identifiers. These are exactly the categories where approximation is invisible and consequential.

What does not work: asking the model to be accurate, or asking whether it is sure. Confidence expressed after the fact is generated the same way as the original answer.

13. What the model does not know about itself

A genuinely counterintuitive point. The model has no privileged access to its own workings. It cannot report its parameter count, describe its architecture from the inside, or explain which internal computation produced an answer. When asked, it produces a plausible response based on what has been written about such systems — which may be accurate, and is not introspection.

The same applies to explanations of its reasoning. A post-hoc explanation is a generated text that plausibly accounts for the answer. It often corresponds to what actually happened, and it is not a readout. Extended thinking is closer to genuine, because the reasoning existed before the answer and influenced it — but even a thinking summary is a rendering, not a trace.

Practically: do not treat a model's account of itself, its limits or its training as authoritative. Check the documentation.

14. Practical implications for prompting

Everything above cashes out into concrete habits.

Provide context rather than relying on recall. The model is far better at using material you supply than at remembering material it saw during training. Paste the document.

Be specific about the output you want. "Summarise this" has a thousand valid interpretations. "Three bullets, each naming a risk and its mitigation, for a non-technical reader" has one.

Show an example when format matters. One good example does more than a paragraph of description, because pattern-matching to a demonstrated shape is exactly what the system is good at.

Let it think on hard problems, and do not on easy ones. The cost is real in both directions.

Ask for reasoning when you need to check the work. Not because the explanation is a trace, but because a visible argument is checkable in a way a bare conclusion is not.

Split complex tasks. Several focused requests outperform one that asks for everything, and each becomes independently verifiable.

Put the important thing at the start or the end of a long prompt, not buried in the middle.

Iterate on the prompt, not just the output. A prompt that reliably produces what you want is reusable; a manually fixed output is not.

15. Twelve misconceptions

  1. "It looks things up." Not unless given a tool. By default it answers from compressed training.
  2. "It remembers me." Only if the application replays or retrieves prior material.
  3. "It learns from our conversation." Within the context, yes; permanently, no.
  4. "It knows when it is wrong." Poorly calibrated on obscure facts — that is exactly where errors occur.
  5. "It can explain its reasoning." It can produce a plausible account, which is not the same thing.
  6. "Thinking always helps." On easy tasks it adds cost and nothing else.
  7. "Bigger context means better recall." Position within a long context still matters.
  8. "It can access the internet." Only through a tool your application provides.
  9. "Rephrasing the same question confirms an answer." The same failure mode produces the same wrong answer.
  10. "It is deterministic." Sampling introduces variation by default.
  11. "Asking for accuracy improves accuracy." Grounding does; instruction alone does not.
  12. "It understands like a person." The representations are genuinely rich and genuinely not human — expecting either extreme will mislead you.

16. A worked example: watching a hard question get answered

Take a realistic request: "Our checkout conversion dropped four percent after last Tuesday's release. Here are the release notes and the last two weeks of funnel metrics. What most likely caused it?"

Tokenisation and context assembly. The request, the release notes and the metrics are converted to tokens and assembled with the system prompt. If the system prompt is stable across requests, caching means it is not reprocessed at full cost — relevant when the same analysis runs daily.

Difficulty assessment. This is genuinely multi-step: correlate a timing signal with a set of changes, weigh alternative explanations, and account for confounders. Adaptive thinking allocates real reasoning here, where the same system would allocate almost none to "what does this error code mean".

Reasoning. The model works through the funnel stage by stage, notes that the drop concentrates at payment rather than being spread evenly, scans the release notes for anything touching payment, finds a change to a third-party integration, and considers alternatives — a seasonal effect, a traffic-mix shift, an instrumentation change that only appears to be a drop. This is where a first-token-immediately answer would have committed to "the payment change" before checking whether the drop was even concentrated there.

The knowledge boundary. Nothing in training knows anything about this company. Every fact in the answer comes from the supplied material — which is exactly the configuration in which the system is most reliable. Had the question been "what is the typical checkout conversion for our industry", the answer would come from compressed training, would be approximate, and would deserve verification.

Tools, if available. With a query tool, the model can check whether the drop appears across all payment methods or one, which distinguishes the integration hypothesis from a broader cause. It requests; the application executes; the result returns. Without tools, it can only name the check and ask you to run it.

The response. A ranked set of hypotheses with the evidence for each, the confounders considered, and the specific check that would confirm the leading one. Some of it — the reasoning about where the drop concentrates — is directly verifiable against the data supplied. That verifiability is the point, and it is why "show your reasoning" is worth asking for even though the reasoning is not a mechanical trace.

What would have degraded it. Supplying only the release notes without the metrics, forcing recall instead of grounding. Asking for a single answer rather than ranked hypotheses, encouraging false confidence. Or running at minimum effort, where a genuinely multi-step correlation gets a first-impression answer.

17. Glossary

TermMeaning
TokenThe unit of text a model processes — roughly 3–4 characters in English.
Pre-trainingThe large, expensive stage where next-token prediction is learned from vast text.
Fine-tuningSmaller subsequent training that shapes behaviour rather than adding knowledge.
AlignmentTraining toward being helpful, honest and appropriately cautious.
Constitutional AIAlignment guided by an explicit written set of principles the model critiques itself against.
Context windowEverything the model can see at once, measured in tokens.
Training cutoffThe point after which the model's built-in knowledge stops.
Extended thinkingReasoning generated before the final answer, improving multi-step accuracy.
Adaptive thinkingThe model deciding how much reasoning a problem warrants, rather than a fixed budget.
EffortA control trading response quality against speed and cost.
Prompt cachingReusing a processed stable prefix so it is not paid for in full each turn.
Tool useThe model requesting an action the application executes on its behalf.
AgentA loop of tool use plus a goal and a stopping condition.
GroundingSupplying source material so answers derive from it rather than from recall.
HallucinationA fluent, confident, incorrect output — approximation where recall was expected.
TemperatureA sampling control governing how varied the output is.

18. Frequently asked questions

Does it actually understand, or is it pattern matching?

The dichotomy is doing more work than it can bear. It has learned rich internal representations that support genuine generalisation to unseen problems, which is more than "pattern matching" usually implies. It also has no body, no persistent goals and no experience, which is less than "understanding" usually implies. Judge it by what it does on your task rather than by which word wins.

Why is it confidently wrong sometimes?

Because fluency and accuracy are produced by the same mechanism. A well-formed sentence about an obscure fact costs the model nothing extra to generate, and there is no separate internal check comparing it against reality. Ground it in supplied material and the failure mode largely disappears.

Is thinking mode just a longer answer?

No — it is computation before committing. The reasoning is generated first and is visible to the model as it produces the answer, so an error caught during reasoning never reaches the output. A longer answer without prior reasoning still commits to its first token immediately.

Should I always use maximum effort?

No. It costs more and takes longer, and on tasks that are not hard it changes nothing. Reserve it for genuinely difficult problems and use lower settings for volume work — the trade is real in both directions.

Can it access our internal systems?

Only through tools your application provides and executes. The model produces a request; your code decides whether to honour it. It has no independent access to anything, which is why permission enforcement belongs in your code rather than in instructions to the model.

Why does it sometimes decline something harmless?

Alignment tunes toward caution, and the boundary is imperfect — a request that superficially resembles something genuinely harmful can be caught. Adding context about what you are doing and why usually resolves it, and the occasional over-refusal is the visible cost of a system tuned to avoid worse failures in the other direction.

Does a bigger context window mean I should stop summarising?

Partly. Very large windows make it practical to include entire documents rather than extracts, which improves grounding. But everything included costs on every turn, and material buried mid-context is attended to less reliably than material at the edges — so include what is relevant and place the critical parts deliberately.

What single change most improves the answers I get?

Supplying the source material instead of relying on the model's memory. Almost every accuracy complaint traces back to asking the model to recall something it only approximately knows, when the authoritative version could have been pasted in.

Key takeaways

  • Next-token prediction is the objective, not the capability. Learning it well requires internalising a great deal.
  • Knowledge is compressed and time-bounded. Ground it in supplied material for anything factual or current.
  • Thinking is real computation applied before committing — valuable on hard problems, wasted on easy ones.
  • Context is not memory. Persistence is always an application feature.
  • The model cannot act on its own. Every tool call is executed by your code, which is where permissions belong.
  • It has no privileged self-knowledge. Its account of itself is generated, not introspected.

None of this makes the system less useful. It makes it predictable — and predictability is what lets you build on something. Once you know that grounding beats recall, that thinking is a trade, and that the model can only ask rather than act, most surprises stop being surprises.

Enjoyed this article?

Get more engineering insights from ELIVTECH — or talk to us about your project.

Get in touch