Skip to main content
Blog

Claude for Developers: Building Reliable AI Features with Anthropic's Models

Last updated AI

Getting a language model to produce something impressive takes an afternoon. Getting it to produce something dependable — the same quality on the ten-thousandth call as on the first, with costs you can predict and failures you can explain — is an engineering problem, and it is mostly not about the model. It is about what surrounds the model: how context is assembled, how output is validated, how tools are exposed, and how you know whether a change made things better.

This guide covers building production features with Anthropic's Claude models: how the interface works, what actually improves output quality, how tool use and structured output should be designed, and the evaluation and operational discipline that separates a feature from a demonstration.


What you will learn
  • The request model, and the parameters that genuinely matter
  • Prompting techniques that survive contact with real inputs
  • Grounding in your own data, and why it beats every other quality lever
  • Tool use and agent loops, designed safely
  • Getting reliably structured output
  • Evaluation, cost control and the operational work around the model
In this article
  1. What you are actually building
  2. The request model
  3. Choosing a model
  4. The system prompt
  5. Prompting that holds up
  6. Grounding in your own data
  7. Structured output
  8. Tool use
  9. Agent loops
  10. Extended thinking
  11. Caching and cost control
  12. Streaming and perceived latency
  13. Evaluation
  14. Guardrails and safety
  15. Operating it
  16. Twelve mistakes
  17. A worked example: a support answering feature
  18. Frequently asked questions

1. What you are actually building

The model is a component, and usually not the interesting one. A production feature is: a way of assembling the right context, a call to the model, validation of what comes back, a fallback when it is wrong, and a record of what happened so you can explain it later.

The proportions surprise people. In a typical feature, the prompt is a small fraction of the work. Retrieval, validation, error handling, evaluation and observability are the rest — and they are where reliability comes from.

The framing that helps throughout: the model generates, your code decides. Anything that must be true should be enforced by deterministic code outside the model, not requested in a prompt. Instructions are requests; validation is a control.

2. The request model

A request consists of a system prompt establishing role and rules, an alternating sequence of user and assistant messages forming the conversation, and parameters controlling generation.

Two properties shape everything. The interface is stateless — the model has no memory between calls, so the entire conversation is resent each time, which is why conversation length drives cost. And the model produces tokens, roughly three-quarters of a word each, billed in both directions, with input typically far cheaper than output.

The parameters worth understanding:

ParameterEffectGuidance
max tokensCeiling on the response lengthSet deliberately; a truncated response mid-sentence is a common avoidable bug
temperatureRandomness in selectionLow for extraction and classification, higher for creative variation
stop sequencesHalt generation at a markerUseful for structured formats
streamDeliver tokens as producedEssential for anything a user waits on

The response includes a stop reason, and checking it is not optional: a response that ended because it hit the token limit is incomplete, and treating it as a complete answer produces silent truncation that is difficult to notice in aggregate.

3. Choosing a model

The family spans capability and cost, and the right choice varies by task rather than by preference. The correct approach is empirical: run your actual evaluation set against several options and compare quality against cost and latency.

The general shape of the decision:

  • High-volume, well-defined tasks — classification, extraction, routing, simple transformation — usually run well on a smaller, faster model at a fraction of the cost. Trying the largest model first and never revisiting is the most common source of unnecessary expense.
  • Complex reasoning, long documents, multi-step agent work and nuanced instruction-following justify a more capable model, where the quality difference is real and the failure cost is high.
  • Mixed workloads benefit from routing: a cheap model handles the majority, with escalation to a stronger one when confidence is low or the task is classified as hard.

Build the routing layer early even if it initially routes everything to one model. It makes model choice a configuration decision rather than a migration, which matters because the landscape changes faster than your codebase should.

4. The system prompt

The system prompt establishes who the assistant is, what it may do, and how it should behave. It is the right place for anything true of every request, and the wrong place for anything specific to one.

What belongs there: the role and domain, the boundaries of what it should attempt, the output format, the tone, and how to handle uncertainty. What does not: the user's actual question, retrieved documents, or anything varying per call — those belong in the user message, clearly separated.

The instruction with the largest effect on reliability is telling the model what to do when it does not know. Without it, the model will produce a plausible answer, because that is what it was trained to do. Instructing it to say explicitly that the provided material does not contain the answer converts a confident fabrication into a useful signal your code can act on.

Keep the system prompt versioned in your codebase and pinned per deployment. A prompt edited directly in a console is a production change with no review, no history and no rollback.

5. Prompting that holds up

Techniques that survive real inputs rather than working on the examples you tested:

Be specific about the output. "Summarise this" produces variable length and focus. "Summarise in three bullet points, each under twenty words, covering the decision, the reason and the next step" produces something consistent enough to render in an interface.

Give examples. Two or three examples of input and desired output improve consistency more than any amount of description, particularly for format and edge-case handling. This is the highest-return prompting technique available.

Delimit untrusted content clearly. Retrieved documents and user input should be wrapped in explicit markers so the boundary between instructions and data is unambiguous. This helps quality and is also the first defence against injected instructions.

Ask for reasoning before the answer on tasks requiring analysis, and put the final answer in a delimited section your code extracts. Reasoning after the conclusion is rationalisation and does not improve accuracy.

Prefill the start of the response where the format matters, which constrains the model into the shape you want from the first token.

State the failure behaviour. What to output when the input is malformed, out of scope or insufficient. Unhandled cases produce creative responses at exactly the wrong moment.

6. Grounding in your own data

The single largest quality lever, and it is not a prompting technique. A model answering from its training data is recalling; a model answering from documents you supplied is summarising, which it does far more reliably.

The pipeline that works has four stages, and the quality of each bounds everything downstream:

  1. Chunking. Split documents into passages that are self-contained. Splitting mid-sentence or mid-table destroys meaning; chunks too large dilute relevance. Preserving structure — keeping a section header with its content — matters more than the exact size.
  2. Retrieval. Hybrid search combining semantic similarity with keyword matching. Pure semantic search misses exact identifiers, part numbers and names; pure keyword search misses paraphrase.
  3. Reranking. Score the candidates against the actual question with a model designed for it. This step consistently produces the largest accuracy improvement per unit of effort and is the one most often omitted.
  4. Context packing. Fit the winners into the budget, remove duplicates, and attach identifiers so the answer can cite them.

Two rules make grounded answers trustworthy. Require citations so every claim maps to a supplied passage, which makes fabrication detectable. And filter by permission at query time, so a user cannot receive content from a document they could not otherwise open — this is an access control decision that must happen in retrieval, not in the prompt.

7. Structured output

Most features need machine-readable output rather than prose. Three techniques, in increasing reliability.

Describe the schema and ask. Works reasonably and fails occasionally, usually by adding explanatory text around the structure.

Prefill and stop. Begin the response with the opening of the structure and stop at its close, which constrains the shape considerably.

Use tool definitions as schemas. Defining a tool whose input schema is your desired output, and requiring the model to call it, produces structured output validated against a schema by the interface itself. This is the most reliable approach and is what production features should use.

Whichever you choose, validate on receipt. Parse against your schema, and on failure either repair, retry with the error included, or fall back — never pass unvalidated output downstream. Structured output that is usually correct becomes a correctness problem the moment it is treated as always correct.

8. Tool use

Tool use lets the model request an action: search a database, look up an order, send an email. You define the tools; the model asks; your code executes and returns the result.

The security property that matters is that the model never touches your systems directly. It emits a request; your code decides whether and how to comply. Everything a model can do is something you explicitly enabled.

Design guidance:

  • Descriptions matter more than names. The description is how the model decides when to use a tool, and vague descriptions produce tools called at the wrong moments.
  • Narrow parameters. Enumerated values rather than free text wherever possible. A tool accepting an arbitrary query string is a tool that will eventually be asked something you did not anticipate.
  • Few tools, clearly distinguished. Twenty overlapping tools produce poor selection. Consolidate.
  • Authorise in your code. The permission check happens when you execute, using the actual caller's identity — never assume the model only requests permitted things.
  • Return errors as results. A tool that fails should return a description of the failure so the model can adapt, rather than throwing and breaking the loop.
  • Keep results small. Tool output enters the context and is resent on every subsequent turn.

9. Agent loops

An agent loop runs until the task is complete: the model requests a tool, your code executes it, the result is returned, and the model continues. The pattern is simple and the safeguards are not optional.

Bound everything. A maximum number of iterations, a total token budget, and a wall-clock timeout. Without these, an unproductive loop consumes real money.

Make actions reversible or confirmed. Anything with external consequence — sending a message, charging a card, deleting data — should either be reversible or require human confirmation. This is a design decision, not a configuration.

Make tools idempotent. Retries happen, and a tool that creates something twice produces a duplicate.

Treat tool results as untrusted. Content retrieved from a document or a web page may contain text attempting to redirect the model's behaviour. The defence is architectural — the model has only the tools you gave it, and permissions are enforced in your code — rather than a prompt asking it to ignore instructions.

Trace every step. Which tools were called with which arguments, and what was returned. Without this, debugging an agent that did something unexpected is guesswork.

10. Extended thinking

Models supporting extended reasoning can work through a problem before answering, with the thinking visible and separately budgeted. It improves accuracy on genuinely hard tasks — multi-step analysis, complex constraints, careful comparison — at the cost of additional tokens and latency.

Where it helps: problems where a careless first answer is likely to be wrong, and where the failure cost justifies the extra cost. Where it does not: classification, extraction, formatting and other tasks that are essentially pattern application, where it adds cost without improving results.

The practical approach is the same as for model selection: measure on your evaluation set. Enable it where the accuracy improvement justifies the cost, and treat the thinking output as diagnostic material for you rather than as something to show the user.

11. Caching and cost control

Prompt caching is the largest available cost reduction for most features, and it is frequently unused.

The mechanism: a stable prefix — system prompt, tool definitions, reference documents — is cached, and subsequent requests reusing it pay substantially less for those tokens. The requirement is that the cached portion is genuinely identical and appears at the start, which means ordering the prompt with stable content first and variable content last. That single structural decision determines whether caching works at all.

Other levers, in order of typical impact:

  • Route by task difficulty. Most requests do not need the most capable model.
  • Semantic caching for repeated questions, keyed by tenant and entitlement so one customer never receives another's answer.
  • Trim conversation history. Summarise older turns rather than resending everything.
  • Bound output length to what the interface actually displays.
  • Batch processing for work that tolerates delay, which is typically much cheaper.

Instrument cost per request from the first day, attributed to the feature and the tenant. Cost that is not visible per feature cannot be optimised, and it arrives as a monthly surprise.

12. Streaming and perceived latency

Generation takes time proportional to output length. For anything a user waits on, streaming transforms the experience: time to first token is what feels like speed, and it is a fraction of total generation time.

Practical considerations: handle partial structured output carefully, since a partially generated structure is not parseable — either stream prose and buffer structure, or design the format so partial results are meaningful. Handle disconnection, since a user closing the tab mid-generation should stop the work rather than continuing to bill for it. And retain the assembled response server-side, because the client's copy is not authoritative and may be incomplete.

For background work, do not stream at all — return immediately with a reference, process asynchronously, and notify. Holding an HTTP connection open for two minutes fails in a dozen ways across proxies and mobile networks.

13. Evaluation

Without evaluation, changing a prompt is gambling. With it, prompt and model changes become ordinary engineering.

The components:

A golden set of representative inputs with known-good outputs, built from real traffic rather than invented examples. Fifty well-chosen cases covering the common paths and the known failure shapes is enough to start, and it beats five hundred invented ones.

Automatic checks where possible: did it return valid structured output, did it cite real documents, did it call the expected tool, does the extracted value match. These are cheap and catch a large share of regressions.

Model-graded checks for qualities that resist exact comparison — is the answer helpful, is the tone right, is it grounded in the sources. Use a rubric, and validate the grader against human judgement on a sample, because an unvalidated grader is a confident opinion.

Run on every change to prompt, model, retrieval configuration or tool definitions, and block promotion on regression. Run on a schedule too, because provider-side model updates can change behaviour without any change on your side.

Grow the set from incidents. Every failure that reaches a user becomes a case, which is how the harness stays relevant rather than testing yesterday's problems.

14. Guardrails and safety

Deterministic checks outside the model, with authority to reject.

On input: strip secrets and personal data that need not be sent, screen for injection patterns, and enforce length and rate limits.

On output: validate structure against the schema, verify citations resolve to real documents, check that claims are grounded where grounding is required, and screen for content categories your policy prohibits.

The rule worth stating plainly: a guardrail that only logs is a metric, not a control. Give it the ability to fail a request, and track how often it fires as a health signal — a sudden change in guardrail activation usually indicates an upstream change worth investigating.

On prompt injection specifically, the durable defence is architectural rather than textual. Treat all retrieved and user-supplied content as data rather than instructions, give the model only tools whose blast radius you accept, and enforce every permission in code that runs outside the model. Prompt wording is a filter, not a wall.

15. Operating it

What to record for every request: prompt version, model, input and output token counts, cost, latency, retrieved document identifiers, tools called, guardrail outcomes, and a request identifier tying it together.

What to alert on: error and refusal rate, latency at the ninety-fifth percentile, cost per request drifting, guardrail activation changing sharply, and grounding rate falling.

What to build before launch: a fallback path for provider unavailability — a cheaper model, a cached answer, or an honest decline — and a way to disable the feature without a deployment. Both are cheap in advance and painful to add during an incident.

Retention deserves a deliberate decision. Prompts and responses may contain personal data, and keeping them indefinitely for debugging is a liability. Decide the period, apply it automatically, and make sure the audit trail you need for explaining a past decision survives the deletion of the material you do not need.

16. Twelve mistakes

  1. Not checking the stop reason. Truncated responses treated as complete.
  2. Unvalidated structured output. Usually correct becomes always assumed.
  3. Prompts edited in a console. Production changes with no review or rollback.
  4. Variable content before stable content. Caching never engages.
  5. The largest model for every task. Substantial unnecessary cost.
  6. No retrieval reranking. The largest available accuracy improvement, skipped.
  7. Permissions checked in the prompt. Access control that a sentence can bypass.
  8. Unbounded agent loops. Real money, spent on nothing.
  9. Tool results treated as trusted. Retrieved content can contain instructions.
  10. Guardrails that only log. Documentation, not control.
  11. No evaluation set. Every prompt change is a gamble.
  12. No fallback path. A provider incident becomes your outage.

17. A worked example: a support answering feature

Consider a feature that drafts answers to customer support questions from the company's own documentation, for an agent to review before sending.

The human-in-the-loop decision comes first, and it shapes everything. Drafting for review rather than answering directly means a wrong answer costs an agent thirty seconds rather than costing a customer relationship. That choice makes the entire feature viable months earlier than a fully automated version would be.

Retrieval is where the quality lives. Documentation is chunked by section with headers preserved, indexed with both semantic and keyword search, and reranked against the actual question. The reranking step is added second rather than first, and its introduction produces a larger accuracy improvement than any prompt change made before or after it — which is the usual pattern and the reason it is worth doing early.

The prompt is ordered for caching. System prompt, tool definitions and the small set of always-relevant policy documents come first and are cached; the retrieved passages and the customer's question come last. This ordering is decided before writing the prompt rather than discovered afterwards, because restructuring later means rewriting it.

Output is structured through a tool definition: the draft answer, an array of citations with document identifiers, a boolean indicating whether the question was answerable from the supplied material, and a suggested category. The boolean is the important field — the system prompt instructs the model to set it false and decline rather than guess, and the interface renders that as "no confident answer, here are the two most relevant documents", which agents find more useful than a plausible paragraph.

Validation runs on every response. The structure is parsed against the schema. Every cited identifier is checked against the documents actually retrieved for that request — a citation to a document not in the context is a fabrication, and catching it is a two-line check that eliminates the most damaging failure mode.

The evaluation set is built from real tickets. Fifty questions with known-good answers, covering the common topics, the known ambiguous cases, and three questions the documentation genuinely does not answer — those last three test the decline path, which is the behaviour most likely to regress silently when a prompt is edited.

Cost lands lower than expected. A smaller model handles the majority of questions adequately, with routing to a stronger model when retrieval confidence is low. Prompt caching covers the stable prefix. The two together reduce cost per answer by roughly an order of magnitude compared with the first working version, with no measurable quality difference on the evaluation set.

What is recorded per request: prompt version, model, retrieved document identifiers, tokens, cost, whether the agent edited the draft before sending, and whether the customer replied again. That last signal is the one that actually measures usefulness, and it becomes the source of new evaluation cases.

18. Frequently asked questions

Should we fine-tune or prompt?

Prompt first, then ground in your data, then consider fine-tuning — in that order, and only when the previous step has stopped paying. Fine-tuning excels at format adherence, tone and domain vocabulary. It is a poor substitute for knowledge, because facts change and a fine-tuned model cannot cite. Most features that considered fine-tuning actually needed better retrieval.

How do we stop it making things up?

Ground it: supply the relevant documents and require the answer to cite them, then verify that every citation resolves to a document actually in the context. Instruct it explicitly to decline when the material is insufficient, and make declining a useful outcome in your interface. Grounding plus citation verification removes most fabrication; prompt wording alone does not.

How do we handle a provider outage?

Design the ladder before you need it: primary model, a fallback model, a cached answer for repeated questions, then an honest decline. Build a way to disable the feature without deploying. An honest "this is unavailable right now" is a considerably better experience than a request that times out after thirty seconds.

What is a realistic cost profile?

Input tokens usually dominate, particularly with retrieval, which is why caching matters so much. The surprises tend to be conversation history resent on every turn, agent loops iterating more than expected, and retrieved context that is larger than necessary. Measure cost per resolved request rather than per call — a cheap answer that triggers a human escalation is the most expensive outcome available.

Can we send customer data?

Subject to your contract, your data classification and your obligations — which are procurement questions with documented answers rather than unknowable risks. Send the minimum necessary, redact identifiers you do not need, log what was sent, and confirm retention and training terms in writing. Many organisations run these features under an agreement that addresses exactly this.

How large should the evaluation set be?

Fifty well-chosen cases from real traffic is a useful start and far better than five hundred invented ones. Cover the common paths, the known failure shapes and the cases where the correct behaviour is to decline. Grow it from incidents — every failure that reaches a user becomes a case, which is what keeps it testing today's problems rather than last year's.

Should we build agents or keep it simple?

Keep it simple until a specific requirement forces otherwise. A single call with good retrieval and validation is easier to evaluate, cheaper, faster and more predictable. Agent loops earn their complexity when the task genuinely requires multiple steps whose sequence depends on intermediate results — and even then, bound them tightly and keep a human at the consequential decisions.

What should we build first?

One narrow capability, with a human reviewing output before it matters, structured output validated on receipt, logging of prompt version and cost, and an evaluation set built from the traffic it produces. That is a week of work and it establishes everything the harder features will need — which is a considerably better position than an impressive demonstration nobody can safely deploy.

Key takeaways

  • The model is a component. Retrieval, validation and evaluation are where reliability comes from.
  • Grounding beats prompting. Supplied documents plus verified citations is the largest quality lever available.
  • Use tool schemas for structured output, and validate everything on receipt.
  • Order the prompt for caching. Stable content first is a structural decision, not an optimisation.
  • Guardrails must be able to reject. Prompt instructions are requests; code is control.
  • No evaluation set means every change is a gamble. Build it from real traffic and grow it from incidents.

A reliable AI feature looks unremarkable from the outside: it answers when it can, says so when it cannot, cites what it used, costs a predictable amount, and can be explained three weeks later. Almost all of that comes from the engineering around the model rather than from the model itself.

Enjoyed this article?

Get more engineering insights from ELIVTECH — or talk to us about your project.

Get in touch