A model that scores well in a notebook and a model that creates value in production are separated by a distance most teams underestimate by about a factor of five. The modelling was never the hard part. The hard part is everything that has to be true for that model to receive correct inputs, produce timely outputs, and keep working when the world changes underneath it.
This guide covers that gap: feature engineering that survives the move to serving, training pipelines that are reproducible, deployment patterns, monitoring that detects the failures that matter, and the organisational shape that makes it sustainable. No code.
What you will learn
- Why notebook-to-production is genuinely hard
- Training-serving skew and how to eliminate it
- Feature stores: what problem they solve and when you need one
- Serving patterns — batch, real-time, streaming
- Monitoring for drift, and what to actually alert on
- Retraining, versioning and rollback
- Why the gap exists
- The shape of a production ML system
- Data: the foundation everything rests on
- Features and the skew problem
- Feature stores
- Training pipelines
- Experiment tracking and model registry
- Evaluation that means something
- Serving patterns
- Latency and cost
- Deployment and rollout
- Monitoring: operational and model
- Drift, and what to do about it
- Retraining
- Feedback loops and their dangers
- Governance and explainability
- Team shape and ownership
- A maturity progression
- Twelve mistakes
- A worked example: a churn model in production
- Frequently asked questions
1. Why the gap exists
Five reasons, and naming them explains most of what follows.
Notebooks hide state. Cells run out of order, variables persist from earlier experiments, and the notebook that produced the good result cannot be reliably re-run. Reproducibility is not a nice-to-have; it is the precondition for everything else.
Training data is not serving data. In training you have a clean historical table. In serving you have a request and must assemble the same features from live systems, quickly, with different code. This mismatch is the single most common cause of a model that scored well and performs badly.
Models decay. Software does what it did last year. A model's accuracy degrades as the world it was trained on changes, silently, without an error.
The failure mode is quiet. A broken service returns errors. A broken model returns confident predictions that are wrong, and nothing in your monitoring notices.
Ownership is ambiguous. The data scientist built it, an engineer deployed it, and when it degrades at three in the morning nobody is on call for it.
2. The shape of a production ML system
| Component | Responsibility | Failure looks like |
|---|---|---|
| Data pipeline | Deliver clean, timely data | Stale or malformed inputs |
| Feature layer | Compute features consistently for training and serving | Training-serving skew |
| Training pipeline | Produce a model reproducibly | A model nobody can rebuild |
| Registry | Track versions and lineage | Nobody knows what is deployed |
| Serving | Return predictions within budget | Latency or availability failures |
| Monitoring | Detect operational and model problems | Silent degradation |
| Retraining | Refresh the model as data changes | Accuracy decay |
Notice the asymmetry: only two of these are about the model. Most of a production ML system is data infrastructure and operations, which is why teams staffed entirely with modellers struggle regardless of how good the modelling is.
3. Data: the foundation everything rests on
Model quality is bounded by data quality, and no architecture compensates for bad inputs.
Validate at ingestion. Schema, ranges, null rates, category values, volume against expectation. A pipeline that silently accepts a column that became all nulls will train a model on it and serve predictions from it.
Detect distribution changes at the data layer, not just at the model layer. A column whose mean shifted by forty percent is worth an alert before anyone looks at prediction quality.
Track lineage. Which raw data produced which features which produced which model. Without it, investigating a bad prediction is archaeology.
Handle late-arriving data explicitly. Events that arrive after the window they belong to are normal. A pipeline that computes an aggregate and never revisits it will disagree with the raw data permanently.
Version your datasets. "Trained on the customer table" is not reproducible. "Trained on this snapshot" is.
4. Features and the skew problem
The central technical difficulty, and worth being precise about.
Training-serving skew is any difference between how a feature is computed during training and during serving. It arises almost inevitably when the two are implemented separately — one in an analytics environment over historical tables, the other in application code against live systems.
The differences that cause it are mundane and consequential: a slightly different definition of a time window; different handling of missing values; a category encoding derived from a different set of values; a rounding difference; an aggregate computed over a different period. Each is small and each degrades the model in ways that look like the model being worse than it tested.
Related and equally damaging: leakage, where a feature encodes information that would not be available at prediction time. A model predicting whether a customer will churn, trained with a feature derived from their cancellation record, will score beautifully and be useless.
The defences, in order of effectiveness:
- Compute features once, in one place, used by both training and serving. This is what a feature store exists to do.
- Use point-in-time correct joins when building training data, so each row's features reflect only what was known at that moment.
- Compare distributions of the same feature in training and in production. A mismatch is a bug, not noise.
- Log features at serving time and reuse those logs as training data. This makes skew structurally impossible, and it is the most reliable approach available.
5. Feature stores
A feature store centralises feature definitions and serves them to both training and serving paths.
It typically provides an offline store — historical values with point-in-time correctness for building training sets — and an online store optimised for low-latency lookup during serving, with the same definition producing both.
The benefits are real: skew is structurally prevented, features are reusable across models, and definitions are documented in one place rather than reimplemented per project.
The honest caveat: a feature store is significant infrastructure, and a team with one or two models does not need one. The problem it solves — many models, many features, many teams reimplementing the same definitions — is a scale problem. Before that scale, a shared library of feature computations used by both paths achieves most of the benefit for a fraction of the effort.
6. Training pipelines
The transition from a notebook to a pipeline is where reproducibility is won or lost.
A production training pipeline should be: parameterised, so a run is defined by its configuration; versioned in data, code and configuration together; reproducible, such that re-running produces the same model within acceptable tolerance; automatable, runnable on a schedule or a trigger; and observable, emitting metrics and artefacts you can inspect afterwards.
The practical steps: pull a versioned dataset, compute or fetch features, split with a strategy appropriate to the problem, train, evaluate against a held-out set and against the current production model, register the resulting model with its metrics and lineage, and stop — deployment is a separate decision.
Two points that cause trouble when skipped. Splitting strategy matters enormously for temporal problems: a random split on time-series data leaks the future into training and produces an evaluation that is meaningless. And evaluate against the current production model, not only against a threshold, because the question is always whether to replace what is running.
7. Experiment tracking and model registry
Experiment tracking records what was tried: parameters, data version, code version, metrics and artefacts. Without it, the answer to "why did we choose this model" is somebody's memory.
A model registry holds trained models with their metadata and stage — candidate, staging, production, archived — and is the answer to what is currently deployed and what it was trained on.
The two together give you the property that matters most operationally: the ability to answer questions about a prediction made three months ago. Which model version, trained on what data, with which features. Investigating a complaint without that is guesswork, and in regulated contexts it is a compliance failure.
8. Evaluation that means something
Offline metrics are necessary and consistently over-trusted.
Evaluate on the right split. For anything temporal, train on the past and evaluate on the future. Random splits flatter the model substantially.
Evaluate on segments, not just the aggregate. A model with good overall accuracy can be systematically wrong for a segment that matters — a region, a customer tier, a device type. Aggregate metrics hide exactly the failures you care about.
Use a metric connected to the decision. Accuracy is rarely the business objective. If false positives and false negatives have different costs, the evaluation should reflect that ratio, and the threshold should be chosen against it rather than left at the default.
Establish a baseline. How does the current rule-based approach perform? A model that beats random and loses to the existing heuristic is not an improvement, and it is remarkable how often this comparison is skipped.
Then test online. Offline metrics predict online performance imperfectly, and the gap is where the interesting problems live. An A/B test measuring the actual business outcome is the only evaluation that settles the question.
9. Serving patterns
| Pattern | How it works | Latency | Suits |
|---|---|---|---|
| Batch | Predict for everything on a schedule, store results | Hours | Scoring, segmentation, recommendations refreshed daily |
| Real-time | Predict on request | Milliseconds | Fraud, ranking, personalisation at page load |
| Streaming | Predict on events as they arrive | Seconds | Anomaly detection, monitoring |
| Embedded | Model runs on the client or device | Immediate | Offline operation, privacy, very low latency |
The advice worth giving: use batch unless you have a reason not to. Batch serving is dramatically simpler — no latency budget, no online feature lookup, trivial to debug because you can inspect every prediction before it is used, and easy to reprocess if something was wrong. A great many problems presented as real-time are satisfied by predictions refreshed nightly.
Real-time is necessary when the prediction depends on information that only exists at request time, or when the decision must be immediate. Those are real requirements, and they are less common than architectures suggest.
10. Latency and cost
For real-time serving, the budget is usually tighter than expected because the model is one component among several.
Where the time goes: fetching features from the online store, often the largest component; preprocessing; the model itself; and network round trips.
What reduces it: precompute features where possible rather than computing at request time; cache predictions for inputs that recur; use a smaller model, since the accuracy difference is frequently smaller than the latency difference; batch requests where the pattern allows; and reduce the model itself through the standard compression techniques where the deployment justifies the effort.
On cost, the observation most teams arrive at late: serving cost usually exceeds training cost over a model's life, often by a large margin. A model trained once and serving millions of predictions daily has its economics determined by inference, not by the training run everyone worried about.
11. Deployment and rollout
A new model is a change with the same risk profile as a code deployment, and it deserves the same care.
Shadow deployment runs the new model alongside the current one on live traffic without using its predictions. Comparison on real data with no risk, and it catches skew and infrastructure problems that offline evaluation cannot.
Canary routes a small fraction of traffic to the new model and compares outcomes before increasing.
A/B test splits traffic and measures the business metric. The only method that answers whether the model is actually better for the thing you care about.
Whichever you use, three requirements: automated rollback triggered by metric degradation rather than by someone noticing; the previous model kept warm so rollback is immediate; and predictions logged with the model version so post-hoc analysis is possible.
12. Monitoring: operational and model
Two distinct categories, and teams routinely implement only the first.
Operational monitoring is ordinary service monitoring: latency, throughput, error rate, resource use, availability of dependencies. Necessary, well understood, and it tells you nothing about whether predictions are any good.
Model monitoring is the part that matters and the part usually missing:
- Input distributions. Are the features arriving now distributed like the training data? Shifts here precede accuracy problems.
- Prediction distribution. A model that predicted five percent positive last month and twenty percent this month has either encountered a changed world or a broken pipeline.
- Feature availability. Nulls and defaults rising is a data pipeline problem showing up as model degradation.
- Accuracy, when labels arrive. The direct measure, and usually delayed — sometimes by weeks, which is why the leading indicators above matter.
- Segment performance. Aggregate accuracy holding while one segment collapses is a real and common pattern.
What to alert on, as opposed to merely chart: sustained input distribution shift beyond a threshold; prediction distribution moving sharply; feature nulls exceeding a bound; and accuracy dropping below a floor once labels are in. Charting everything and alerting on the leading indicators is the right balance.
13. Drift, and what to do about it
Two kinds, with different implications.
Data drift is a change in the input distribution — your customers are different from the ones the model was trained on. The relationship between features and outcome may be unchanged.
Concept drift is a change in that relationship itself. The same inputs now imply a different outcome, because behaviour, the market or the environment changed.
Data drift is easy to detect and often tolerable. Concept drift is harder to detect — it requires labels — and is the one that genuinely breaks a model.
The responses: retrain on recent data, which handles most drift; add features capturing what changed, if the change is structural; adjust the decision threshold, which is often sufficient when the distribution moved but the ranking did not; or rebuild, when the problem has genuinely changed shape.
An important caution: not all drift requires action. Seasonal variation looks like drift and is expected. A monitoring system that alerts every December has trained its users to ignore it.
14. Retraining
Three strategies, and the choice should be deliberate.
Scheduled. Retrain on a fixed cadence. Simple and predictable; retrains when unnecessary and may not retrain fast enough when something changes.
Triggered. Retrain when monitoring detects degradation. Efficient, and it requires monitoring you trust.
Continuous. Retrain constantly on recent data. Suits fast-moving domains and demands the most infrastructure and the most careful guardrails.
Whichever, the non-negotiable rule: a retrained model is a candidate, not a deployment. It must be evaluated against the current production model and promoted only if better. Automatic retraining that automatically deploys will eventually deploy a model trained on a broken pipeline, and it will do so at a moment when nobody is watching.
15. Feedback loops and their dangers
A subtle failure mode worth understanding because it is invisible in ordinary monitoring.
When a model's predictions influence the data it is later trained on, the loop can reinforce its own errors. A recommender shows certain items, those items get engagement because they were shown, the model learns they are popular, and it shows them more. The model becomes increasingly confident about a pattern it created.
The same shape appears in fraud detection, where flagged transactions are investigated and unflagged ones are not, so labels only exist for what the model already suspected; and in credit decisioning, where outcomes are only observed for approved applications.
Mitigations: reserve a small fraction of traffic for random or exploratory decisions, so unbiased data continues to arrive; log what was shown and not shown, not just outcomes; and be explicit about which labels are observed and which are missing by construction, because treating the latter as negatives is a systematic error.
16. Governance and explainability
Requirements vary by domain and some apply broadly.
Lineage and reproducibility — being able to say which model made a decision, what it was trained on, and to reproduce it.
Explainability, at two levels: why the model behaves as it does in general, and why it produced a particular prediction. Techniques exist for both, and the appropriate level depends on the consequence of the decision.
Fairness evaluation across relevant groups, which requires deciding what fairness means for your case — the definitions conflict mathematically, so choosing one is a decision rather than a calculation.
Human review for consequential decisions, with a genuine path to override.
Documentation of the model's intended use, its known limitations and the populations it was evaluated on. A model deployed outside its intended use is a common and avoidable failure.
17. Team shape and ownership
The organisational question determines outcomes as much as the technical one.
The pattern that fails: data scientists build models and hand them to engineers to deploy. The handoff loses context, the engineers cannot debug model behaviour, the scientists cannot debug production, and nobody owns the outcome.
The patterns that work share a property: the people who build a model are involved in operating it. Whether that is scientists with engineering support, engineers with modelling skills, or a platform team providing the infrastructure so product teams own their models end to end, the ownership must be continuous.
Concretely, someone must be on call for a model in production, and that someone must be able to understand why it is behaving as it is.
18. A maturity progression
Most organisations pass through these stages, and skipping ahead rarely works.
Stage one: manual. Models trained in notebooks, deployed by hand, monitored by nobody. Fine for a first model; unsustainable past two.
Stage two: reproducible. Training is a pipeline. Data and models are versioned. Deployment is repeatable. Monitoring is operational only. This stage delivers most of the risk reduction available.
Stage three: monitored. Model monitoring exists, drift is detected, retraining is triggered rather than remembered. Predictions are logged with versions.
Stage four: automated. Retraining, evaluation and promotion are pipelines with guardrails. Rollout is staged with automatic rollback. Multiple models are managed by a small team.
The advice: reach stage two before doing anything else. Reproducibility is the foundation, and elaborate automation on an irreproducible base is worse than none.
19. Twelve mistakes
- Deploying from a notebook. Nothing after it is reproducible.
- Separate feature code for training and serving. Skew is then inevitable.
- Random splits on temporal data. Leaks the future; flatters the evaluation.
- No baseline comparison. A model that loses to the existing heuristic looks like progress.
- Aggregate metrics only. Hides the segment that is failing badly.
- Operational monitoring only. Silent model degradation with green dashboards.
- Automatic retrain plus automatic deploy. Eventually deploys a broken model unattended.
- Real-time serving when batch would do. Enormous complexity for no requirement.
- Not logging predictions with model versions. Post-hoc analysis becomes impossible.
- Ignoring feedback loops. The model reinforces its own errors invisibly.
- No rollback path. A bad model stays live while someone retrains.
- Unclear ownership. Nobody is on call for the thing making decisions.
20. A worked example: a churn model in production
A subscription business predicting which customers are likely to cancel in the next thirty days, so retention offers can be targeted.
Serving pattern: batch. Predictions are refreshed nightly for the whole customer base and written to a table the retention system reads. This was a deliberate choice against an initial real-time design — nothing about the decision requires immediacy, and batch made every subsequent problem easier: predictions can be inspected before use, a bad run can be reprocessed, and there is no latency budget or online feature store to build.
Features and skew. Features are computed by a shared pipeline used for both training and scoring — usage in the last seven, thirty and ninety days, support contacts, payment failures, plan changes, and tenure. Because both paths call the same code, skew is structurally prevented. The first version did not do this: training used an analytics warehouse query and scoring used application code, and the two definitions of "active in the last thirty days" differed by whether the current day was included. The model scored well and underperformed in production by a margin nobody could explain for three weeks.
Leakage, caught during review. An early feature set included the number of days since the last login — which for cancelled customers was computed after cancellation, making it a near-perfect predictor and completely unavailable at prediction time. The evaluation showed an implausibly good result, which is the signal worth treating as a bug rather than a success. Point-in-time correct joins fixed it and the honest performance was substantially lower.
Evaluation. Split by time — train on the previous year, evaluate on the following two months — because a random split leaks future behaviour. Evaluated by segment, which revealed that the model performed well on monthly plans and barely better than chance on annual ones, since annual customers churn at renewal for different reasons. The decision was to scope the model to monthly plans rather than deploy something that would misallocate offers for a third of the base.
Threshold chosen against cost. A false positive means an unnecessary discount; a false negative means a lost customer. Those costs differ by roughly a factor of six, so the threshold was set from that ratio rather than left at the default, which changed the operating point substantially.
Rollout. Shadow for two weeks, comparing predictions against actual outcomes with no offers sent. Then an A/B test where half the flagged customers received the offer and half did not, measuring retention rather than model accuracy. The test is what established the model was worth running — the offline metrics could not have.
Monitoring. Input distributions per feature, prediction rate, null rates, and monthly accuracy once labels arrive. Alerts on sustained input shift and on the prediction rate moving sharply, because both are leading indicators available immediately, whereas accuracy is thirty days delayed by construction.
What monitoring caught. Four months in, the prediction rate jumped from eight percent to nineteen percent overnight. Accuracy would not have revealed this for a month. The cause was an upstream change that altered how support contacts were logged, so the contact-count feature collapsed to near zero for most customers, which the model read as disengagement. A data-layer alert on null rate would have caught it even faster, and was added afterwards.
Feedback loop, acknowledged. Customers receiving retention offers are less likely to churn, so the model's own predictions alter the outcomes it is later trained on. The mitigation is a held-out five percent who never receive offers regardless of prediction, which provides unbiased labels for retraining and costs a small amount of retention to preserve the model's validity.
Retraining. Monthly, triggered by schedule, evaluated against the current production model on the same held-out period, and promoted only if better. Twice in the first year the retrained model was worse and was not promoted, which is the mechanism working rather than failing.
21. Frequently asked questions
What is the single biggest cause of models underperforming in production?
Training-serving skew — features computed differently in the two paths. The differences are mundane, like a window boundary or a null-handling rule, and the effect is a model that tested well and performs worse for reasons nobody can locate. Computing features once in shared code, or logging serving-time features and training on those, eliminates it structurally.
Do we need a feature store?
Not with one or two models. It solves a scale problem — many models and teams reimplementing the same feature definitions — and it is substantial infrastructure. Before that scale, a shared library of feature computations used by both training and serving gets most of the benefit for a fraction of the effort.
Batch or real-time serving?
Batch unless something genuinely requires immediacy. Batch has no latency budget, no online feature lookup, is trivial to inspect and debug because every prediction exists before it is used, and can be reprocessed when something was wrong. A great many problems framed as real-time are satisfied by nightly refresh.
How often should we retrain?
Determined by how fast the domain changes, and best driven by monitoring rather than by a calendar. Whatever the cadence, a retrained model must be evaluated against the current production model and promoted only if better — automatic retraining that automatically deploys will eventually ship a model trained on a broken pipeline.
What should we monitor beyond latency and errors?
Input feature distributions, prediction distribution, feature null rates, segment-level accuracy once labels arrive, and overall accuracy against a floor. The first three are leading indicators available immediately; accuracy is often delayed by weeks, which is exactly why the leading indicators matter.
How do we know if drift needs action?
Distinguish data drift from concept drift. A changed input distribution may be harmless if the relationship to the outcome is stable; a changed relationship is what actually breaks the model. Also account for seasonality, since a monitor that alerts every December trains its users to ignore it.
Who should own a model in production?
Someone who can both understand why it is behaving as it is and change it — which usually means the people who built it stay involved in operating it. The handoff pattern, where scientists build and engineers deploy, reliably produces a model nobody can debug at three in the morning.
Where should a team start?
Reproducibility. Get training out of notebooks and into a parameterised pipeline, version data and models, make deployment repeatable, and log predictions with model versions. That is stage two of the maturity progression and it delivers most of the available risk reduction. Elaborate automation on an irreproducible base is worse than none.
Key takeaways
- Most of a production ML system is not the model. It is data infrastructure and operations.
- Training-serving skew is the dominant failure. Compute features once, in shared code.
- Batch serving unless proven otherwise. Enormously simpler in every dimension.
- Model monitoring is separate from operational monitoring, and it is the one usually missing.
- Retrained models are candidates. Evaluate against production and promote only if better.
- Reproducibility first. Everything else is built on it.
The teams that get value from machine learning are not usually the ones with the best models. They are the ones whose models are reproducible, monitored, owned by someone, and connected to a decision that measurably changes an outcome — which is a set of engineering and organisational properties rather than a modelling achievement.
Enjoyed this article?
Get more engineering insights from ELIVTECH — or talk to us about your project.
Get in touch