A failed delivery is usually a data problem wearing a logistics costume. Somebody typed a building name where a street should be, abbreviated a locality in a way no reference file recognises, or transposed two digits in a postal code. ELIVTECH built a platform that resolves those inputs into a canonical, geocoded address in a single API call, and treats every human correction as training data for the matching that follows.
The client provides address validation as a service to commerce and logistics companies who call it during checkout, order ingestion and route planning. Its original implementation matched against a reference table with progressively looser queries, which meant an unusual input either resolved slowly or not at all, and the latency was unpredictable enough that customers were reluctant to place it in a checkout path. Reference data was refreshed by replacing the table wholesale, taking the service offline. ELIVTECH rebuilt matching on Elasticsearch with a purpose-built analysis chain, moved reference ingestion to a Python pipeline that publishes behind an index alias, and added a feedback loop so corrections confirmed by end users improve subsequent matches.
At a glance
The challenge
A matching service too slow and too unpredictable to be trusted inside a checkout.
Latency that varied with input quality
Clean inputs matched on the first query while messy ones fell through a cascade of progressively looser attempts. The worst cases took orders of magnitude longer than the best, and because customers had to budget for the worst case, the service could not be placed anywhere a person was waiting.
Real input that no reference file anticipates
People write addresses with local abbreviations, landmark references, inconsistent transliteration and building names used instead of numbers. Exact and near-exact matching handled the tidy minority and failed the rest, which was precisely the population that generated failed deliveries. The inputs the service most needed to resolve were the ones it was least equipped to interpret.
Reference updates required downtime
New reference data was loaded by replacing the live table, so refreshes were scheduled as outages and therefore performed rarely. Between refreshes the service matched against increasingly stale data, and newly developed areas were the addresses most likely to be missing.
Corrections discarded after use
When an operator manually fixed an address, the correction was applied to that order and forgotten. The same input arrived the following week and failed identically. The platform was generating the exact signal that would improve it and throwing that signal away.
Results without a confidence signal
The API returned a best guess with no indication of how certain it was. A confident match and a speculative one looked identical to the calling system, so clients could not decide when to accept automatically and when to ask the customer to confirm.
No isolation between client workloads
All clients shared one capacity pool with no rate limiting. A single customer running a bulk cleansing job consumed the service for everyone, and the customers affected were often those calling it from a live checkout where latency was least forgivable.
Our solution
One analysis chain built for messy input, publishing behind an alias, learning from every correction.
Failure Analysis and Address Modelling
Four weeks analysing real unresolved inputs to categorise how addresses actually fail: abbreviation, transliteration variance, landmark substitution, component transposition and genuine absence from reference data. Each category needed a different treatment, and separating them prevented the usual outcome where loosening matching to catch one class silently degrades precision for all the others.
Elasticsearch Matching and Confidence Scoring
Reference addresses are indexed with a custom analysis chain combining component-aware tokenisation, phonetic matching for transliteration variance, and curated synonym sets for local abbreviations. One query with weighted clauses replaced the cascade, so a difficult input costs roughly what an easy one costs. Every result carries a normalised confidence score and the components that did or did not match, which lets a calling system decide between accepting silently, prompting the customer, or routing to manual review.
Python Ingestion with Alias-Based Publishing
Reference sources are ingested by a Python pipeline that standardises components, deduplicates across overlapping sources, geocodes to a common projection and validates the result against the previous edition before publishing. Each build writes to a fresh index and goes live by swapping an alias, so refreshes happen during trading hours and a bad edition is reverted by pointing the alias back. Where two sources disagree on a geocode, the pipeline keeps both with a recorded precedence rather than silently choosing, so a downstream correction can be traced to the source that produced it.
Feedback Loop and Tenant Isolation
Confirmed corrections are captured as candidate synonyms and alternate forms, reviewed, and folded into the next index build, so the service improves from the traffic it handles. Client workloads are separated into interactive and bulk paths with per-tenant rate limits, keeping a bulk cleansing job from consuming the capacity a checkout call depends on. Bulk submissions are accepted onto a queue and processed asynchronously against the same index, so the two workloads share matching logic without sharing a latency budget.
Technology stack
Results
Measured once all client traffic was served by the new matching and publishing pipeline.
Before and after: platform engineering measures
- Cascading queries with latency that varied by input quality
- Abbreviations and transliteration variance largely unmatched
- Reference refresh performed as a scheduled outage
- Operator corrections applied once and discarded
- Best-guess results returned with no confidence signal
- One shared capacity pool with no rate limiting
- Bulk jobs degrading latency for live checkout callers
- One weighted query with predictable cost across input quality
- Phonetic matching and curated synonyms handle local variance
- Alias-based publishing refreshes reference data during trading hours
- Confirmed corrections folded into the next index build
- Every response carries a confidence score and matched components
- Interactive and bulk paths separated with per-tenant limits
- A bad reference edition reverted by pointing the alias back
Project timeline
Failure Analysis and Address Modelling
Categorisation of unresolved inputs from production logs, canonical address component schema, reference source assessment and licensing review, and the target Elasticsearch analysis design.
Matching Engine
Custom analyser build with component-aware tokenisation, phonetic matching and curated synonym sets, weighted single-query design, confidence score calibration against a labelled evaluation set, and component-level match reporting.
Ingestion Pipeline
Python ingestion across reference sources, component standardisation, cross-source deduplication, geocoding to a common projection, edition validation against the previous build, and alias-based publishing with rollback.
API Platform and Feedback Loop
Laravel API surface with authentication and per-tenant rate limiting, interactive and bulk path separation, correction capture and review workflow, and the promotion path from confirmed correction into the next index build.
Evaluation and Client Migration
Regression evaluation of match quality across input categories before each publish, load testing on mixed interactive and bulk profiles, monitoring on latency percentiles and confidence distribution, then client-by-client migration.
Key takeaways
What shaped the engineering decisions
- One weighted query beats a cascade: Progressive loosening was rejected because it makes latency a function of input quality, which is the one thing a checkout integration cannot tolerate. A single query with weighted clauses gives difficult inputs roughly the cost of easy ones.
- Categorise failures before tuning matching: Loosening thresholds to catch one failure class silently degrades precision elsewhere. Separating abbreviation, transliteration, transposition and genuine absence meant each got a targeted treatment with measurable effect.
- Publish behind an alias, always: Replacing the live table made refreshes outages, which made them rare, which made the data stale. Building into a fresh index and swapping an alias turned a quarterly outage into a routine operation with an instant revert.
- Validate an edition before it goes live: Each build is checked against the previous edition for unexpected volume shifts and match-rate regressions, because a silently degraded reference file is worse than an old one that everyone knows is old.
- Return confidence, not just an answer: Without a score, callers cannot distinguish a certain match from a speculative one. Returning a calibrated confidence and the matched components lets each client set its own threshold for auto-accepting.
- Separate interactive from bulk: Sharing one pool meant a cleansing job could degrade a checkout call. Splitting the paths and applying per-tenant limits made the latency guarantee something the platform enforces rather than hopes for.
Want results like these?
Let's discuss how ELIVTECH can drive measurable outcomes for your business.
Start your project