You say something to a device and it answers. The gap between those two events is about a second, and in that second sound becomes text, text becomes intent, intent becomes an action and a reply, and the reply becomes sound again. Each of those steps is a substantial engineering problem, and each has a characteristic way of failing that explains most of what frustrates people about voice interfaces.
This article walks the full round trip — wake word, capture, transcription, understanding, generation, speech synthesis — and covers the parts that determine whether a voice product feels responsive or maddening. No code, and no pretending the hard parts are solved.
What you will learn- Every stage between speaking and hearing a reply
- Why wake word detection runs on the device and everything else usually does not
- How speech recognition works and where it reliably fails
- The shift from intent classification to language models
- Why latency dominates the experience, and how it is managed
- Privacy, accents, noise and the other unglamorous realities
- The round trip
- Wake word detection
- Capture and endpointing
- Audio preprocessing
- Speech recognition
- Where recognition fails
- Understanding: the classical approach
- Understanding: the language model approach
- Dialogue state and context
- Taking action
- Generating the response
- Speech synthesis
- Latency, and why it dominates
- Streaming and interruption
- Errors and recovery
- Privacy and what is actually sent
- Accents, languages and fairness
- Designing a voice interface
- Twelve mistakes
- A worked example: one request end to end
- Frequently asked questions
1. The round trip
| Stage | What happens | Where it runs | Typical time |
|---|---|---|---|
| Wake word | Detect the activation phrase | Device | Continuous |
| Capture | Record until the user stops | Device | Duration of speech |
| Preprocessing | Noise reduction, echo cancellation | Device | Milliseconds |
| Recognition | Audio to text | Usually cloud | Streaming; finishes shortly after speech |
| Understanding | Text to intent and parameters | Cloud | Tens to hundreds of milliseconds |
| Action | Query or command execution | Cloud or device | Highly variable |
| Response generation | Producing what to say | Cloud | Hundreds of milliseconds |
| Synthesis | Text to audio | Cloud or device | Streaming; first audio quickly |
The whole chain has a budget of roughly a second before the pause becomes noticeable and about two before it feels broken. Since several stages are irreducible, nearly every design decision in a voice system is about spending that budget.
2. Wake word detection
A device listening for its name is running a small model continuously on the audio stream, on the device, comparing against one specific phrase.
It runs locally for two reasons. Privacy: streaming all audio to a server continuously would be unacceptable, and the local model means nothing leaves the device until the phrase is detected. Practicality: continuous upload of audio from millions of devices is not a viable architecture.
These models are deliberately tiny — small enough to run on a low-power chip indefinitely — which constrains their accuracy and produces the two failure modes everyone has experienced.
False accepts: the device activates on something that sounded similar. Annoying, occasionally alarming, and the reason a visible indicator matters.
False rejects: you say the phrase and nothing happens. More common in noise, with unusual accents, and at distance.
The threshold trades one against the other directly, and there is no setting that eliminates both. Systems typically bias toward false accepts, on the reasoning that an unnecessary activation is less frustrating than being ignored — a judgement not everyone agrees with.
A detail worth knowing: many devices maintain a short rolling buffer of the preceding seconds, so once the wake word fires, the audio just before it is available. This is why saying the phrase and the request as one continuous sentence works.
3. Capture and endpointing
Once activated, the device records — and must decide when you have finished. This is endpointing, and it is one of the largest contributors to a system feeling clumsy.
The simple approach is a silence threshold: stop after some duration of quiet. It fails in both directions. Too short and it cuts you off mid-thought when you pause to think. Too long and every interaction has a dead second at the end.
Better systems use the partial transcript as a signal — whether what has been said so far forms a complete request — combined with prosody, since people's pitch and pace change at the end of an utterance. This is measurably better and still gets it wrong, particularly for anyone who speaks with natural pauses.
The most modern systems avoid the problem for some interactions by processing continuously and allowing interruption, which removes the need for a hard endpoint. That is more expensive and only viable where the audio channel is already open.
4. Audio preprocessing
Raw microphone audio is worse than people expect, and cleaning it up is what separates a device that works across a room from one that requires you to lean in.
Noise reduction suppresses steady background sound — fans, traffic, air conditioning.
Echo cancellation removes the device's own output from the input. Without it, a speaker playing music cannot hear anything else. This is what allows you to interrupt a device mid-sentence, and it is harder than it sounds because the device must subtract a signal that has been altered by the room.
Beamforming uses several microphones to focus on the direction the speech is coming from and attenuate everything else. This is why devices have microphone arrays rather than one microphone, and it is the largest single contributor to far-field performance.
Gain control normalises volume so that a quiet speaker and a loud one both arrive at a usable level.
All of this runs on the device, and the quality of it explains a great deal of the difference between hardware that works in a real kitchen and hardware that works in a demo.
5. Speech recognition
Turning audio into text. The approach has changed substantially and the change matters.
Classical systems combined an acoustic model mapping sound to phonemes, a pronunciation dictionary, and a language model scoring which word sequences are plausible. Three components, each trained separately, each a source of error.
Modern systems are end-to-end: a single neural model taking audio and producing text directly, trained on enormous quantities of paired audio and transcripts. Simpler, substantially more accurate, and better at handling accents and noise because it learns those variations rather than having them handled by a pronunciation dictionary somebody wrote.
Two properties matter for how the rest of the system is built.
Recognition is streaming. Text appears as you speak rather than after you finish, which means downstream processing can begin before the utterance ends. This is a large part of how the latency budget is met.
Early words get revised. A streaming recogniser produces a best guess that it updates as more context arrives — "recognise speech" may briefly have been "wreck a nice beach". Systems that act on partial transcripts must tolerate revision.
Recognition can run on the device for constrained vocabularies and simple commands, and on-device quality has improved considerably. Cloud recognition remains better for open-ended speech because the models are larger than a device can hold.
6. Where recognition fails
Predictably, and in ways worth designing around.
Proper nouns. Names of people, places, businesses and products. The model has seen common words many times and an unusual surname rarely or never. This is the single largest source of error in practical systems, and it is why contact names and music libraries are supplied to the recogniser as context — a technique that measurably improves accuracy on exactly the words that matter most.
Homophones. Distinguishable only by context, and sometimes not then.
Numbers and identifiers. Strings of digits, reference codes, postcodes. Little linguistic context to constrain the guess, and errors are consequential.
Overlapping speech. Two people talking. Separating them is an active research area rather than a solved problem.
Accents and dialects underrepresented in training data. The most consequential fairness issue in the field, and section seventeen returns to it.
Domain vocabulary. Medical, legal and technical terms, unless the model was adapted for them.
The practical mitigation across most of these is context injection: telling the recogniser what words are likely in this situation. A system that knows the user is choosing from a list of twelve options can weight heavily toward those twelve.
7. Understanding: the classical approach
For most of the field's history, understanding meant intent classification and slot filling.
A fixed set of intents is defined — set a timer, play music, check the weather. A classifier assigns the utterance to one. Then named entities are extracted into slots: the duration, the artist, the location.
The advantages are real and explain why this approach persists: it is fast, cheap, predictable, and easy to test. You know exactly what the system can do, and adding an intent is a bounded piece of work.
The disadvantages are equally real. Anything outside the intent list fails, and it fails as either a wrong match to the nearest intent or a blanket apology. Multi-part requests break it. Follow-ups requiring context break it. And the effort of defining, training and maintaining intents grows with every capability added.
8. Understanding: the language model approach
The current generation replaces the classifier with a language model that receives the transcript, the conversation history, and a description of the available capabilities, and decides what to do.
What this changes:
No fixed intent list. The model interprets what was asked and maps it to available actions, including phrasings nobody anticipated.
Compound requests work. "Turn off the lights and set an alarm for seven" is two actions, and the model can issue both.
Context is natural. "What about tomorrow?" resolves against what came before without a special mechanism.
Responses are generated rather than templated, so they can be conversational and can explain things a template could not.
What it costs:
Latency. Generation takes meaningfully longer than classification, and the budget is a second.
Cost per interaction, which matters at scale.
Predictability. The set of things it might do is no longer enumerable, which makes testing harder and makes guardrails necessary rather than optional.
Most serious systems now run hybrid: a fast classifier handles the frequent, simple, latency-sensitive commands — timers, lights, volume — and anything it does not confidently match falls through to a language model. This gets the responsiveness where it matters and the flexibility where it is worth paying for, and it is the architecture worth defaulting to.
9. Dialogue state and context
Conversations depend on what came before, and holding that context is what makes an assistant feel like one.
What needs tracking: recent turns; entities referred to, so "it" and "that one" resolve; the state of any multi-step task in progress; and situational context — where the user is, what device they are on, what time it is.
The classical approach maintains an explicit structured state and updates it per turn. Predictable, and brittle when a conversation goes anywhere unanticipated.
The language model approach keeps the conversation in context and lets the model resolve references. More natural, and it degrades over long conversations as context fills.
Two decisions that need making deliberately. How long does context persist? Too short and the assistant forgets what you just said; too long and a reference from ten minutes ago intrudes on an unrelated request. Most systems expire context after a period of silence, and the right period is short — a minute or two rather than a session.
What persists across sessions? Preferences and history make the assistant more useful and make it a privacy question. This should be visible and controllable rather than implicit.
10. Taking action
Once the request is understood, something must happen — a query answered, a device controlled, a message sent.
Architecturally this is tool use: the system has a set of available capabilities, each described with what it does and what it needs, and the understanding layer selects and invokes them.
Three considerations specific to voice:
Latency varies enormously. A local device command is instant; an external service call may take seconds. For anything slow, saying something immediately — an acknowledgement — is far better than silence, because in a voice interface silence is indistinguishable from failure.
Confirmation for consequential actions. Recognition errors are more likely than typing errors, and voice has no equivalent of seeing what you typed before pressing send. Anything that spends money, sends a message or cannot be undone should be confirmed, and the confirmation should restate what will happen rather than asking a bare yes-or-no.
Failures need to be spoken usefully. "I could not do that" is useless. "The lights are not responding — the hub seems to be offline" tells the person what to do next.
11. Generating the response
What the assistant says is a design problem more than a technical one, and it is where many otherwise competent systems feel wrong.
Voice responses must be much shorter than written ones. Text can be skimmed; speech cannot. A paragraph that reads fine takes forty seconds to hear and cannot be re-read. The discipline is to answer first and elaborate only if asked.
Lists do not work. Nobody retains seven spoken options. Two or three, or a summary with an offer to hear more.
Formatting must be spoken. Numbers, dates, times, abbreviations and units all need converting to how a person would say them, and getting this wrong is immediately jarring.
Confirm what was understood when acting on something consequential, because it is the user's only opportunity to catch a recognition error.
Vary the phrasing. The same acknowledgement every time becomes grating quickly, which is a small thing that materially affects how a product feels after a week.
12. Speech synthesis
Text to audio, and the quality here has changed more than any other stage in recent years.
Older systems concatenated recorded fragments, which produced the characteristic robotic seams. Current systems are neural, generating audio directly, and the best are difficult to distinguish from a recording.
What matters in practice:
Time to first audio matters far more than total generation time, because synthesis streams. The user hears the beginning while the rest is still being produced, so a system can start speaking within a couple of hundred milliseconds of deciding what to say.
Prosody — rhythm, stress and intonation — is what separates natural from merely intelligible. It is also where synthesis still fails, particularly on questions, emphasis and unfamiliar names.
Pronunciation control is necessary for names, technical terms and anything ambiguous. Every serious deployment maintains a pronunciation dictionary for its own vocabulary.
On-device synthesis is now viable for many cases and eliminates a network round trip, which is a meaningful share of the latency budget.
13. Latency, and why it dominates
Voice is unusually sensitive to delay because conversation has rhythm. In human speech, a gap beyond roughly half a second is meaningful — it signals hesitation or confusion. A device that pauses two seconds before answering feels wrong in a way a web page loading in two seconds does not.
Rough targets that make an interaction feel natural: something audible within half a second of the user finishing, and the substantive answer within one and a half.
Where the budget goes, and what can be done about it:
- Endpointing delay — the wait to confirm the user has finished. Reducible with better endpointing, not eliminable.
- Network round trips. Each one is tens to hundreds of milliseconds. Minimising the number of them is often the largest available win.
- Recognition. Largely hidden by streaming, since it finishes shortly after the speech does.
- Understanding. Fast with a classifier, meaningfully slower with a language model.
- Action execution. Highly variable and often the largest component.
- Synthesis. Hidden by streaming if you measure time to first audio.
The techniques that actually help: process partial transcripts so understanding starts before the user finishes; speculatively execute likely actions before the request is complete; speak a filler immediately for anything slow; handle common commands on the device with no network at all; and stream everything, so the user hears the start rather than waiting for the whole.
14. Streaming and interruption
The current generation of voice systems is moving toward continuous, interruptible interaction rather than discrete turns, and it changes the experience considerably.
In the turn-based model, you speak, it processes, it responds, and interrupting is either impossible or awkward. In a streaming model, audio flows continuously in both directions and either party can interrupt.
This requires echo cancellation good enough to hear the user over the device's own output, immediate stopping when the user starts speaking, and handling the awkward case where both start at once — which humans resolve with sub-second timing that systems find genuinely difficult.
The payoff is that interactions stop feeling like commands and start feeling like conversation, and the ability to interrupt a long answer removes one of the most common frustrations with voice interfaces.
15. Errors and recovery
Every stage can fail, and how the system behaves when it does determines whether people keep using it.
Not activated. The user must repeat themselves with no feedback. A visible indicator helps enormously; without one there is no way to distinguish "not listening" from "listening and confused".
Misrecognised. The system heard something else. Recovery requires the user to know what was heard, which is why restating the understood request before a consequential action matters.
Misunderstood. Transcribed correctly, interpreted wrongly. The most confusing failure for users, because everything looked right.
Action failed. Understood correctly and could not be done. The easiest to handle well and frequently handled worst.
Principles that hold up: fail informatively, naming what went wrong; offer a next step rather than a dead end; do not repeat the same failure message, since a second identical response tells the user the system is not adapting; and escalate to a different modality when voice is not working — showing a list on a screen, or offering to send a link.
16. Privacy and what is actually sent
The concern is widespread and the reality is more specific than the folklore.
In a typical always-on device: the wake word model runs locally on a rolling buffer that is continuously overwritten; nothing is transmitted until the wake word fires; once it fires, audio from shortly before the phrase through the end of the utterance is sent; and that recording may be retained depending on settings.
So the common fear — continuous transmission of everything — is not how these systems work, for practical as much as ethical reasons. The genuine concerns are different and worth taking seriously.
False activations transmit unintended audio. A device that mishears its wake word during a private conversation sends that fragment. This happens, and it is the realistic version of the concern.
Retention. Recordings kept for model improvement, sometimes reviewed by people. Whether this is disclosed clearly and controllable is the real question.
Inference from queries. What you ask reveals a great deal, and that is data with commercial value regardless of the audio.
What responsible design looks like: clear indication when listening and transmitting; local processing wherever feasible; short retention by default; genuine deletion controls; and being explicit about human review rather than burying it.
17. Accents, languages and fairness
Recognition accuracy varies substantially across speakers, and the variation is not random.
Systems perform best on the accents most represented in training data — typically standard varieties of widely spoken languages. Performance degrades for regional accents, second-language speakers, older and younger speakers, and anyone with a speech difference.
This is a data problem rather than a technical limitation, and it has consequences beyond inconvenience. A voice system that works less well for some populations is a product that serves them worse, and where voice is the interface to something important, that is an access issue rather than a usability one.
What helps: training data that deliberately covers the range of speakers; evaluation broken down by speaker group rather than reported as a single accuracy number; on-device adaptation to an individual speaker over time; and always providing an alternative to voice.
That last point deserves emphasis. Voice should be an option, not the only path. Any system where a task can only be completed by speaking excludes people, and the exclusion is invisible in aggregate metrics.
18. Designing a voice interface
- Answer first, elaborate second. Speech cannot be skimmed.
- Keep responses under about fifteen seconds unless the user asked for more.
- Never present a long list. Two or three options, then offer more.
- Confirm consequential actions by restating what will happen.
- Acknowledge immediately when something will take time.
- Accept many phrasings. People will not learn your syntax.
- Make failures actionable, and never repeat an identical error message.
- Show state visibly where a screen or indicator exists.
- Allow interruption wherever the architecture permits.
- Provide a non-voice alternative for everything.
- Test with real speakers in real environments. A quiet office with the development team is not a test.
19. Twelve mistakes
- Responses written for reading. Fine on a page, interminable aloud.
- Reading out lists. Nobody retains them.
- Silence during slow actions. Indistinguishable from failure.
- No confirmation on consequential actions. Recognition errors are more likely than typing errors.
- Aggressive endpointing. Cuts off anyone who pauses to think.
- Ignoring proper nouns. The largest source of recognition error, and fixable with context.
- Unformatted numbers and dates. Immediately jarring.
- Identical acknowledgement every time. Grating within a week.
- Context that persists too long. An old reference intruding on a new request.
- Testing only in quiet rooms with the development team. Misses noise, distance and accent entirely.
- Reporting a single accuracy number. Hides the populations the system fails.
- Voice as the only path. Excludes people, invisibly.
20. A worked example: one request end to end
Someone in a kitchen, with a radio on, says: "Hey device — add oat milk to the shopping list and remind me to call Priya at six."
Wake word. The local model, running continuously on a rolling buffer, matches the activation phrase. The indicator lights. Crucially, the buffer already holds the two seconds before the match, so the beginning of the request is not lost — this is why the sentence works as one continuous utterance rather than requiring a pause.
Preprocessing. The microphone array beamforms toward the speaker and attenuates the radio, which is coming from a different direction. Noise reduction handles the extractor fan. Echo cancellation is idle, since the device is not producing output. Without the beamforming, the radio would be roughly as loud as the speech at this distance and recognition accuracy would drop sharply.
Recognition, streaming. Text appears as the person speaks. "Add oat milk" is transcribed confidently. "Priya" is the uncertain part — an uncommon name, minimal context. Because the device has supplied the user's contact list as recognition context, "Priya" is weighted heavily and comes back correct. Without that, it would plausibly have been "prea" or "pria", and the reminder would have been useless.
Endpointing. The person pauses briefly after "shopping list" — a natural mid-sentence breath. A naive silence threshold would have cut the utterance there and processed half the request. The endpointer uses the partial transcript to recognise that the sentence is grammatically incomplete and the pitch has not fallen, and waits. After "six" the pitch falls and the transcript parses as complete, so capture ends.
Understanding. This is a compound request — two distinct actions in one utterance — which a classical intent classifier would have handled badly, matching one intent and dropping the other. The language model interprets it as two actions with their parameters: add an item to a named list, and create a reminder with a person and a time. It also resolves "six" to the next occurrence of six in the evening, using the current time as context.
Action. Both are executed. The list addition is local and instant. The reminder is created against the calendar service and takes about four hundred milliseconds. Nothing is spoken yet, and the total elapsed time since the person stopped speaking is under a second.
Response generation. A written interface would confirm both items in detail. Spoken, that would take too long. The response is: "Added oat milk, and I'll remind you to call Priya at six." Under three seconds to say, confirms both actions, and — importantly — restates the name, which is the user's only chance to catch a recognition error before the reminder fires at a useless time.
Synthesis. Streams, so the first audio plays about two hundred milliseconds after the response text was decided. The name has a pronunciation entry so it is not mangled. Total time from the person finishing to hearing the first word: roughly nine hundred milliseconds, which sits inside the range where the exchange feels like a conversation rather than a transaction.
What would have broken it. Endpointing on the mid-sentence pause would have processed only the first half. Omitting contact names from recognition context would have produced a wrong name silently. Using a fixed intent list would have handled one action and dropped the other. Confirming both actions in full detail would have taken eight seconds and been worse than saying nothing. And handling the whole thing in the cloud with a separate round trip per stage would have pushed the total past two seconds, at which point the person repeats themselves and the interaction fails entirely.
21. Frequently asked questions
Is the device always listening?
It processes audio continuously on the device to detect the wake word, on a buffer that is constantly overwritten, and transmits nothing until the phrase is detected. The realistic privacy concern is not continuous transmission — which would be impractical as well as unacceptable — but false activations sending unintended fragments, and how long recordings are retained afterwards.
Why does it mishear names so often?
Because a model that has seen common words millions of times has seen an unusual surname rarely or never, and there is little surrounding context to constrain the guess. The fix is supplying likely names — contacts, playlists, device names — to the recogniser as context, which measurably improves accuracy on exactly the words that matter most.
Why does it cut me off?
Endpointing decided you had finished. Simple silence thresholds do this whenever you pause to think. Better systems use the partial transcript and intonation to distinguish a mid-sentence breath from a completed utterance, and they still get it wrong for anyone whose speech rhythm differs from the training distribution.
Should we use a language model or intent classification?
Both. A fast classifier for the frequent, latency-sensitive commands — timers, lights, volume — and a language model for anything it does not confidently match. This gets responsiveness where users notice it and flexibility where it is worth paying for, and it is the architecture most serious systems converge on.
What latency should we target?
Something audible within half a second of the user finishing, and the substantive answer within about a second and a half. Beyond two seconds people assume failure and repeat themselves. Since several stages are irreducible, the practical levers are minimising network round trips, streaming everything, and acknowledging immediately when an action will be slow.
How much should run on the device?
Wake word always. Preprocessing always. Recognition and synthesis increasingly can, and doing so removes network round trips that are a meaningful share of the latency budget. Open-ended understanding generally still needs the cloud because the models are larger than devices hold — though handling the twenty most common commands locally covers a surprising proportion of real usage.
How do we handle accents fairly?
Train on data that deliberately covers the range of speakers, and evaluate by speaker group rather than reporting one aggregate accuracy figure — an aggregate hides exactly the populations you are failing. Support on-device adaptation to individual speakers, and always provide a non-voice path, because a task that can only be completed by speaking excludes people invisibly.
What is the most common design mistake?
Writing responses as though they will be read. A paragraph that scans fine on a screen takes forty seconds to hear, cannot be re-read, and cannot be skimmed. Answer first in one sentence, elaborate only if asked, and never read out a list of more than three items.
Key takeaways
- The whole round trip has a one-second budget, and nearly every design decision is about spending it.
- Wake word runs on the device; almost everything else has historically not, and that is changing.
- Proper nouns are the biggest recognition problem, and supplying context is the fix.
- Hybrid understanding wins — a fast classifier for common commands, a language model for the rest.
- Write for the ear. Short answers, no lists, restate consequential actions.
- Voice must never be the only path, because recognition quality is not equal across speakers.
A voice interface that feels effortless is the product of a great deal of unglamorous work at every stage — microphone arrays, endpointing heuristics, context injection, streaming at every boundary, and responses written by somebody who listened to them rather than read them. None of it is visible when it works, which is exactly the point.
Enjoyed this article?
Get more engineering insights from ELIVTECH — or talk to us about your project.
Get in touch