While much of the current focus in GenAI is placed on text-based systems, voice remains a comparatively underexplored modality. However, this is starting to change: in the past two years, major labs have released native audio models and real-time voice APIs, from OpenAI's Realtime API to Google's Gemini Live, and more applications are being built on top of them. Thus, understanding the building blocks of these systems is becoming a useful skill for engineers.
This article outlines some of the main architectures used for developing applications with voice agents, and the practices that make them work in production. It starts by laying out the architectural patterns: the three main kinds of pipelines and the two turn-taking patterns. Then it turns to deployment, using latency as the lens: why it dominates the user experience, which metrics to track, how to choose a transport, and finally a reference implementation of a cascading pipeline that serves as a starting point to discuss the best practices worth keeping in mind when shipping this type of application to production.
Architectural decisions for building voice agents
There are multiple approaches to developing applications with voice agents. These encompass both the kinds of models being used and the way interactions between the user and the system are handled.
When considering which approach to implement, one of the main properties to keep in mind is latency. Spoken interaction is immediate by nature: in ordinary human conversation, only around 200 ms pass between one participant finishing and the next one starting. To preserve that experience, voice agents must answer as quickly as possible. At the same time, however, keeping the system reliable and deterministic is just as important, since an answer that arrives fast is of little use if the system doesn't behave as expected.
The next section covers three pipeline architectures for voice applications: cascading, speech-to-speech, and hybrid approaches. All three trade off the same three properties: latency, control, and maturity. The section after that deals with turn-taking, i.e., deciding when a speaker has finished and the other party should begin, along with the patterns available for handling it.
Common Pipelines
Cascading pipeline
This is probably the most common pipeline powering voice agent applications. It consists of a three step pipeline:

First, a Speech to Text (STT) model is used to generate a transcription of the user's audio. The result is then fed to an LLM which produces a text response. Finally, this response is fed to a Text to Speech (TTS) model which is in charge of generating the spoken version of the response.
The reason why this is one of the most common pipelines is that it ensures maximum control over the quality of the response. Having an LLM producing the answer guarantees that best practices for prompting and AI engineering can be applied, potentially including tool calling, RAG, or prompt chaining, among others. It also keeps the system debuggable and each stage independently swappable, which is why it remains the default choice for production systems.
However, the main drawback of this approach concerns latency. Having a three step pipeline makes this approach slow. At the time of writing this article, however, many models have been developed which deliver low latency results, making it possible to develop cascading pipelines well below 1 second between the message sent by the user and the first byte returned by the TTS model.
Speech to speech
Another approach consists of using just one model for the whole system:

These models take audio tokens as input and return audio tokens as output, without ever producing an intermediate text representation of the conversation. The whole exchange stays in the audio domain, and that is where both the advantages and the limitations of this approach come from.
The reasoning capabilities of these kinds of models are weaker than those of text-based models. Thus, and contrary to the main advantage of the cascading pipeline, the quality of the conversation delivered by this system may be lacking, and makes it specially well suited for simple interactions where deep reasoning or agentic capabilities are not required for the success of the use case.
But, as only one model is needed for the whole interaction, this implementation guarantees the most natural flow of conversations, with latencies below the 0.5 seconds threshold. This makes it a very good fit for conversations with low cognitive complexity, since they are straightforward to handle while also delivering top latency performance.
Full-duplex models
Most speech-to-speech systems are still half-duplex: they operate in strict turns, listening and then speaking, with turn boundaries managed from the outside. Full-duplex models, on the contrary, process incoming and outgoing audio simultaneously: they are always listening while they speak, the way a phone call carries both directions at once. This makes turn-taking, backchannels ("uh-huh") and interruptions native behaviors of the model rather than things the surrounding system has to engineer.
Notable examples include GPT-Live (commercial), Kyutai's Moshi (the open-source reference that first demonstrated the approach), and NVIDIA's PersonaPlex, which keeps full-duplex naturalness while restoring the customizable voice and role that earlier full-duplex models gave up. The trade-offs are the same as for speech-to-speech in general: weaker reasoning and less control than a text LLM, only more pronounced, which is why full-duplex remains a frontier choice rather than a default.
Half-cascading or Hybrid approaches
Apart from the models presented above, some other variants exist which collapse some of the steps of the cascading pipeline by leveraging multimodal models which can operate both with text and audio tokens.
On one hand, models like GPT-4o or Qwen2-Audio are able to process audio tokens directly and return text responses, effectively merging the STT and LLM stages.

On the other, some models are capable of receiving text input and producing audio tokens directly, merging the LLM and TTS stages.

Of the two, the audio-in / text-out variant is the more compelling. Feeding audio directly to the model preserves paralinguistic information (e.g., tone, emotion, emphasis, hesitation) that a plain transcript discards, and it removes one hop of latency. The trade-off is that you give up independent control over the transcription stage and inherit whatever audio understanding the multimodal model happens to have.
A word of caution, though: this is the least battle-tested of the three approaches. In EVA-Bench, a recent end-to-end benchmark covering all three architectures, only two of the twelve evaluated systems were hybrid, against seven cascading and three speech-to-speech. Both hybrids also scored within the cascading range on experience metrics, which the authors read as a sign that hybrid systems may not inherit the latency advantages of end-to-end speech-to-speech models, though they note that the sample is too small to be conclusive. For a system that needs to get started simply and reliably, hybrid approaches are better thought of as an optimization to reach for later rather than a starting point.
The following table summarizes the three pipelines along the dimensions introduced earlier: latency, control, and maturity.

Turn-taking patterns
One final decision to make when developing a system of this kind relates to how turn-taking will be managed.
Push-to-Talk
In this pattern, the main idea is to handle turns by displaying a button on the UI of the application where the user can click whenever they want to talk. Once they are done speaking, they can click the same button to let the system know that it's the AI's turn to start talking.
This is probably the most straightforward way to implement turn-taking since it doesn't delegate the detection of turns to any additional model. Instead, it's all handled by the interface of the application: the button is the endpointer. However, this may seem less natural, given the fact that the user needs to click every time they want to start speaking or finish doing it.
Barge-ins
A second approach consists of letting the user speak on top of the agent and expecting the agent to stop talking and switch to listening. This is the most natural approach since it exactly mirrors the way conversations happen between humans. Additionally, this approach has the advantage of allowing the user to stop the agent whenever its response is irrelevant.
Full-duplex models are best suited to handle these interactions. Because they listen continuously, they can halt their response as soon as the user starts speaking. Cascading pipelines have no equivalent built-in mechanism and need an additional component: most commonly a Voice Activity Detection (VAD) model (see this article for an example), running continuously alongside the pipeline. VAD signals when speech is present; turning that into a turn boundary — deciding the user has actually finished rather than paused — takes an additional silence threshold or a dedicated turn-detection model.
This approach, however, introduces a series of challenges worth pointing out. With push-to-talk, the mute / unmute button delimits the turns and interruptions are handled by the interface at no cost. Once the microphone is always open, all of that becomes the system's responsibility, and many complications might appear:
- Backchannels: people say "mm-hm", "yeah" or "right" while listening, without intending to take the turn, and a naive VAD stops the agent on any sound.
- History reconciliation: at the moment of the interruption the LLM has generated more text than the TTS has spoken, so the conversation history has to be truncated to what the user actually heard.
- Buffer flushing: audio sits in the TTS output, in the playback buffer and in flight over the transport, and all of it has to be dropped quickly or the agent keeps talking after being cut off.
- Interruption latency: the delay between the user starting to speak and the agent falling silent is a metric of its own.
Knowing when a user is actually done rather than merely pausing mid-thought is one of the hardest UX problems in voice, so it's reasonable to start with push-to-talk and earn barge-in later.
Deployment and best practices
Having covered the architectural patterns, we now turn to what it takes to ship one of them. This section focuses on a use case implementing a cascading pipeline with push-to-talk turn handling, the lowest-lift combination that still delivers strong performance.
The common thread for everything that follows is latency. In voice, latency isn't just one quality among many; it is central to the user experience. A response that would be perfectly acceptable in a chatbot feels broken when a human is waiting for a spoken reply in real time. Every decision below (e.g., which metrics to track, which transport to pick, how to structure the implementation) comes back to defending a latency budget.
Metrics: measuring the right kind of latency
Engineers coming from chatbot applications will be familiar with Time to First Token (TTFT): the delay until the LLM emits its first token. In voice, that metric is necessary but not sufficient, because the user doesn't consume tokens; they consume audio. The metric that actually maps to the user's experience is Time to First Audio (TTFA): the delay between the user finishing their turn and the first byte of spoken response reaching their ears.
One thing to keep in mind is that, in cascading and hybrid pipelines, TTFA is the sum of the latencies of every stage: STT, LLM, TTS, and the transport overhead between them. This means that there are multiple entry points to optimize when aiming for lower latency, but it also means that the latencies accumulate: no stage can compensate for a slow neighbour.
Additionally, when measuring the overall latency of the whole system, doing so at the tail is usually a better idea than at the median. A single turn that stalls for three seconds breaks the user experience, even if every other turn in the exchange was fast. Optimizing the p95 of TTFA is a better proxy for perceived quality than optimizing the median.
Although metrics to measure latency dominate, these are not the only ones worth tracking. Each step of the selected pipeline includes metrics of their own. Some examples are:
- Word error rate: used to measure the performance of STT models. In production, what usually matters is narrower, whether names, numbers and domain-specific terms survive transcription, since those are what the LLM acts on.
- Speaker similarity: used to measure how closely the voice produced by the TTS model matches the reference speaker. Pronunciation and naturalness belong to the same family.
- Throughput: tokens per second for the LLM, and real-time factor for the TTS model. Audio has to be generated faster than it is played back, or the response stutters mid-sentence.
- Noise levels: measured on the incoming audio, since background noise degrades transcription quality and determines whether noise suppression is worth adding to the pipeline.
Each of these is worth collecting in two places: offline, against a fixed evaluation set before shipping, and online, against production traffic, where accents, noise and phrasing rarely resemble the test set
Transport: WebRTC vs. WebSocket
Because latency accrues in the connections between components, the choice of transport is a first-class decision. The requirement is a bidirectional channel that keeps the delay between client and server both low and predictable. There are two realistic options for moving audio between the browser and the services:
- WebSockets is a bidirectional protocol running over TCP. It's reliable, ordered, and simple to set up. TCP delivers a byte stream in order, however, so if a packet is lost, everything that arrived behind it is held in the receive buffer until the missing one is retransmitted. This is known as head-of-line blocking, and it means that a single lost packet stalls playback for a full round trip, significantly impacting perceived latency.
- WebRTC is a stack for real-time media, exposed in the browser as an API and carrying media over UDP. Because UDP makes no delivery guarantee, a missing packet does not hold back the ones behind it: the stream continues and the decoder conceals the gap. This is the preferable trade-off for audio, where a lost packet is imperceptible to the listener but a stalled one is not.
When working on cascading pipelines, both the STT and the TTS connections carry audio: the first sending the user’s speech for transcription, and the second one bringing a synthesized response back from a provider to the browser. Both are equally sensitive to latency, so on principle WebRTC is the better fit for both, since TCP's reliability is useless when it makes audio packets arrive late. In practice, however, the choice is often dictated by what each provider exposes: streaming STT vendors are overwhelmingly WebSocket-first, while WebRTC tends to be the natural option for playing the response back, where gap-free, jitter-buffered audio is what the user directly judges.
In most cases, then, ending up with a different transport on each connection is not a quality decision, but a consequence of what the providers make available. This is why the diagram in the next section doesn't specify any transport: the reasonable default is to use whatever each provider supports best, and to reach for WebRTC when the connection is under your control and latency is critical.
A reference implementation
The following sequence diagram outlines one concrete way to wire up a cascading, push-to-talk pipeline. This isn’t intended to be a rigid prescription but a proposal for a design which achieves a good tradeoff between effort and performance. The best practices in the rest of this section assume this implementation.

The flow starts when the user opens a session. The frontend asks the backend for the session tokens it needs to reach the external providers, and the backend returns them. From that point on, whenever the user speaks, the frontend opens a streaming connection to the STT provider and receives the transcript as it is produced. That transcript is sent to the backend, which requests a completion from the LLM and streams it back to the frontend as the tokens arrive. The frontend then opens a streaming connection to the TTS provider and plays the resulting speech stream back to the user. Note that the frontend is in charge of a significant portion of the whole system. Both STT and TTS services are interacted with from the frontend. The backend limits itself to returning session tokens and brokering the LLM stream.
Let the frontend own the transports
The first practice worth highlighting is making the frontend handle the full-duplex transports to the STT and TTS providers, which makes scaling much easier. Some popular frameworks like Pipecat showcase end to end model orchestration purely in Python. The problem with this approach is that, when deploying such an implementation, the user will need to establish a stateful WebRTC connection to this Python server (which could be served using a framework like FastRTC). This will, in turn, require some specific scaling decisions, like using sticky sessions as a load balancing strategy or implementing a worker-per-session approach.
When the frontend handles these connections, it's the user's browser that talks to the providers' servers directly. Your services stay nice and stateless, while the browser handles the complicated transports.
Stream everything, and aggregate into sentences
Second, streaming is important. Waiting for a full stage to complete before starting the next one throws away the latency headroom the whole pipeline was designed to protect.
One detail to keep in mind: when streaming LLM responses to TTS models, a common practice is to aggregate tokens into sentences before streaming them to the TTS model. This improves the TTS model's performance, especially regarding voice cadence in sentence completion. If latency is a top priority, then eagerly streaming tokens is also possible, with potential downside on voice generation quality.
Prompt the LLM for speech, not text
An LLM prompted for a chatbot will happily emit markdown, bulleted lists, URLs, and long paragraphs all of which sound terrible when read aloud by a TTS model. It's worth explicitly instructing the LLM that its output will be spoken: no formatting, numbers and units spelled out, and short, conversational turns. It also helps to inform the model of its input and output modalities so it can adjust its register accordingly. This is one of the cheapest, highest-impact changes available.
Signal slow operations
When a turn requires a tool call or any other slow operation, the user is left listening to silence, which reads as the system being broken. A good practice is to have the agent emit a short filler message ("let me check that for you") while the work happens in the background, paired with a watchdog that times out gracefully if the operation takes too long. Relatedly, STT or TTS failures should be handled by recovering conversationally rather than surfacing a raw error.
Start Simple, Add Complexity Where It Matters
Building a voice agent is largely a matter of deciding how much complexity the use case actually justifies. A cascading pipeline with push-to-talk turn handling covers a good number of them, keeps every stage under control, and is the easiest thing to debug when something goes wrong. Speech-to-speech models, full-duplex conversation, barge-in and hybrid pipelines all buy naturalness, and all of them charge for it in control, maturity or engineering effort.
Latency is what ties these decisions together. It is the reason the architecture is worth choosing carefully, the thing the metrics are there to protect, and what the transport and streaming choices ultimately defend. Start from the simplest combination that fits within the latency budget, measure it properly, and add sophistication only where the experience clearly asks for it.


