ResourcesBlogs
Anatomy of a Warm Transfer

Anatomy of a Warm Transfer

Shashank Tyagi, Software Engineer - AI

How our voice agent hands off a live call without losing context

Ather and Helix are our AI customer engagement platform for pharma, supporting healthcare professionals and patients across voice, chat, and web. Any call it handles can reach a point where a human has to take over: a question outside the approved content, something that sounds like an adverse event, or a caller who simply asks for a person. How we hand that call over matters as much as anything the agent said before it.

Most AI voice agents do what telephony calls a cold transfer: the agent connects the caller to a human and steps aside. The human picks up blind, and the caller explains everything a second time. That is a poor experience anywhere, and when the conversation involves clinical detail already disclosed once, it is an unnecessary risk. It is also the only transfer most agent frameworks make easy.

A warm transfer is what every contact center does instead, and it is the standard we hold Ather to. The agent puts the caller on hold, calls the human, privately explains who is calling and why, and only then connects them. The human knows the situation before saying hello. This article is an anatomy of that transfer, and of why each part of it is harder than it sounds.
‍


In a warm transfer, the caller is on hold while the agent and the human talk. Each of the three participants hears something different.
‍

Two ways to build it

There are two ways to give three people a call in which only two of them can hear each other at any moment.

The first is to create a second room. The agent and the human talk there while the caller waits in the original, and when the briefing ends, the human is moved across. It is the design that suggests itself first, because it matches how the conversation feels: the private part happens somewhere else.

The problem is what somewhere else means in a media system. A room lives on one server node, and two rooms can sit on different nodes, so moving a live participant between them can mean migrating an active audio session across machines. That makes the behavior of a transfer depend on how many servers happen to be running.

The second way is to keep everyone in one room and change who can hear whom. Nobody moves. There is no boundary to cross, so nothing migrates, and a transfer behaves identically whether the cluster has one node or twenty.

That is the design we built. The rest of this article is what it takes to make it hold up.

‍

Hearing is a subscription

In our media server, every participant publishes one or more audio tracks and separately subscribes to the tracks they want. Being in a room does not mean hearing everyone in it. A participant hears exactly what they are subscribed to, and the server can change that set at any time, for anyone, without anybody leaving or joining.

That is the property we built on. A warm transfer is four phases, and the only thing that changes between them is who is subscribed to what.

First, caller and agent hear each other: an ordinary call. On transfer we move to hold, unsubscribing the caller from the agent's voice and subscribing them to a hold-music track while we dial the human. When the human answers we enter the briefing: agent and human hear each other, the caller still hears only music. When the human is ready we connect, subscribing caller and human to each other and unsubscribing the agent from both.

No participant is created, removed, or moved at any point. The agent stays in the room throughout, hearing nobody once the handover completes, because it is what takes the caller back if the human drops off.


The agent is unsubscribed during hold so it does not transcribe its own hold music. It stays in the room after connect because it is what returns the caller if the human hangs up.

‍

A fifth transition is not a phase. We call it cancel, it is legal from all four states, and it restores the original subscriptions so the caller is talking to the agent again. Every failure below resolves to it.

What made it hard

Making this hold up in production ran into three constraints that were not obvious at the start.

Hold does not exist. Our media layer has no concept of hold, and carrier hold signals are discarded before they reach us. So hold is something we construct. The agent publishes two tracks from the start of every session: its voice, and a track carrying hold music. Placing the caller on hold means unsubscribing them from the first and subscribing them to the second. We publish the music at session start rather than at transfer time, so the transfer never waits on a publish. The agent itself must also be unsubscribed from everything during hold. If it can hear its own hold music, the speech-to-text pipeline dutifully transcribes it, and the model spends the entire hold reasoning about noise.

Subscription changes do not persist. This shaped the design most. Subscribing to every track is the default, and removing a subscription is a single instruction applied to live session state on one server. It is not stored as a rule. Whenever that state is rebuilt it comes back with the defaults, and it is rebuilt more often than we expected: a deploy does it, so does an autoscaling event, a browser leg dropping and recovering, and the routine reconnection WebRTC performs on its own. None are failures, and two of them we schedule ourselves.

No error is raised and nothing is logged. The call continues normally, and the only change is that the caller on hold can suddenly hear the agent briefing the human about them. It is unlikely to show up in a test run and very likely to show up in production.

So we stopped treating the subscription call as the source of truth. We store the intended routing outside the media server, keyed by participant identity rather than track, because track identifiers change whenever a participant republishes. After every event that could have rebuilt session state, a reconcile step reads that intent and reapplies it. The stored intent is the durable record; the media server is brought back into line with it as often as necessary.

‍

The human is a variable. Everything so far assumes the human does what we expect. In practice a transfer can fail in at least seven ways: they never answer, decline, or are busy; the call goes to voicemail; they answer and say nothing, hang up during the briefing, or hang up after the connection is made; or the connect step itself fails.

Rather than handle each individually, we made all seven resolve to cancel. Whatever the human does, the caller ends up back with the agent, in the state the call was in before the transfer began. No combination of human behavior leaves the caller alone in a silent room.

One case produces no signal at all. A human who answers and says nothing generates no event to react to, and the briefing cannot end, because ending it depends on the human saying something the model can act on. Only a deadline catches this, so a timer runs for every briefing, and if it expires the caller is handed back.


Every failure path has the same exit, so no sequence of events leaves the caller without either a human or the agent.

‍

Details that only appear once a human is on the line

A few things surfaced during the build that the design alone would not have predicted.

The order of subscription changes matters. Placing a caller on hold, we unsubscribe from the agent before subscribing to the music, or there is a moment where they hear both. Connecting, we reverse it and subscribe to the human before removing the music, or they hear dead air between losing one and gaining the other.

The briefing agent has to be told who it is talking to. The only live voice it hears is the human's, but the model has spent the whole call treating whoever it hears as the caller. We remap the roles in the briefing prompt, so it knows it is speaking to a colleague who needs a summary, not a patient who needs an answer.

A parked agent must not act on silence. After the handover it remains in the room hearing nothing, so from its own point of view the call has gone quiet. Its idle prompts and inactivity timeout are correct behavior in a normal call. Here they would fire against a conversation the agent cannot hear, and the timeout would end the call for caller and human together. Both now check whether the call has been handed off, and ending the agent's part removes only the agent, not the room.

Why it holds together

The three pieces depend on each other. A briefing without audio isolation is a disclosure, since the caller hears themselves being described. Isolation without a persisted intent disappears on the next deploy and tells nobody it has gone. Persisted intent without a single recovery path protects the audio but strands the caller when the human hangs up. Only with all three is a transfer something we are comfortable putting in front of a patient.

This is also not a novel architecture. Before building, we looked at how several established telephony platforms implement warm transfer: each keeps all parties in a single bridge and controls which audio each receives. The two-room design is the outlier. We did not set out to follow a convention; the constraints pushed us toward the pattern telephony had already settled on.

What we can build on this

The four phases are not four features. They are four rows in one table saying which tracks each participant should hear, and the routing controller's only job is to make the room match the current row. Every other conversation shape is another row in that table, not another architecture.

Some rows are already obvious. A human can join unheard to supervise, stepping in only if they choose. A human who has finished their part can hand the caller back to the agent, the recovery transition used deliberately. All three can be audible at once, so a human can keep the agent on the line to look something up, then dismiss it. None of these change how audio moves; they change who is subscribed to whom.

The briefing can change shape too. Today it is a spoken summary, the contact-center convention. But our agent generates it from our own conversation, so a safety team and a scheduling desk can get different things from the same call, and a human at a console can get the written record alongside the spoken one.

Because the mechanism is our own routing layer rather than a feature we switched on, it is available to every voice flow we build, and behaves the same on the first server and the twentieth.

‍

Anatomy of a Warm Transfer

Shashank Tyagi, Software Engineer - AI
Sep 29, 2026

Heading

Increase in patient engagement

Heading

Reduction in appointment cancellations

Heading

Improvement in treatment adherence

Ather and Helix are our AI customer engagement platform for pharma, supporting healthcare professionals and patients across voice, chat, and web. Any call it handles can reach a point where a human has to take over: a question outside the approved content, something that sounds like an adverse event, or a caller who simply asks for a person. How we hand that call over matters as much as anything the agent said before it.

Most AI voice agents do what telephony calls a cold transfer: the agent connects the caller to a human and steps aside. The human picks up blind, and the caller explains everything a second time. That is a poor experience anywhere, and when the conversation involves clinical detail already disclosed once, it is an unnecessary risk. It is also the only transfer most agent frameworks make easy.

A warm transfer is what every contact center does instead, and it is the standard we hold Ather to. The agent puts the caller on hold, calls the human, privately explains who is calling and why, and only then connects them. The human knows the situation before saying hello. This article is an anatomy of that transfer, and of why each part of it is harder than it sounds.
‍


In a warm transfer, the caller is on hold while the agent and the human talk. Each of the three participants hears something different.
‍

Two ways to build it

There are two ways to give three people a call in which only two of them can hear each other at any moment.

The first is to create a second room. The agent and the human talk there while the caller waits in the original, and when the briefing ends, the human is moved across. It is the design that suggests itself first, because it matches how the conversation feels: the private part happens somewhere else.

The problem is what somewhere else means in a media system. A room lives on one server node, and two rooms can sit on different nodes, so moving a live participant between them can mean migrating an active audio session across machines. That makes the behavior of a transfer depend on how many servers happen to be running.

The second way is to keep everyone in one room and change who can hear whom. Nobody moves. There is no boundary to cross, so nothing migrates, and a transfer behaves identically whether the cluster has one node or twenty.

That is the design we built. The rest of this article is what it takes to make it hold up.

‍

Hearing is a subscription

In our media server, every participant publishes one or more audio tracks and separately subscribes to the tracks they want. Being in a room does not mean hearing everyone in it. A participant hears exactly what they are subscribed to, and the server can change that set at any time, for anyone, without anybody leaving or joining.

That is the property we built on. A warm transfer is four phases, and the only thing that changes between them is who is subscribed to what.

First, caller and agent hear each other: an ordinary call. On transfer we move to hold, unsubscribing the caller from the agent's voice and subscribing them to a hold-music track while we dial the human. When the human answers we enter the briefing: agent and human hear each other, the caller still hears only music. When the human is ready we connect, subscribing caller and human to each other and unsubscribing the agent from both.

No participant is created, removed, or moved at any point. The agent stays in the room throughout, hearing nobody once the handover completes, because it is what takes the caller back if the human drops off.


The agent is unsubscribed during hold so it does not transcribe its own hold music. It stays in the room after connect because it is what returns the caller if the human hangs up.

‍

A fifth transition is not a phase. We call it cancel, it is legal from all four states, and it restores the original subscriptions so the caller is talking to the agent again. Every failure below resolves to it.

What made it hard

Making this hold up in production ran into three constraints that were not obvious at the start.

Hold does not exist. Our media layer has no concept of hold, and carrier hold signals are discarded before they reach us. So hold is something we construct. The agent publishes two tracks from the start of every session: its voice, and a track carrying hold music. Placing the caller on hold means unsubscribing them from the first and subscribing them to the second. We publish the music at session start rather than at transfer time, so the transfer never waits on a publish. The agent itself must also be unsubscribed from everything during hold. If it can hear its own hold music, the speech-to-text pipeline dutifully transcribes it, and the model spends the entire hold reasoning about noise.

Subscription changes do not persist. This shaped the design most. Subscribing to every track is the default, and removing a subscription is a single instruction applied to live session state on one server. It is not stored as a rule. Whenever that state is rebuilt it comes back with the defaults, and it is rebuilt more often than we expected: a deploy does it, so does an autoscaling event, a browser leg dropping and recovering, and the routine reconnection WebRTC performs on its own. None are failures, and two of them we schedule ourselves.

No error is raised and nothing is logged. The call continues normally, and the only change is that the caller on hold can suddenly hear the agent briefing the human about them. It is unlikely to show up in a test run and very likely to show up in production.

So we stopped treating the subscription call as the source of truth. We store the intended routing outside the media server, keyed by participant identity rather than track, because track identifiers change whenever a participant republishes. After every event that could have rebuilt session state, a reconcile step reads that intent and reapplies it. The stored intent is the durable record; the media server is brought back into line with it as often as necessary.

‍

The human is a variable. Everything so far assumes the human does what we expect. In practice a transfer can fail in at least seven ways: they never answer, decline, or are busy; the call goes to voicemail; they answer and say nothing, hang up during the briefing, or hang up after the connection is made; or the connect step itself fails.

Rather than handle each individually, we made all seven resolve to cancel. Whatever the human does, the caller ends up back with the agent, in the state the call was in before the transfer began. No combination of human behavior leaves the caller alone in a silent room.

One case produces no signal at all. A human who answers and says nothing generates no event to react to, and the briefing cannot end, because ending it depends on the human saying something the model can act on. Only a deadline catches this, so a timer runs for every briefing, and if it expires the caller is handed back.


Every failure path has the same exit, so no sequence of events leaves the caller without either a human or the agent.

‍

Details that only appear once a human is on the line

A few things surfaced during the build that the design alone would not have predicted.

The order of subscription changes matters. Placing a caller on hold, we unsubscribe from the agent before subscribing to the music, or there is a moment where they hear both. Connecting, we reverse it and subscribe to the human before removing the music, or they hear dead air between losing one and gaining the other.

The briefing agent has to be told who it is talking to. The only live voice it hears is the human's, but the model has spent the whole call treating whoever it hears as the caller. We remap the roles in the briefing prompt, so it knows it is speaking to a colleague who needs a summary, not a patient who needs an answer.

A parked agent must not act on silence. After the handover it remains in the room hearing nothing, so from its own point of view the call has gone quiet. Its idle prompts and inactivity timeout are correct behavior in a normal call. Here they would fire against a conversation the agent cannot hear, and the timeout would end the call for caller and human together. Both now check whether the call has been handed off, and ending the agent's part removes only the agent, not the room.

Why it holds together

The three pieces depend on each other. A briefing without audio isolation is a disclosure, since the caller hears themselves being described. Isolation without a persisted intent disappears on the next deploy and tells nobody it has gone. Persisted intent without a single recovery path protects the audio but strands the caller when the human hangs up. Only with all three is a transfer something we are comfortable putting in front of a patient.

This is also not a novel architecture. Before building, we looked at how several established telephony platforms implement warm transfer: each keeps all parties in a single bridge and controls which audio each receives. The two-room design is the outlier. We did not set out to follow a convention; the constraints pushed us toward the pattern telephony had already settled on.

What we can build on this

The four phases are not four features. They are four rows in one table saying which tracks each participant should hear, and the routing controller's only job is to make the room match the current row. Every other conversation shape is another row in that table, not another architecture.

Some rows are already obvious. A human can join unheard to supervise, stepping in only if they choose. A human who has finished their part can hand the caller back to the agent, the recovery transition used deliberately. All three can be audible at once, so a human can keep the agent on the line to look something up, then dismiss it. None of these change how audio moves; they change who is subscribed to whom.

The briefing can change shape too. Today it is a spoken summary, the contact-center convention. But our agent generates it from our own conversation, so a safety team and a scheduling desk can get different things from the same call, and a human at a console can get the written record alongside the spoken one.

Because the mechanism is our own routing layer rather than a feature we switched on, it is available to every voice flow we build, and behaves the same on the first server and the twentieth.

‍

Download icon with a downward arrow pointing to a horizontal line inside a blue circular button.

Thank you! The case study will be emailed to you shortly. Please check your spam and junk folders

Oops! Something went wrong while submitting the form.
Download icon with a downward arrow pointing to a horizontal line inside a blue circular button.

Thank you! The case study will be emailed to you shortly. Please check your spam and junk folders

Oops! Something went wrong while submitting the form.