A frustrated customer decides how a call is going to go in the first few seconds, and the main indicator they use is the sound of the voice that greets them.
That’s the part most automation gets backward. CX teams pour effort into what the system says, and treat how it sounds as decoration. In a tense moment, the sound is the message.
There is science behind this. Researchers studying voice-to-voice communication found that people automatically reproduce the emotional tone of whoever they’re talking to. The effect is fast, unconscious, and hard to suppress. Put plainly: calm is contagious. A steady, warm voice gives an upset caller something to settle into. A flat, robotic one gives them a reason to escalate, or to ask for a human.
So a voice is the first chance to contain a call, and the first chance to de-escalate it.
Trusting the voice
Customers extend trust to a voice before they’ve heard a single fact. We do it with each other all the time. A measured tone reads as competence. Warmth reads as care. Good pacing reads as “this is under control.”
The catch is that a voice only earns that trust if it holds up under pressure. One mispronounced name, one burst of dead air, one stretch of mechanical delivery, and the spell breaks. The caller stops listening to the words and starts listening for the exit. In a tense call, that’s the moment you lose them to an agent, or lose them entirely.
So the bar for an automated voice in a key moment is high. It has to sound human, stay natural under load, and shift its delivery to match the weight of the conversation.
A voice that reads the room
This is what Lexis is built for. Lexis is Omilia’s generative text-to-speech model: expressive, real-time speech that runs natively inside OCP and never leaves the platform.
A few things make it fit for those tense conversations:
- It adapts to the context. Prosody shifts with the content of the conversation, not just the words on the page. A sensitive account call doesn’t sound like a quick order confirmation.
- It never leaves dead air. Lexis streams audio as it generates, with first audio in under 45 milliseconds and barge-in support, so turn-taking feels like a real exchange instead of a recording.
- It says every word right. Custom lexicons and SSML control pronunciation down to the term, so names, products, come out correct the first time. Output transcribes back at under 3% word error rate.
- It can sound like you. A pre-tuned persona such as the Trusted Advisor brings a measured, reassuring tone to high-stakes calls out of the box, or you can clone a custom brand voice almost immediately from a short clip.
The result is a voice customers accept. And a voice they accept keeps them in the conversation instead of bailing to an agent, which protects the single biggest line of contact-center cost.
The takeaway
Generating trust isn’t only a skill you train into agents. It’s a property you can design into the voice itself. Get the sound right and the calm spreads before the first answer even lands. The customer relaxes because the voice already told them they’re in good hands.
That’s the quiet advantage of treating voice as a brand asset rather than an afterthought. The right voice doesn’t just deliver the resolution. It sets the emotional tone that makes resolution possible.
Want to hear Lexis in your own brand voice? Let’s talk!
About the Author
Conall Murtagh, Product Marketing Manager
Conall is a trusted product marketing professional with over 20 years of experience in B2B technology. With a cross-industry track record of shaping how complex products are understood in the market, he has a genuine interest in how technology can help the world communicate better. Based outside Madrid, he is a keen amateur chef and enjoys running, cycling and spending time with his wonderful family.


