These Execs Think Voice AI Still Hasn't Had Its ChatGPT Moment —

The theory that voice will be the next major computing interface has picked up strong momentum. Investors have poured billions of dollars into voice AI startups spanning model makers, enterprise customer service providers, meeting note-takers, and AI-powered dictation tools. Every week brings a new model or product release claiming to sound human — and to converse like one.


In practice, however, that promise remains unfulfilled. Shawn Wen, CTO of enterprise voice AI platform PolyAI, argues that despite the arrival of full-duplex models — which can listen and speak simultaneously — voice AI has yet to reach its "ChatGPT moment."


"We have reached the milestone of developing full-duplex models. The next challenge is to make reasoning very fast, so that the models can fetch answers quickly and the conversation feels natural," Wen said on stage at the HumanX conference last month.


He added that AI agents in customer service should not sound robotic, and should give callers enough confidence that their problems can actually be solved.


"I think the next stage will be slightly different because once the voice is good enough, and the customer is willing to engage with them for the first two or three turns, they start to build confidence," he said. "Over time, they will feel like: I probably don't have to talk to a human if the agent can solve my problem."


Understanding and Transparency Remain Open Problems


While voice AI models have improved, assistants frequently fail to understand users, and meeting notetakers often surface the wrong transcript or summary. Wen believes ASR (Automatic Speech Recognition) models routinely miss important keywords, which undermines their ability to capture full context.


Alex Gay, CMO of meeting notetaker Otter, agreed. He said speaker identification, intent capture, and combining that with organizational knowledge are key steps toward meaningful automation. The company is also developing digital twins that could represent people in meetings. For that technology to work, he said, the output voice must convey the same emotive expressions as talking to a human in a meeting.


"If you think about the meetings you're in right now, the best conversations you have are where you can debate and have strategic discussions, and where you feel like there's a relationship that underpins it," Gay said. "If you aren't able to have that with an avatar, then it's just a Q&A chatbot."


Why Transcription Accuracy Matters


Gay emphasized that transcription was never Otter's endpoint — it was simply the foundation for driving productivity gains. But if the original transcription lacks the accuracy users need, every downstream action built on it becomes flawed.


"The minute that starts to take action, that is wrong. You lose trust in the platform," he said. "It is critical for us to continue to improve that ASR model because all of the downstream impacts are significant." He also identified language as a key area where voice models still need improvement.


The Trust and Disclosure Question


The rise of new voice tools raises a parallel question of transparency: tools should clearly disclose to customers when they are being recorded or speaking with an AI. Otter says it aims to instill trust in meeting participants — even when its bot is not present, the company wants to explore methods such as notifying everyone in the chat that the meeting is being recorded. PolyAI's Wen likewise stressed the importance of making it clear to people on enterprise calls that they are talking to an AI.


As of 2026, the voice AI sector continues to attract heavy investment and rapid model iteration, yet the gap between demo-quality conversations and reliable, trustworthy enterprise deployment persists. Executives across the space agree that speed of reasoning, ASR accuracy, emotional expressiveness, and transparent disclosure will determine whether voice AI finally gets its breakthrough moment.

via TechCrunch AI

Related