I was in the AI speech assessment field for ~ 9 years so let me share a few things re: AI speaking practice.
Even before this LLM era, many models had already been working quite well in identifying and correcting your speaking - pronunciation, intonation (this includes chinese as well), fluency, speed, etc. (My team was offering such service in several languages)
The only downside was the UX, which heavily relied on old chatbot (i.e. dumb) to only focus on repetition. There was no actual interaction or proper adaptive feedback.
Now I can only imagine with gen-AI and agentic AI taking over, this UX gets massively upgraded. This is shown a plethora of new speaking apps nowadays, though of which most are still pretty much a sham.
The difficulty of tackling the UX in this is that natural spoken conversation is fundamentally unpredictable — and most systems are still engineered around the assumption that it isn’t.
The feedback loop kills the flow. Real conversation doesn’t pause for correction, but most apps still interrupt mid-sentence or dump a scorecard at the end that you’ve already emotionally disconnected from. A good teacher knows when to jump in. Most AI still doesn’t.
There’s also a difference between assessing and actually listening.
Assessment is evaluation against a rubric. Listening means tracking intent, context, and register — and knowing what to flag based on what matters right now. Current systems largely can’t make that distinction.
And spontaneous input still breaks the pipeline. Most apps are secretly scripted — happiest when you say something predictable. The moment a learner goes off-piste, quality degrades fast. True conversational AI for speaking practice needs ASR, NLU, and pedagogical logic all firing simultaneously, not sequentially.
But the biggest gap? There’s no model of the learner over time. Pronunciation errors aren’t random — they’re systematic and personal. Without tracking your specific patterns across sessions, you’re just doing expensive karaoke.
What we actually need is an AI that can hold a real conversation, catch in the moment that you consistently drop certain sounds under cognitive load, and surface that at the right time without making you feel like you’re being graded.
That’s not a model problem anymore. That’s a product and interaction design problem — and most teams building these apps don’t have anyone who deeply understands both sides.