Speak CEO: The 99.9% accuracy leap that built an AI tutor
Curated by the Inblix editorial team
For most of the last decade, language apps treated speaking like an afterthought. Connor Zwick saw that gap in 2015 when a scrappy side project—training a model on scraped YouTube data—accidentally beat the state of the art in accent detection. That moment, he says, made it obvious that deep learning could “completely smash state of the art” if you fed it enough data. Speak was born from that insight, zeroing in on a painful reality: existing speech recognition models were laughably bad at understanding anyone with an accent.
The early years were a grind of building robust speaking experiences that actually worked. Zwick maintains that shipping AI products requires deep technical intuition, not just market vision. You have to know the difference between 90% and 99.9% accuracy—a distinction he calls “a completely different ballgame”—and be able to predict when that reliability curve will bend. That instinct lets Speak place bets on features that are cost-prohibitive today but will be cheap tomorrow, or design around model weaknesses they’re confident will vanish.
Now the frontier has shifted. Zwick points to OpenAI’s real-time API and audio multimodality as the breakthrough that makes a “superhuman AI speaking tutor” plausible. It’s not just about transcribing words. Instant tone detection, pronunciation analysis, and open-ended feedback that mirrors a learner’s emotional register—that’s the holy grail. But he’s equally bullish on something less flashy: reasoning. He argues that what separates elite human teachers isn’t just empathy, but the ability to design curricula and adjust learning plans. Agentic reasoning, he predicts, will ultimately close that gap.
Still, Zwick is careful to frame AI as a scaling mechanism, not a replacement. Billions of people are trying to learn English, and the world simply doesn’t have enough skilled teachers. The mission isn’t to sideline humans but to make conversational practice accessible to anyone with a phone. It’s a pragmatic bet on supply and demand—and on a team culture he describes as relentlessly curious, the kind that won’t let a tool like ChatGPT settle for a “blah” response.
💡 Key Takeaways
- Speak’s origin traces back to an accidental breakthrough in accent detection using scraped YouTube data, proving deep learning’s power with enough scale.
- Connor Zwick argues that AI product leaders need hands-on technical intuition to predict when model accuracy will jump from unreliable to transformative.
- OpenAI’s real-time audio API enables feedback that goes beyond transcription—analyzing tone, pronunciation, and intent—which Zwick calls the holy grail for AI tutors.
- The next major unlock for language learning isn’t voice, but reasoning: AI that can design curricula and adapt teaching strategies like elite human instructors.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.