Founder Spotlight
Davut Akça
Founder
Voxis (voxislive.com)
Watching a foreign film, playing a game, sitting on a call with someone who doesn't speak your language, and there it is again: that bar of text at the bottom of the screen. You're reading instead of watching, translating instead of listening. Davut Akça kept hitting that wall, and he didn't want better captions. He wanted the audio itself to switch languages, live.
That frustration became Voxis.
We invited Davut to answer our usual three questions.
How did you end up in voice?
I didn't start out building for voice. I started out frustrated by subtitles. Watching foreign films, playing games, sitting in calls with people who didn't share my language: I kept hitting the same wall. Text on screen breaks the moment. You're reading instead of watching, translating instead of listening. I wanted the actual audio to change languages in real time, not just have a caption bar underneath. That took me into real-time speech pipelines. I worked on capturing whatever's playing on a Windows machine, streaming it to a simultaneous-interpretation model, and playing back translated speech fast enough that it still feels live. Voxis is the result. No subtitles, just the room speaking your language back to you a few seconds behind. It turned out "just add captions" was the easy, wrong answer. The harder problem, real-dubbed audio live, was the one worth solving.
What's your struggle or moment of joy with your voice?
The struggle is that I'm building on borrowed infrastructure. Every engine I route through, whether it's Gemini Live, Qwen, or OpenAI, is someone else's hosted preview API. So a chunk of my job is absorbing their instability, like a spent balance, a rate limit, or a silent disconnect, and building failover so the user never hears the seam. The joy was smaller and more human: the first time a real stranger paid for it. Not a friend testing a demo, but a person in Romania who found Voxis, tried it, and handed over money for minutes because it did the one thing it promised, making a video sound like it was speaking their language. That's the moment solo-building stops being abstract and becomes "Someone needed this."
Where do you think voice is going?
Away from text as the default translation layer. Subtitles and transcripts were always a compromise, a workaround because translation was too slow, too costly, or too robotic just to be spoken. Simultaneous interpretation models are closing that gap fast enough that translated audio, live, in your own voice or theirs, stops being a novelty and starts becoming infrastructure. It will be under a game, a meeting, a stream, and eventually a phone call. The next fight isn't about whether voice can be translated in real time. It's about who owns the last few seconds of latency and how natural the result sounds. I think the products that win won't market themselves as translators at all. They'll just be the thing that made language stop being a wall.