The Global Voice AI Gathering report is out 🥳
    Industry Conversation

    Founder Spotlight

    Thomas Kluiters

    Co-founder and Head of AI

    Reson8

    LinkedIn
    Thomas Kluiters

    Meet Thomas Kluiters, co-founder and Head of AI at Reson8. Three years ago his girlfriend, a doctor, was losing hours to keyboard notes and asked him to fix it. He started with a medical transcription company, Juvoly, then built Reson8's own speech models when off-the-shelf speech-to-text could not handle Dutch medical consultations. Reson8 now builds speech-to-text for European languages. In our fifth BotCast, Thomas argues European founders should be more ambitious, backs speech-to-text plus text-to-speech over speech-to-speech, and predicts voice becomes the default interface within five years.

    Source: BotCast Ep 5, Voice AI Space.


    Who are you?

    I'm Thomas Kluiters. I am the co-founder and Head of AI at Reson8. We build customizable speech models, primarily focused on speech-to-text. Currently we're one of the leading speech-to-text providers for the European languages, both for batch processing and turn-based speech models.

    How did you end up using voice AI?

    Three years ago, my girlfriend, who was a doctor, complained a lot about having to type a lot on her keyboard during her consultations with patients. She just asked me, "Why don't you solve this?" So I founded a company called Juvoly, and we used existing speech-to-text models to transcribe the conversations between patients and doctors. But we realized that the accuracy of a lot of speech-to-text providers just wasn't good enough for Dutch medical consultations. So we started building our own speech models, and I realized that the way we're building these speech models is really working, and the way we can customize them is pretty scalable. So why don't we just do this for the rest of the world? That's why I started Reson8, solving exactly this.

    A moment of joy and pain using voice tech.

    I think both the joy and the pain have a common denominator: getting models to a really good point. I would spend an excruciating amount of time on getting our models better, only for them to do worse in benchmarks. And at the same time, the most wonderful thing there is is sitting on top of benchmarks, both internally and externally, and your customers being really excited about your new model coming out, where you're actually solving specific problems, maybe rare names or medication names. And whenever you're working weeks on end trying to find a strategy that works and it's all been for nothing, spending so much money on compute, it is the worst feeling in the world.

    What lesson would you share with other builders?

    Definitely being a bit more ambitious. Whenever I speak to first-time founders, generally they're just too afraid or too scared to think bigger. Especially in Europe, we tend to keep things to ourselves and try to minimize risk. I think we should be risk takers. We should be ambitious. My primary advice here is: think about what you want to achieve, multiply it by maybe a thousand, and really aim for and reach for the stars. Don't just think about the country you're working in, think about the continent, think about the world, not just where you are right now. Truly try to reach for the stars.

    Where do you see the industry in 12 months and in five years?

    I think the next 12 months are going to be very important to see just how well speech-to-speech models are going to be adopted in production. There's a large group of people who think speech-to-speech is going to be the winning architecture of choice, whereas there's a lot of people as well, me included, who think that speech-to-text and text-to-speech are going to be the better way forward. So it's very curious to see what's going to happen in the next 12 months.

    For five years from now, I think voice is going to be the default interface to whatever we have. We're going to use our keyboard, we're going to use our touch screen, and we're going to be using a lot of voice. I think it's going to be more normal to be speaking with our devices. So I think voice AI will be everywhere within five years, maybe even three years, who knows.

    More Industry Conversations