Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I personally tested all the STT models for my real-time translator (https://fliptalk.ai). From language detection and accuracy in a noisy environment to the most important point: latency.

At the moment, Soniox STT v5 is definitely the best, and I'm impressed by its performance. It's good that Google released Gemini-3.5-Transcribe, and it beats every other model on accuracy, but it definitely needs a bit more work on latency, which is the most important factor for STT apps.



Soniox website has a live comparison demo: https://soniox.com/compare-stt

For me, Gemini-3.5-Transcribe actually has slightly lower latency. Kudos to Soniox for both paying their competitor and letting them win. But yes, Soniox is much cheaper.


Interesting, I uploaded a voice recording from a meeting I had recorded with a relatively cheap microphone.

Soniox came out really good. OpenAI started getting some things very wrong and even introduced some German. Google did okay but cut off the start by several seconds.

What's Soniox doing (left most) that's making it so good ? It was also the only one that could distinguish between the speakers.


Oh, thanks for pointing me to Soniox. It is really good. Would also pick up the words with different languages, identify and output in the right language. Looks interesting. It was much faster too, but that I cannot say much since it was on their own website.


  > latency, which is the most important factor for STT apps.
Perhaps latency is more important than accuracy for a real time translation app (I actually disagree with this - imagine e.g. the hilarity when requesting "a new display" being translated as "a nudist play"), but certainly not for all applications. My pet app transcribes personal voice notes to self, it could run all night.


I agree. For my use case, I chose to prioritize latency over accuracy, but it's always difficult to find the right balance between the two. There is no easy answer.


Depends on your use case. If you're transcribing meeting notes, latency is a non-issue.


Realtime + Voice AI usecases is where latency is most important. I use Handy on my desktop and i can tolerate a latency of a few seconds every now and then. Your P99 should on TTFB should be really low to compete for voice ai realtime


Thank you for this!

I was using Cartesia, while their TTS is amazing their STT pricing has kind of irked me.

Interested in know how good soniox latency and EUD is on STT compared to Cartesia. Cartesia's is really in real world conversations




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: