HNHacker News
TopNewBestAskShowJobs

abdik

67 karma · joined February 6, 2017

submissionscomments
abdik··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
not really counter-positioned, they are one of the providers we route to. elevenlabs sells models - voices, stt, and their agent platform on top of them. we don't sell any models: we measure all of them (elevenlabs included, their scribe is near the top of our stt board) and route each request to whatever wins for your language and constraints. so for them the best outcome is that you use their models; for us the best outcome is that you use the right one. sometimes that is the same thing.
abdik··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
i don't have a thesis on voice as a form factor for search. the demand we serve already exists: businesses answer phones. receptionists, outbound campaigns, clinic front desks and these calls happen at scale today, and the teams running them are the ones picking speech models. whether voice wins new interfaces is a separate bet
abdik··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
thanks! yes, some of them do, cuz speed is a per-provider capability, not universal. And, you can see it in the gateway code (minimax, hume, xai tts adapters all handle a speed param). that unevenness is actually a routing constraint by itself: "voices that support rate control" narrows the candidate list the same way language or latency does. and agree on the voice modes, the gap between the demo and a dependable agent is exactly why we started this. what are you building with it?
abdik··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
People who built or building voice agents immediately get this.

openrouter is the openrouter for audio models.

the conflation is "audio models" vs "voice ai", and the mental model that untangles it: think batch requests. text in, audio out (streamed, even). audio in, transcript out.

three questions inside the word "router":

1 - what gets picked: a model/voice, a provider of the same model, or the whole stack (stt + llm + tts) per call

2 - where it lives: an external http gateway, the agent platform (vapi/retell/livekit configs), or inside the live session

3 - when: session start, or mid-call

a voice agent is not a batch request. it's a live duplex session: turn-taking, barge-in, telephony legs, session state. the latency physics diverge too: a middleman hop in the media path is paid once by a batch request and on every conversational turn of a live call, so the media path wants a direct connection to the provider. a gateway that terminates at http can route the requests inside a call; it can't route the call.

abdik··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
yes, it does support, when you are creating an api key, you can point out narration or transcription use case, then you will be able to see. let me know how it goes or what use cases are thinking of for non-realtime models?
abdik··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
looks cool, checking it out!
abdik··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
exactly, we see the same thing, around 95% cases are still cascaded, even tho STS has been improving a lot
abdik··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
Synthetic-data fine-tuning is the other credible answer to domain vocabulary. Curious whether you re-benchmark the fine-tune when new base models ship?
abdik··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
What purrcat259 said, and I think that the page should explain it, we will add a tooltip.

for some languages CER is more relevant than WER. Thai and Mandarin have no word boundaries, so we score them by character, and Japanese gets a reading-based CER.

abdik··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
agree with omneity here. Whisper's initial-prompt trick is exactly that, and several hosted vendors have equivalents (custom vocabulary / keyword prompting). Domain vocabulary is where STT models separate the most in our runs. for example, on medical terms the field spreads from about 8% to 19% WER across models: https://benchmarks.speko.ai/blog/what-a-voice-agent-hears.

We often find that models that wins on clean speech are often not the one that wins on your terms, so test with your own vocabulary rather than a headline number.

abdik··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
Actually, we measured exactly this recently. The strongest open-weights speech-to-speech model we have run is NVIDIA's NemotronLabs VoiceChat 11B - no provider serves it, so we hosted it ourselves and ran the same scripted call every model on our board gets. Remarkably stable, 40+ sessions with zero errors - but by turn thirteen it was answering nine tries in ten without saying a word. Stable engine, but degrades on long calls.

Wider picture: we rebuilt the speech-to-speech test around a hard 20-turn concierge call, and today only the closed models finish it cleanly. Board: https://benchmarks.speko.ai/s2s

shared some write-up: https://benchmarks.speko.ai/blog/can-s2s-replace-the-cascade

And jakswa's Gemma datapoint matches what we measure, the fastest rows on our LLM board are the small models like Gemma 4 is performing incredibly well for voice agents

abdik··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
cool app, and agreed that on-device keeps eating the single-user cases, we are seeing dictation and translation are exactly where local models shine. We benchmark the open models on the same boards as the hosted ones, but there are still a few: https://benchmarks.speko.ai/open.

production phone agents are a different shape today: the call terminates server-side, three models plus turn-taking under one latency budget, and per-language quality still swings a lot from our tests

abdik··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
Turn-taking specifically does not need a listening panel, but it is measured mechanically. 200+ real human clips, and we score end-vs-wait decisions: did the model decide the caller finished speaking, or just paused mid-thought. The best detector gets 94.0% of those right; a plain VAD silence timer gets 46.9%. Results are published here: https://benchmarks.speko.ai/turntaking

You are right about human input for naturalness, that one we did not automate away with yet. We run blind A/B listening rounds with native speakers.

abdik··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
Yes, on the hosted side (agents platform): full sessions come with VAD and turn-taking handled - we set them up and tune them for your use case, so that is the closest thing to conversation in a box. If you run your own orchestration, the gateway is just the routing layer and turn-taking stays in your framework - in our own stack we run Pipecat's Smart Turn in-process and tune the commit threshold on real calls. We also share our benchmarks here: https://benchmarks.speko.ai/turntaking

On fast dumb models answering while a smarter one takes over: we are experimenting with exactly that split - a small fast model holds the conversation while a larger one works behind it. Today it runs as two pinned routes, not one packaged API. Most turns in a phone call do not need a frontier model, and the fastest models on our LLM board are all small, so this is where routing earns its keep. We publish benchmarks on LLMs here: https://benchmarks.speko.ai/llm

abdik··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
there is a good progress on on-devise models, but not ready for production yet to fit in devices. But as soon as there is are some good results, we are going to benchmark them and put in https://benchmarks.speko.ai/
abdik··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
thanks! actually, we have the filipino already, can you check out and share your feedback?
abdik··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
Fair pushback. On end to end: we measure those too, same methodology: https://benchmarks.speko.ai/s2s. If the single models win, we route to them the same way, so we do not care which architecture (s2s or cascaded) wins. For now, what we see in production so far is that most teams still want to control each piece: swap the STT for medical vocabulary, keep the LLM, keep the voice.

On "promptfoo of voice models": that is closer to how it started. At my last company we ran these evals manually, we would even hire native-speaking raters, benchmark, switch if it wins. The evals are the value, agreed. The routing is what makes them actionable: teams told us swapping always looked like an R&D project, so scores alone did not change what ran in production.

On prompt-based voice gen and reference-audio cloning: agreed, that is what we see too. It makes continuous measurement more important: the same style prompt behaves differently per language and per content type, so we rank the voices themselves, tagged by use case: https://benchmarks.speko.ai/tts-voices

abdik··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
The main difference from gateway is we help with picking the right voice stack, which seems to be a big problem for users: we benchmark the models continuously and route based on those measurements for your language and constraints, and the boards are public at https://benchmarks.speko.ai/

Second difference is where it runs. Our gateway is open source and runs in your own container, including with self-hosted livekit/pipecat. You get a temporary token before the session starts, and then your orchestration connects directly to the provider.

Vapi is a managed platform: you use their infra to use the voice AI stack. In our case you can have your own infra and switch between models, so you are not locked into a vendor. A lot of teams we talk to build their own infra as they mature, and that is where the router comes handy.

abdik··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
Thank you!
abdik··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
Yes, i added it. that's the right link.
abdik··on Show HN: Cleanvoice – Automated Podcast Editing
The logo is similar to ours https://www.lovo.ai/
abdik··on Ask HN: What's your ideal city in a 100% remote world?
Check out https://nomadlist.com/
abdik··on Ask HN: Who is hiring? (November 2018)
YY9mtbXH2xpcyTvPeFVqh6Guo7ISe47HGtfcrqt11nsDy9hJDAQ50er6KpCCmILl1ztJ6xdC/7vPSTyiTEQfUYP05ZMSsv7e5IAa3xO0U4VZr/9rTEEub/a0epxZujTJSlazNsdYlFRMrUDekVsqIxq6bjlf3v5lQdIZxQMGOscp4cbfWgMZuf0yFiZb0t7S2W4I5UsHyxmR/dZg7eHx2An+CTeirvIFE/KQHjAZLhqnkl0GfN3371SWHw7VqFZ4Ov/C/RqGPOfUme3veggve04JxGigd5ObeWZV9vLgknEO3ftdxelOd9VUqLlwLvpzV3gURY310UPzpiW/6mKG9A==