67 karma · joined February 6, 2017
openrouter is the openrouter for audio models.
the conflation is "audio models" vs "voice ai", and the mental model that untangles it: think batch requests. text in, audio out (streamed, even). audio in, transcript out.
three questions inside the word "router":
1 - what gets picked: a model/voice, a provider of the same model, or the whole stack (stt + llm + tts) per call
2 - where it lives: an external http gateway, the agent platform (vapi/retell/livekit configs), or inside the live session
3 - when: session start, or mid-call
a voice agent is not a batch request. it's a live duplex session: turn-taking, barge-in, telephony legs, session state. the latency physics diverge too: a middleman hop in the media path is paid once by a batch request and on every conversational turn of a live call, so the media path wants a direct connection to the provider. a gateway that terminates at http can route the requests inside a call; it can't route the call.
for some languages CER is more relevant than WER. Thai and Mandarin have no word boundaries, so we score them by character, and Japanese gets a reading-based CER.
We often find that models that wins on clean speech are often not the one that wins on your terms, so test with your own vocabulary rather than a headline number.
Wider picture: we rebuilt the speech-to-speech test around a hard 20-turn concierge call, and today only the closed models finish it cleanly. Board: https://benchmarks.speko.ai/s2s
shared some write-up: https://benchmarks.speko.ai/blog/can-s2s-replace-the-cascade
And jakswa's Gemma datapoint matches what we measure, the fastest rows on our LLM board are the small models like Gemma 4 is performing incredibly well for voice agents
production phone agents are a different shape today: the call terminates server-side, three models plus turn-taking under one latency budget, and per-language quality still swings a lot from our tests
You are right about human input for naturalness, that one we did not automate away with yet. We run blind A/B listening rounds with native speakers.
On fast dumb models answering while a smarter one takes over: we are experimenting with exactly that split - a small fast model holds the conversation while a larger one works behind it. Today it runs as two pinned routes, not one packaged API. Most turns in a phone call do not need a frontier model, and the fastest models on our LLM board are all small, so this is where routing earns its keep. We publish benchmarks on LLMs here: https://benchmarks.speko.ai/llm
On "promptfoo of voice models": that is closer to how it started. At my last company we ran these evals manually, we would even hire native-speaking raters, benchmark, switch if it wins. The evals are the value, agreed. The routing is what makes them actionable: teams told us swapping always looked like an R&D project, so scores alone did not change what ran in production.
On prompt-based voice gen and reference-audio cloning: agreed, that is what we see too. It makes continuous measurement more important: the same style prompt behaves differently per language and per content type, so we rank the voices themselves, tagged by use case: https://benchmarks.speko.ai/tts-voices
Second difference is where it runs. Our gateway is open source and runs in your own container, including with self-hosted livekit/pipecat. You get a temporary token before the session starts, and then your orchestration connects directly to the provider.
Vapi is a managed platform: you use their infra to use the voice AI stack. In our case you can have your own infra and switch between models, so you are not locked into a vendor. A lot of teams we talk to build their own infra as they mature, and that is where the router comes handy.