You point is valid for textual LLM, with large CoT, not Omni which will respond quick with a voice. In this case, token price is a good enough proxy.
What matter the most and isn't told by token price is the latency. You expect a voice LLM to respond very quick. If it takes 5s to response to a simple "Hello, what the weather today?", them not much people will use it.