We took a different path at Komilion: classify the prompt upfront (regex fast path + lightweight LLM classifier) and route to the cheapest model that benchmarks well for that query type. Simpler, but works without model-specific training.
The 70% cost reduction figure matches what we see in practice. The insight that most queries are "easy" is the key — once you stop sending FAQ-level questions to frontier models, the savings are dramatic.
Curious if you have looked at combining both approaches — upfront classification for obvious cases, then internal state probes for the ambiguous ones in the middle.