144 karma · joined January 11, 2022
Still, I agree that routing in isolation is not thaaat useful in many LLM domains. I think the usefulness will increase when applying to multi-step agentic systems, and when combining with other optimizations such as end-to-end learning of the intermediate prompts (DSPy etc.)
Thanks again for diving deep, super helpful!
On our roadmap, we plan to support:
- an API which returns the neural scores directly, enabling model selection and model-specific prompts to all be handled on the client side
- automatic learning of intermediate prompts for agentic multi-step systems, taking a similar view as DSPy, where all intermediate LLM calls and prompts are treated as latent variables in an optimizable end-to-end agentic system.
With these additions, the subtleties of the model + prompt relationship would be better respected.
I also believe that LLMs will become more robust to prompt subtleties over time. Also, some tasks are less sensitive to these minor subtleties you refer to.
For example, if you have a sales call agent, you might want to optimize UX for easy dialgoue prompts (so the person on the other end isn't left waiting), and take longer thinking about harder prompts requiring the full context of the call.
This is just an example, but my point is that not all LLM applications are the same. Some might be super sensitive to prompt subtleties, others might not be.
Thoughts?
In contrast, our router sits at a higher level of the stack, sending prompts to different models and providers based on quality on the prompt distribution, speed and cost. Happy to clarify further if helpful!
Best email would be: daniel.lenton@unify.ai
Cheers!
For those who don't want to always route, another core benefit of our platform is simple custom benchmarking on your task across all existing providers: https://youtu.be/PO4r6ek8U6M
If you then just want to use that provider rather than a router config, then that's fine!