But provider switching is built in some of these - and the folks behind envoy built: https://github.com/katanemo/archgw - developers can use an OpenAI client to call any model, offers preference-aligned intelligent routing to LLMs based on usage scenarios that developers can define, and acts as an edge proxy too.
You're right that archgw handles routing at the infrastructure level, which is perfect for centralized control. any-llm simply gives you the option to handle routing in your application code when that makes sense (For example, premium users get Opus-4). We leave the architectural choice to you, whether that's adding a proxy, keeping routing in your app, or using both, or just using any-llm directly.
I have to do this already with practically all software I write, so the comexity is already baked in. Sure, if you don't already have a database or cache, maybe a proxy is simpler, but otherwise it's just extra infrastructure you need to manage.
> The proxy server is the right design decision if you are truly trying to build something production worthy and you want it to scale.
I've been doing stuff like the above (not for LLMs but similar use cases) for years "at scale" without issues. But in any case, you need to store state the moment you scale beyond a single proxy server anyway. Plus, most products never achieve a scale where this discussion matters.