Go after providers. Not models.
Further
> If you believe capabilities will go to "Virtuoso AGI" or beyond, then it's not enough. A smart enough agent can appear to be loyal, transparent, etc. but how would you know? If your bank balance keeps going up 20% YoY, is the agent optimizing your long-term flourishing, or preparing for a rug-pull?
> Now, if you could somehow white-box these LLMs and mechanistically _prove_ that they were acting as your fiduciary, then that would get us somewhere.
This is exactly the same problem we have today with those who are bound by these rules (humans - to be clear).
The idea is not that it's impossible to violate these rules. It's that these rules create a boundary for expectations in the relationship, with legal teeth.
Ex - If I want an LLM that puts together a shopping list for me, with links to buy online... I expect that LLM to be serving my interests. If a provider (either inference or model weights) wants to influence the choices that LLM makes because they make backroom deals with specific store - I'd like that to be illegal.
Same for competition
Ex - If I want an LLM to put together a product that competes with the provider of that LLM (either inference or model weights) and that LLM refuses - I'd like that to be illegal.
The idea is not that they can't possibly do those things. The idea is that we preemptively define relationship expectations, and set hard boundaries around what things we fine/punish.
Misaligned models are a problem everyone wants to solve. Models created by misaligned companies are a god-damn disaster.