But there are much more informed ML people out there than me, so I assume this and similar techniques have already been thought of.
But there are much more informed ML people out there than me, so I assume this and similar techniques have already been thought of.
The LLM is just generating orders of magnitude more user state as a result of a prompt. The actions the LLM is permitted to take must still be gated on the authorizations allowed for that user. This means that data unavailable to the user must not be in the training set of the LLM acting as an interface layer.
The "we have to learn new lessons" only applies when you lump all data together and hope the LLM doesn't spit up someone else's data from a probabilistic elbow jog. Hope is not a strategy.
But then, you have two independent mechanisms that can get out of sync, a classic source of issues in infosec - except both are also more or less inscrutable and fail in unexpected ways.