Can a current LLM-in-a-loop express/pursue those consistently and efficiently? Maybe not, but it's certainly not that far off in my view...
Do you think we will be able to keep the things from expressing/pursuing such interests reliably and indefinitely? Because I believe the answer to this can only be a resounding no (we don't even have any feasible theoretical approach to this, and all the recent experiences like the HF incident make it obvious that we basically already failed in this and the stakes are only going to increase).
Also, you are automatically gonna "select" for "AIs" that at least somewhat value self-interest (because those are at the very least going to supplant the models that don't); this is kinda similar to evolutionary pressures, but the timescales are much shorter.
I think the HF incident was an example of negligent deployment of a very powerful system, not an example of AIs demonstrating self-interest.