This isn't important. What is important is that we can't prevent LLMs from doing things they're not supposed to. Which means connecting their output to anything important is a very bad idea for now.
This isn't important. What is important is that we can't prevent LLMs from doing things they're not supposed to. Which means connecting their output to anything important is a very bad idea for now.
1. You don't train/tell the model anything you're not willing to share with the user.
2. The model gives suggestions to a user, and--even if that suggestion is accepted as-is--you treat it as potentially untrustworthy data supplied by that user.
This may be true for images, but at least for text, treating an LLM as an untrusted person in your threat models will at least allow you to apply defense in depth to the downstream systems consuming from the LLMs
Tldr: engineer other systems to treat LLMs as a potentially bad actor by default.