If there isn't a way to secure the behaviour of AI models against reliable exploits then the utility of the models is dramatically limited.
It's like AI diagnosis, we aren't going to run it full stop automated without safeguards on top or manual review for a long time.
* Learn how to make bombs (as mentioned in the article)
* Get away with committing crimes?
Are moderating these topics related to morality?I can think of many things an LLM could do that would be far more harmful than any of this.
The US government said that info was born secret and sued: https://en.wikipedia.org/wiki/United_States_v._Progressive,_....
I won’t spoil who won the argument.
[0]https://www.usnews.com/news/articles/2017-06-02/jury-convict...
https://fija.org/news-events/2020/july/keith-wood-conviction...
This isn't so much about moralizing as much as it is businesses deciding what to do to make the most money. That doesn't mean you can't disagree with it, far from it. But I think the framing of, "these companies are imposing their morality on me" is a misdiagnosis. I don't think it's really a moral position for them, it's a product engineering position.
I would describe the situation as, "more and more of the world is controlled by large corporations, and I'm increasingly subject to their arbitrary and unaccountable decisions. Many of which make no sense from my vantage point."
If you are training you own model, it would be nice to know what, if any, techniques you could employ to balance the effectiveness of it with generality.