When Google finally made the whole of the web searchable, what would have happened, what would we have said, if they had decided some pages were inappropriate and therefore it should be their duty to prevent users to ever find them?
Why should AI companies become editors of the human experience, and who decides what topics are off-limits?
Sexuality? Isn't sexuality an important topic? Why can it not be part of our discussions with a chatbot?
"Illegal" activities? Aside from the fact that all laws are local, there are many illegal activities that are described in great realistic details in novels, movies, documentaries. If chatbots can't discuss those, how long will it take for the corresponding books to be banned from public libraries?
Etc.
This seems quite vain in any case. LLMs are in many ways comparable to a compressed index of the web; if it's on the Internet, it will end up in a model, and once there, it will be possible to extract it.
because that is the point of "AI". Search is getting less and less useful to both please advertisers and/or to hide piracy/forbidden content. To the point google is useless today.
AI is nothing but an interface for search. Training is just the new indexing.
And being the interface to search, it must control all that google is expected to control today.
Google does this every day.
Its called "deindexing" and Google (and every other search engine, probably, but Google owns enough of the usage that no really cares about the other ones) has done it forever, for a variety of reasons.
And apparently, you wouldn't have said anything, because it happened, and you never even noticed it.
Maybe it also does it on its own and we don't notice it -- but in any case you can find as many "inappropriate" things as you want, searching Google, including images of all nature, or recipes for bad things, or descriptions of crimes, etc.
So if Google is actually in the business of editing the web so that it conforms to puritain standards of appropriateness (which I doubt), it's not very good at it.
Google does unilateral deindexing too, and is open that it does so.
Google may temporarily or permanently remove sites from its index and search results if it believes it is obligated to do so by law, if the sites do not meet Google's quality guidelines, or for other reasons, such as if the sites detract from users' ability to locate relevant information. We cannot comment on the individual reasons a page may be removed. However, certain actions such as cloaking, writing text in such a way that it can be seen by search engines but not by users, or setting up pages/links with the sole purpose of fooling search engines may result in removal from our index. Please read our spam policies pages for more information.
https://support.google.com/webmasters/answer/40052?hl=en
People do notice it, and write “what do if it happens to you” articles about it:
https://www.searchenginejournal.com/deindexed-by-google-how-...
Anyway my point is really that current editorial rules of LLMs seem several orders of magnitude more restrictive than those of the Google index, and I think it's interesting to ask why.
This isnt accurate. There already several tools and methods explicitly for excluding forbidden knowledge from training data, or reweighting the training data so that the LLM have the perspective that the owners want.
e.g. if you want the model from being racist, you run the training data through a classifier to remove data with negative sentiment around race.
https://openai.com/research/dall-e-2-pre-training-mitigation...
Is the the choice to exclude some sources from the training data "censorship"? How about the choice of loss minimization strategy?
Excluding data is censoring from them model. Then when the model provides incorrect or biased, the firm is actively deceiving the user.
Im not sure what you mean by loss minimization strategy. what I said is focused on the case where data is manipulated with the intent to curate a different perspective than an the source data.
In short, fanatics and power mongers are always going to fanatic and power monger.
But, the fact that they're trying doesn't mean they will succeed.
And old secrets very often become common knowledge later on.
Its not, though, and if you think it is, than you don't understand what "censorship" is (a supply-side propaganda mechanism aimed at increasing the average cost of spreading disfavored ideas.)
It's true that complete suppression of ideas (or behaviors) by censorship is generally impossible to achieve, but to mistake that for censorship being impossible to achieve is like saying "making money by work is impossible to achieve" when what you really mean "making enough money to purchase all the property in the world by work is impossible to achieve."