This should be mandatory, enforced, and come with strict fines for companies that do not comply.
This should be mandatory, enforced, and come with strict fines for companies that do not comply.
It's literally just running an algorithm over your data and spitting out the results for you. Fundamentally it's no different from spellcheck, or automatically creating a table of contents from header styles.
As long as the results stay private to you (which in this case, they are), I don't see what the concern is. The fact that the algorithm is LLM-based has zero relevance regarding privacy or security.
I don't want any results from AI. I don't even want to see them. And there is too much of a grey area. What if they use how I use the results to improve their AI. I hate AI also and want nothing to do with its automations.
If I want a document summarized, I will read it myself. I still want to be human and do things AT A REASONABLE LEVEL with my own two hands.
OK, sure. But then just don't use it.
The problem is that you're calling for a legal policy against it to be "mandatory, enforced, and come with strict fines".
Have your own personal preferences, that's great. But I don't want you imposing your preferences on the products I use. I want companies and the market to decide.
An auto-summary feature that is enabled by default is not something we should be asking for government regulation over, any more than we should be asking the government to prohibit wavy red lines unless they're explicitly opted into.
That's why I seriously recommend everyone everywhere regularly replace their blinker fluid and such.
Early search engines had a problem, which was that when they crawled willy nilly, people would block their IP addresses. Inventing this concept of `robots.txt` worked because search engines wanted something: to avoid IP blocks, which they couldn't easily get around. And site hosts generally wanted to be indexed.
Today it's WAY harder to block relevant IP addresses, so site hosts generally can't easily block a crawler that wants its data: there is no compromise to be found here, and the imbalance of power is much stronger. And many site hosts generally don't want to be crawled for free for AI purposes at all. Pretty much anyone who sets up an `ai.txt` uses it to just reject all crawling, so there is no reason for any crawler to respect it.
That's a game of whack-a-mole that always lets a few miscreants through. I used to find that an acceptable amount of error until I learned that crawlers were gathering data to be used to train LLMs. That's a situation where even a single bot getting through is very problematic.
I still haven't found a solution to that aside from no longer allowing access to my sites without an account.
IMO the bigger concern is that this data is not just used to train models. It is stored, completely verbatim, in the training set data. They aren’t pulling from PDFs in realtime during training runs, they’re aggregating all of that text and storing it somewhere. And that somewhere is prone to employees viewing, leakage to the internet, etc.
Isn't this like encryption, though?
I'm fairly sure that the cryptography community basically says: if someone has a copy of your encrypted data for a long time, the likelihood over time for them to be able to read it approaches 100%, regardless of the current security standard you're using.
Who could possibly guarantee that whatever LLM is safe now will be safe at all times over the next 5-10-20 years? And if they're guaranteeing, they're lying.
To which the reply was "they'll just make the LLM able to better defend itself".
And my point was "the attackers will learn to build better prompts, too".