https://noyb.eu/en/twitters-ai-plans-hit-9-more-gdpr-complai...
https://noyb.eu/en/twitters-ai-plans-hit-9-more-gdpr-complai...
"He used one of my eggs to irreversibly make a cake"
It's true, but it would be kind of amazing if it weren't
From this knowledge, you can ask the AI "Who has driven to Mexico" and it might know that Ben Adams did, and reply with that.
HOWEVER it's also baked into the model and can't be surgically removed after a complaint. That's the irreversibility part. You can't undo isolated training. You need to provide it a new data set and train it all over again. They won't do that because it's too costly.
The problem with the above example is of course that it can also contain sensitive or private user details.
I've easily extracted the complete song lyrics to the letter from GPT-4 even if OpenAI try to put up guardrails against it due to the copyright issues. AI is really still in the wild west phase...
"He used his eyes to irreversibly read this post"
Additional Twitter data is in my eyes mostly low quality content, that's nothing I would want in a AI model.
How does it matter even if the quality is high or low? The point is user data was used without consent.
That may be considered a feature.
ChatGPT seems reasonably concise, Gemini's answers tend to be verbose (without adding meaningful content).
In particular, requesting an analysis of the problem first before jumping to conclusions can be more effective than just asking for the final answer directly.
However, this analysis phase, or similar one, could just be done hidden in the background, but I don't think any are doing that yet. From the user point of view that would be just waiting, and from API point of view those tokens would also cost. Might just as well entertain the user with the text it processes in the meanwhile.
That doesn't seem to be the case any more and there has been speculation this is down to the star method being used for training newer models. I say speculation because I don't believe people have come out and said they are using star for training. OpenAI referred to Q* somewhere but they wouldn't be drawn on whether that * is this "star" and although google were involved in publishing the star paper they haven't said gemini uses it (I don't think).
(But regardless, many people raised an issue of OpenAI training from sources they shouldn't be allowed to access, so they're definitely a problem as well)
They provide a product in the EU, therefore they must either follow EU law or exit the EU market. Just like an EU company that provides a product in the US has to follow US law.
The line of 'following the law of another country' is grey area on the internet, given that it goes both ways:
EU online companies providing services to US users fail to provide the free speech guarantees that the US laws afford their citizens. That's because all EU countries have more strict laws limiting free speech. Should the EU companies break their own countries' law to satisfy the US audience?
Exactly what is "free speech guarantees" in the context of a private business?
So it seems there are states where a europeans social medium should abide by rules that would most likely contradict european laws, right?
Could you sharpen up this claim? Like suppose I run a microblogging site but I delete libellous posts and incitements to violence in accordance with my local European law. Am I violating a US law by allowing Americans to use the site?