That won't be possible anymore, anything on the internet would be considered public, it's no longer considered private if you didn't keep it private.
> To protect the privacy of the individual and to strengthen the freedom of the information which concern the public the yet last point was added.
The privacy of individuals is important, regardless where they store their private data. Their account information -- what they buy, their medical information and so on is stored on servers and could be hacked.
I think the second point is equally clear, and further up on the list.
I agree that what people buy, their medical information and so on should be private, hence it should only be offline and not stored/handled on computers connected to the internet at all, the internet should be for public data exclusively, is my argument in the initial comment. Medical information would be only on effectively airgapped computers, as that data should be private, as you say.
To me private file stay private, even if they get leaked onto the open or closed internet, because the public has no right to know them, they are private data of an individual. They might no longer be secret, but they are still private.
> To me private file stay private, even if they get leaked onto the open or closed internet
We have very different definitions of what "private" means. Once it's leaked, it's no longer private, and pretending it can go back to being "private" after being on the public internet, is doing no one any favors.
If private data got leaked, like revenge porn, it is a breach and that private data that belongs to an individual is still private, and still needs to be protected. This is what GDPR and other legislation is about. If secrecy and privacy is the same, someone that isn't able to protect their data sufficiently will not have any privacy, thus any leaking of data is now the fault of the person that got their data leaked, not of the person that broke the trust and leaked it.
Your conclusion seems rather extreme to me. So of course data that got leaked, and is no longer secret is still private, because 'private' means who should be in control of that data, not about if the person has control or not.
I also don't follow your point about waffles and pancakes, because this is a pretty big disagreement we have here. To me this dialog is more like you are saying "I don't like laws", and I say "While I agree that some laws are stupid, other laws are pretty useful, for instance people shouldn't be allowed rob other houses, even if they are able to do that or even where invited." And then, instead of agreeing, you sort of say, "No, I really mean that. If someone isn't able to defend their home properly or give out invitations to someone, it is okay to steal from them."
Personally, I'd prefer a fixed term. I know enough independent authors making a living from selling their books that I'm willing to allow the fixed term to be large, like 50 years from date of completion of the work. (With a good definition of "completion" so someone can't cheat by editing a couple lines per year to keep something copyrighted indefinitely). The simpler the rule is, the easier it is to understand, and the harder it is to cheat it. The more complicated you make a rule, the more loopholes get found.
One can make the case that this period should be more limited, or that the combination should be capped, but life+X is the right formulation, I think.
Why can't they do what the rest of us do? Earn and save money during your working life and leave _that_ for your heirs. Let copyright die with the author.
I know this them-and-us thinking is fashionable in the tech world but the reality is that the majority of creative people don’t earn much and never have, and copyright was developed not to give them extra power over the rest of us but to create a framework for creative work to earn them an income at all.
You should read about it.
But I do think some potential post-mortem protection is essential for creative work to remain viable, and that means that any post-mortem buyer of an artist's estate has to be able to get value from recent work for a period of time.
This whole discussion is somewhat fantastical now anyway, because copyright is fucked.
But the intent was always to make working artists' lives possible; the various copyright extensions have always been for the benefit of corporate copyright holders, and it is unfair to vilify individual working artists for that.
That's quite a narrow definition of what motivates creative work
It is working artists we are talking about; working artists work for money.
That money, in the post-patronage era, comes from exercising copyright. The reason the copyright can’t simply die with them is that this tends to dissuade the creation of long-gestating work.
Copyright was developed to make it possible for artists, writers, musicians etc. to work for long periods on work of significance with no income, on the basis of the future, deferred earnings of the work, without their work being stolen from them, and it gives them the limited right to direct how their work is monetised on their behalf, including establishing publishing rights etc.
Some protection after death is a key component of that, because people do die while they are still working.
I think it is serious. In which case, I gotta say, it really seems like you didn't spend much time thinking about this. "A 1 year grave period for everyone to pull stuff off they don't want to be a part of" - How does that work when the Internet is already full of unauthorized reproductions, most of which people aren't even aware of? Even ignoring practical considerations, when literally everyone is basically stuck using the Internet for everything, this seems a bit unfair to anyone who isn't onboard, akin to The Onion's Google Opt-out Village. But there are so many practical issues with this, it would be easier to list the number of problems this doesn't have. You accidentally leak something to the Internet and it becomes commons? What happens when other people leak things to the Internet? How about revenge porn?
Not minor stuff that can easily be papered over, this literally reintroduces the problem of needing to care about the provenance of data again, in a way that can't be automated, which makes the whole thing entirely moot. All just to make training data for AI models easier to distribute?
I'm all for intellectual property reform, maybe even fairly radical. But this just seems like it wasn't thought out.
If this was satire, well, I took the bait. Oddly convincing despite being hard to believe.
It wasn't entirely serious, but also not entirely un-serious. But yes, I spent maybe 20-30 seconds thinking about then barfed up the text that makes the comment, so yes, obviously many issues and not really workable in practice.
I'm glad it made you seriously think about it and also flip-flopp back and forth about it, made it worth posting the comment so happy to hear :)
That's a fair amount of computational and labor overhead mind you, as you'll need to verify and prune the quality of your mountain of synthetic data, but certainly possible.
Though this assumes the legal system is a rational actor playing by the set of rules it claims to. In fact, I highly suspect you could get very unlucky and get an unfavorable ruling against you, because you stepped on a big pile of money's toes in the process of doing this.
Are LLMs what we need to make all data public domain? This way it could be used for that purpose
The first broadly useful fully open source models will do this.
We already have open data / open code / open weights for some domain-specific cases, such as audio models trained on large open datasets, eg. Tacotron / LJSpeech from waaay back in the day, though that is certainly not SOTA anymore.
Distillation could possibly be considered an early case of this as raw AI outputs are themselves not copyrightable unless humans enrich, filter, or transform them. Granted, that does not handle the cases where the outputs are sufficiently similar to copyrighted original works.
That said, I don't necessarily disagree with you. Talkie[1] presents an interesting case for it being at least possible to do this entirely on public domain material.
But even that used Claude somewhere in the course of its training pipeline (it's listed as a contributor on their GitHub), so again, how granular you want to get with that is still a question.
Decentralized unstoppable storage, combined with decentralized unstoppable training, sorta like SETI for AI training. The seed of this tech already exists with IPFS and others like it.
We know (some? all?) of the big labs have skirted copyright laws at one point or another. Truly open models would just build on what is publicly available.