EDIT: scuba diving NOT scooba diving
EDIT: scuba diving NOT scooba diving
If I read your hobby website about photography and use it to take 1% better pictures, do I owe you 1% of what my clients pay me?
I think that probably most people would say no, assuming you could even determine that 1% in a way that both parties agreed was fair. I think generally, we have an understanding that some stuff is put out into the world for other humans to learn from and use to make themselves better, and that they don’t owe the original authors anything other than the price of admission.
I guess it comes down to this: do we think that training a model is:
- like storing and later reproducing a version of some collected data, or
- like learning from collected data, and synthesizing new info?
Is there even a meaningful distinction, for a computer?
(Is there even a meaningful distinction for a human…?)
Things get interesting at corporate scale. There are fat VC funds, executives, board of directors and what not - making far more and far more comfortable than an individual trying to get better at their craft to put food on the table. And on top of that, you don't give me access to the product that was refined on my input.
It is like someone learning photography from my website but later taking a really masterpiece shot but asking me for money each time I want to view the photo in their studio.
There are no easy answers, I concur.
Thanks for your comment though, really. :)
> do I owe you 1% of what my clients pay me?
I would still derive some immaterial gain or satisfaction from you reading my website specifically and using what you learnt to improve yourself. As I expect most people would, so it's still a give and take relationship. LLMs sever that link.
It is doubtful many people will be as willing to continue "putting stuff out into the world" if they know that they are only contributing to some sort of (arguably semi-dystopian) hive-mind.
IMHO whether what they are doing or not is justifiable from a legalistic perspective is tangential and not that relevant if we're talking about free/non-commercial content.
Do they though? I mean, do you personally have a link to the people that are consuming the content you post publicly?
I find all the vitriol around LLMs being trained on public data to be a bit weird. If you don't want that data being used then don't publish it for the world to see? Why get mad when you are the one freely publishing the data in the first place? That's like posting your content on a bulleting board in the dorm common room and telling the trust-fund kids they can't read it because they are rich and you don't want them learning anything from you that might make them richer. Maybe a bad analogy, but I feel like it's a fair approximation of the vitriol I see.
If I watch a youtube video my browser is also in a way scraping youtube and storing a (temporary) copy of the video. Does it make sense to protect the protect the owner's right's at this point? Absolutely not. Instead we wait to see if I share that downloaded video or content from it again, or somehow reuse it in my own products. Only then does the law step in.
You can go to a library to borrow a book, but you can't go to the library and copy all the books for your own use.
It's just impractical to photocopy every page of every book in a library.
Otherwise, intellectual property laws can perhaps apply.
It'd be a hard push to claim it's fair use, a wholesale copying of other's works.
You can actually copy all the book but the things is you can't publish it as your own book after you copied. Because obviously it is not your work.
Why? What's stopping me from doing that? The only limitation is time.
AI training is basically only extractive and has the potential to severely disrupt the actual field that made the AI systems possible at all. It's a much more mechanical process that the human interaction of studying a master. It doesn't develop any human skills.
Even if the processes were the same (and I don't think they are, as someone who has actually done computational psychology research), I would still think the AI companies are doing something they know is harmful to actual creative people that generate real value.
I believe no. Most people would make a distinction between “normal” and “rich”. They would give normal people free access, but the rich should pay for it.
It’s like a billionaire asking for a free hot dog. It’s like “come on, you can easily pay $100, which could even sponsor it for the next 100 people”.
Here it’s not the AI itself that’s exploiting you. It’s the rich people that make the AI that get even richer - partly thanks to your free work.
The reason copyright law exists in the first place is due to the difference of scale between copying books by hand and using a machine to do it, so I think "it's different because a machine is doing it" is a completely rational stance to take.
Far in the future - if ever - where we have biological grade artificial beings which you can't program, control and limit in the classical software development sense, this could be rethought.
Until then, we don't need to humanize machines.
But you don't get any of that from an LLM.
The moral thing to do would be to use opt-in training data.
It is an interesting question. I would have no qualms paying for a textbook or university course for curated learning (worth noting OpenAI has paid datasets too), but paying for (or being paid for) relatively diffuse and low quality content through hobby blogs seems at odds with my expectations as an individual, and as a society we were never (en masse) concerned about things like Google's search excerpt answers...
With a gardened proprietary paywalled model, what I wrote ends up as some constituent of giant arrays of floating point numbers which I must pay to use.
If humans could perfectly remember information, I’m sure copyright would be very different.
A contrarian take to support the original commenter is that if the site owner had ads, i probably got him or her some increment in site visits and helped in some small way with monetization, site ranking and boosted his / her public persona, credibility.
When GPT bot visits, none of that happens. Much worse - people who might have visited the hobby site and contributed to traffic and ad revenue will now start getting their answers from the OpenAI chatbot and never visit this hobby site.
That's exploitation and I think that's what most of the responses on this thread miss.
When a private company takes the sum of human knowledge without permission, attribution or payment and then monetizes it via the back door whilst cutting of any connection between the intended consumer and publisher, then we're dealing with a system I'd describe as criminal. It cannot be morally defended as "fair" in any major economical or political system.
The fact that they call it "Open" AI shows the level of trolling involved.
This is a bit pedantic but the term is "scuba diving". Scuba is an acronym that's short for "self contained underwater breathing apparatus". It doesn't work if you don't spell it right.
If so, then pay.
You can't be serious. Thank god our world doesn't work like that.
What do you think why writers and actors have included AI in the reasons of their strike?
Because they are about to become obsolete, and they believe that screaming as loudly as they can is going to stop that.
Their chances of success are roughly the same as if they were protesting against the law of gravity.
AI will not make writers “obsolete”, that is utterly absurd. Would you say reality TV made tv writers obsolete? No? Oh well.
You get what you pay for. That includes what you pay for as a producer…
Of course. And those so-called "computers" won't make human calculators obsolete. After all, they are as large as an entire room, and by the time they are ready to receive input, a human with his slide rule has already computed three and a half entire logarithms!
Human creative professions have 5-10 years left, if they are very lucky.
So in that sense do developers have ~2 years left? Code is much more rigid than acting or creative writing and AI seems to be getting there first. I mean if the all-powerful AI can make modern movies than clearly it can handle writing all code right?
I think they're _very_ worried and rightfully so. I assume it would be very difficult to cancel an AI.
We make our own rules. We decide what to allow and what to value. If technology changes something, it's because we let it.
Yes, I might run a course or something. You are still not entitled to pay.
If you can train your bot on my blog post about scuba diving without my permission and then people can ask your bot for scuba diving advice instead of reading my blog, that doesn't seem very fair.
[1]: https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....
No, you don’t.
That’s a factor weighing in favor of fair use, but the fair use factors are not defined in such a way that that is a necessary factor.
But you do get paid in kind - you "gave" information for the AI to train on, the aI gives you information back, contextualised to your needs. Sometimes those 1000 tokens are worth much more than $0.06
You still need to be able to pay for inference costs, it's crowded and expensive on GPUs nowadays.
I think you'll find if you do try to push Google Search too far, its not quite "limitless" either.
Hasn't happened in a long time.
openai is worth $29,000,000.00 you contributed 0.00000000001
punches numbers in calculator
thus the value of your free credits is 0.001 cents. minus any accounting fees.
You might have missed a few zeros.