I guess technically it's supposed to play some role in making sure OpenAI "benefits humanity". But as we've seen multiple times, whenever that goal clashes with the interests of investors, the latter wins out.
I guess technically it's supposed to play some role in making sure OpenAI "benefits humanity". But as we've seen multiple times, whenever that goal clashes with the interests of investors, the latter wins out.
That entity will scrape the internet and train the models and claim that "it's just research" to be able to claim that all is fair-use.
At this point it's not even funny anymore.
It does suit the modus operandi of a number of American companies that start out as literally illegal/criminal operations until they get big and rich enough to pay a fine for their youthful misdeeds.
By the time some of them get huge, they're in bed with the government to dominate the market.
Or we're all talking about and envisioning some specific little subset of artists. I suspect you're trying to pretend that someone with a literal set of paintbrushes living in a shitty loft is somehow having their original artwork stolen by AI despite no high resolution photography of it existing on the internet. I'm not falling for that. Be more specific about which artists are losing their livelihoods.
If we take your argument to it's logical conclusion, all progress is inherently bad, and should be stopped.
I deposit instead that the real problem is that we tied people's ability to afford basic necessities to how much output they can produce as a cog in our societal machine.
Yes, because if you depend on some overarching organisation or person to give it to you, you are fucked 100% of the time due this dependency.
If AI replaces millions of jobs, it will be a net negative in job availability for working class people.
I agree with your last point, the way the system is set up is incompatible with the looming future.
There is no similar net creation of jobs for society if jobs are eliminated by AI, and it's even worse than that because many of the jobs are specialized, high-skill positions that can't be transferred to other careers easily. It goes without saying that it also includes millions of low-skill jobs like cashiers, stockers, data entry, CS reps, etc. Generally people who are already struggling to get enough hours and feed their families as it is.
Recording devices permitted artists to sell more art.
Many of the uses of AI people get most excited about seem to be cutting the expensive human creators out of the equation.
Current AI can greatly elevate what a beginning artist can produce. If you have a decent grasp of proportions, perspective and good ideas, but aren't great at drawing, then using AI can be a huge quality improvement.
On the other hand if you're a top expert that draws quickly and efficiently it's quite possible that AI can't do very much for you in a lot of cases, at least not without a lot of hand tuning like training it on your own work first.
So yeah it had a profound effect, but we got consent for the parts that fundamentally relied on other people.
It's what everyone else does. The entitlement has to stop.
You're advocating for destroying all AI or ensuring a monopoly by corporations. Whose side are you actually on?
Irrelevant. The law does not care about feasibility of breaking it.
If I decide to run a hit man business, that's also infeasible. Dealing with the arrests and fines would be too much. The conclusion then is not to bend the law to make murder legal. The conclusion is my business is illegitimate, and it's the civic duty of my Country to make sure it fails.
> Whose side are you actually on?
The people making the content that corps are profiting big off of. They should pay a license.
Nope. The law will side with whoever pays the most. Once OpenAI solidifies its top position, only then will regulations kick in. Take YouTube, for example—it grew thanks to piracy. Now, as the leader, ContentID and DMCA rules work in its favor, blocking competition. If TikTok wasn’t a copyright-ignoring Chinese company, it would’ve been dead on arrival.
If it's only OK to scrape, lossy-compress, and redistribute book-paragraphs when it gets blended into a huge library of other attempts, then that's only going to empower big players that can afford to operate at that scale.
I've no idea if it could be valid when it comes to OpenAI, but it does seem to be a general concept designed to counter wrongdoers who take a little value from a lot of people?
The crazy thing is that there hasn't been an injunction to make them stop.
burning the bridge so nobody else can legally scrape, that's the line.
Assets like the Internet Archive, though, should be protected at all costs.
Update: ML doesn't copy information. It can merely memorise some small portions of it.
https://www.copyright.gov/title37/201/37cfr201-14.html
§ 201.14 Warnings of copyright for use by certain libraries and archives.
....
The copyright law of the United States (title 17, United States Code) governs the making of photocopies or other reproductions of copyrighted material.
Under certain conditions specified in the law, libraries and archives are authorized to furnish a photocopy or other reproduction. One of these specific conditions is that the photocopy or reproduction is not to be “used for any purpose other than private study, scholarship, or research.” If a user makes a request for, or later uses, a photocopy or reproduction for purposes in excess of “fair use,” that user may be liable for copyright infringement.
This institution reserves the right to refuse to accept a copying order if, in its judgment, fulfillment of the order would involve violation of copyright law.
You can make a copy. If you (the person using the copied work) are using it for something other than private study, scholarship, research, or reproduction beyond "fair use", then you - the person doing that (not the person who made the copy) are liable for infringement.It would be perfectly legal for me to go to the library and make photocopies of works. I could even take them home and use the photocopies as reference works write an essay and publish that. If {random person} took my photocopied pages and then sold them, that would likely go beyond the limits placed for how the photocopied works from the library may be used.
A more fitting metaphor would be something like... If you had the ability to read all the books in the library extremely quickly, and to make useful mental connections between the information you read such that people would come to you for your vast knowledge, should you be allowed in the library?
The anti-AI stance is what is baffling to me. The path trotten is what got us here and obviously nobody could have paid people upfront for the wild experimentation that was necessary. The only alternative is not having done it.
Given the path it has put as in, people either are insanely cruel or just completely detached from reality when it comes to what is necessary to do entirely new things.
"Hugely beneficial" is a stretch at this point. It has the potential to be hugely beneficial, sure, but it also has the potential to be ruinous.
We're already seeing GenAI being used to create disinformation at scale. That alone makes the potential for this being a net-negative very high.
Perhaps the biggest “needs citation” statement of our time.
Not in any weirdly-self-aggrandizing "our tech is so powerful that robots will take over" sense, just the depressingly regular one of "lots of people getting hurt by a short-term profitable product/process which was actually quite flawed."
P.S.: For example, imagine having applications for jobs and loans rejected because all the companies' internal LLM tooling is secretly racist against subtle grammar-traces in your writing or social-media profile. [0]
To continue one of the analogies: Plenty of people and industries legitimately benefited from the safety and cost-savings of asbestos insulation too, at least in the short run. Even today there are cases where one could argue it's still the best material for the job--if constructed and handled correctly. (Ditto for ozone-destroying chlorofluorocarbons.)
However over the decades its production and use grew to be over/mis-used in so very many ways, including--very ironically--respirators and masks that the user would put on their face and breathe through.
I'm not arguing LLMs have no reasonable uses, but rather that there are a lot of very tempting ways for institutions to slot them in which will cause chronic and subtle problems, especially when they are being marketed as a panacea.
We have a term for that, it's called "luddite". Those were english weavers who would break in to textile factories and destroy weaving machines at the beginning of the 1800s. With the extreme rare exception, all cloth is woven by machines now. The only hand made textiles in modern society are exceptionally fancy rugs, and knit scarves from grandma. All the clothing you're wearing now are woven by a machine, and nobody gives this a second thought today.
https://en.wikipedia.org/wiki/I%27m_alright,_Jack
Except, we are all Jack.
Maybe it would have been better for humanity if the Luddites won.
It is not possible to rehabilitate the Luddites. If you insist on attempting to do so, there are better venues.
No, that's apples-to-oranges. The goals and complaints of Luddites largely concerned "who profits", the use of bargaining power (sometimes illicit), and economic arrangements in general.
They were not opposing the mechanization by claiming that machines were defective or were creating textiles which had inherent risks to the wearers.
I have never thought of being anti-AI as “Luddite”, but actually this very description of “Luddite” does sound like the concerns are in fact not completely different.
Observe:
Complaints about who profits? Check; OpenAI is earning money off of the backs of artists, authors, and other creatives. The AI was trained on the works of millions(?) of people that don’t get a single dime of the profits of OpenAI, without any input from those authors on whether that was ok.
Bargaining power? Check; OpenAI is hard at work lobbying to ensure that legislation regarding AI will benefit OpenAI, rather than work against the interests of OpenAI. The artists have no money nor time nor influence, nor anyone to speak on behalf of them, that will have any meaningful effect on AI policies and legislation.
Economic arrangements in general? Largely the same as the first point I guess. Those whose works the AI was trained on have no influence over the economic arrangements, and OpenAI is not about to pay them anything out of the goodness of their heart.
The Luddites were actually a fascinating group! It is a common misconception that they were against technology itself, in fact your own link does not say as much, the idea of “luddite” being anti-technology only appears in the description of the modern usage of the word.
Here is a quote from the Smithsonian[1] on them
>Despite their modern reputation, the original Luddites were neither opposed to technology nor inept at using it. Many were highly skilled machine operators in the textile industry. Nor was the technology they attacked particularly new. Moreover, the idea of smashing machines as a form of industrial protest did not begin or end with them.
I would also recommend the book Blood in the Machine[2] by Brian Merchant for an exploration of how understanding the Luddites now can be of present value
1 https://www.smithsonianmag.com/history/what-the-luddites-rea...
2 https://www.goodreads.com/book/show/59801798-blood-in-the-ma...
They had very rational reasons for trying to slow the introduction of a technology that was, during a period of economic downturn, destroying a source of income for huge swathes of working class people, leaving many of them in abject poverty. The beneficiaries of the technological change were primarily the holders of capital, with society at large getting some small benefit from cheaper textiles and the working classes experiencing a net loss.
If the impact of LLMs reaches a similar scale relative to today's economy, then it would be reasonable to expect to see similar patterns - unrest from those who find themselves unable to eat during the transition to the new technology, but them ultimately losing the battle and more profit flowing towards those holding the capital.
> "our tech is so powerful that robots will take over"
> "lots of people getting hurt by a short-term profitable product/process which was actually quite flawed."
You response assumes the former, but it's my understanding the Luddite's actual position was the latter.
> Luddites objected primarily to the rising popularity of automated textile equipment, threatening the jobs and livelihoods of skilled workers as this technology allowed them to be replaced by cheaper and less skilled workers.
In this sense, "Luddite" feels quite accurate today.
We don't have to imagine such things, really, as that's extremely common with humans. I would argue that fixing such flaws in LLMs is a lot easier than fixing it in humans.
I currently work in the HR-tech space, so suppose someone has a not-too-crazy proposal of using an LLM to reword cover-letters to reduce potential bias in hiring. The issue is that the LLM will impart its own spin(s) on things, even when a human would say two inputs are functionally identical. As a very hypothetical example, suppose one candidate always does stuff like writing out the Latin like Juris Doctor instead of acronyms like JD, and then that causes the model to end up on "extremely qualified at" instead of "very qualified at"
The issue of deliberate attempts to corrupt the LLM with prompt-injection or poisonous training data are a whole 'nother can of minefield whack-a-moles. (OK, yeah, too far there.)
I just don't think your original comment was entirely fair. IMO, LLMs and related technology will be looked at similarly as the Internet - certainly it has been used for bad, but I think the good far outweighs the bad, and I think we have (and continue to) learn to deal with the issues with it, just as we will with LLMs and AI.
(FWIW, I'm not trying to ignore the ways this technology will be abused, or advocate for the crazy capitalistic tendency of shoving LLMs in everything. I just think the potential for good here is huge, and we should be just as aware of that as the issues)
(Also FWIW, I appreciate your entirely reasonable comment. There's far too many extreme opinions on this topic from all sides.)
Some not-problems, presented as though they are:
"How can we prevent the untimely eradication of Polio?"
"How can we prevent bot network operators from being unfairly excluded from online political discussions?"
"How can we enable context-and-content-unaware text generation mechanisms to propagate throughout society?"
For example, MKUltra tried to solve a problem: "How can I manipulate my fellow man?" That problem still exists today, and you bet AI is being employed to try to solve it.
History is littered with problems such as these.
Yes, we are clearly talking about things to mostly still come here. But if you assign a 0 until its a 1 you are just signing out of advancing anything that's remotely interesting.
If you are able to see a path to 1 on AI, at this point, then I don't know how you would justify not giving it our all. If you see a path and in the end using all of human knowledge up to this point was needed to make AI work for us, we must do that. What could possibly be more beneficial to us?
This is regardless of all issues the will have to be solved and the enormous amount of societal responsibility this puts on AI makers — which I, as a voter, will absolutely hold them accountable for (even though I am actually fairly optimistic they all feel the responsibility and are somewhat spooked by it too).
But that does not mean I think it's responsible to try and stop them at this point — which the copyright debate absolutely does. It would simply shut down 95% of AI, tomorrow, without any other viable alternative around. I don't understand how that is a serious option for anyone who roots for us.
Oh, the humanity! Who will write our third-rate erotica and Russian misinformation in a post-AI world?
I don’t think that the consumer LLMs that openai is pioneering is what need optimism.
AlphaFold and other uses of the fundamental technology behind LLMs need hype.
Not OpenAI
AlphaFold is a game changer for medical R&D. Everyone should be hyped for that.
They also are leveraging these same ML techniques for detecting kelp forest off the coast of Australia for preservation.
Alphabet isn’t a great company, but that does not mean the good they do should be ignored.
Much more deserving than chatgpt. Productifyed LLMs are just an attempt to make a new consumer product category.
I think you raise some interesting concerns in your last paragraph.
> enormous amount of societal responsibility this puts on AI makers — which I, as a voter, will absolutely hold them accountable for
I'm unsure of what mechanism voters have to hold private companies accountable. Fir example, whenever YouTube uses my location without me ever consenting to it - where is the vote to hold them accountable? Or when Facebook facilitates micro targeting of disinformation - where is the vote? Same for anything AI. I believe any legislative proposals (with input from large companies) is very likely more to create a walled garden than to actually reduce harm.
I suppose no need to respond, my main point is I don't think there is any accountability thru the ballot when it comes to AI and most things high-tech.
Firstly, *skeptics.
Secondly, being skeptical doesn't mean you have no optimism whatsoever, it's about hedging your optimism (or pessimism for that matter) based on what is understood, even about a not-fully-understood thing at the time you're being skeptical. You can be as optimistic as you want about getting data off of a hard drive that was melted in a fire, that doesn't mean you're going to do it. And a skeptic might rightfully point out that with the drive platters melted together, data recovery is pretty unlikely. Not impossible, but really unlikely.
Thirdly, OpenAI's efforts thus far are highly optimistic to call a path to true AI. What are you basing that on? Because I have not a deep but a passing understanding of the underlying technology of LLMs, and as such, I can assure you that I do not see any path from ChatGPT to Skynet. None whatsoever. Does that mean LLMs are useless or bad? Of course not, and I sleep better too knowing that LLM is not AI and is therefore not an existential threat to humanity, no matter what Sam Altman wants to blither on about.
And fourthly, "wanting" to stop them isn't the issue. If they broke the law, they should be stopped, simple as. If you can't innovate without trampling the rights of others then your innovation has to take a back seat to the functioning of our society, tough shit.
> The anti-AI stance is what is baffling to me
I don't see s lot of anti AI but instead I see a concern for how it's just being managed and controlled by the larger companies with resources that no start up could dream. Open AI was to release it's models and be well.. Open but fine they're not. But their behaviour of how things are proceeding are questionable and unnecessarily aggravating.
I don't think this is the "ends justify the means" argument you think it is.
Automobiles allow people to travel great distances over short periods of time, increase physical work capacity, allow for building massive structures, and allow for farming insane amounts of food.
Both the internet and automobiles have positively affected my life, and I assume the lives of many others. How are any of these aimless questions?
And who is the one calling for action?
Sorry for being dense, but I'm trying to understand if I'm the "strong" or the "weak" in your analogy.
The work of artists, authors, etc.
I know currently the legal situation is messy, but that's exactly the point, anyone who can't engage in lengthy legal battle and defend their position in court are being sacrificed. The companies behind LLMs are spending hundreds of millions of dollars in lobbying and exploiting loopholes.
Let's be real without the data there wouldn't be LLMs, so it crazy that some people are downplaying its significance or value, while on the other hand they're losing sleep over finding fresh sources to scrape.
The big publishers seem to have given up and decided it's best to reach agreement with their counterparts, while independent authors are given the finger.
> The work of artists, authors, etc.
Define "a lot"? Most people barely know how to use their email. Even among the minority who do actively use "AI" and excited about it, outside of ML engineers they aren't well-informed or aware what data is used for training, or even what training means and how these models work to begin with.
> People are excited and enthusiastic about AI and are actively reaping the benefits of progress.
Except the terms were already violated in the initial training phase before the services were even public and saw adoption. That's like pointing at a rape victim who got some form of compensation later, saying:
see how she's "reaping the benefits"
So let's not play the people wanted it card.By the time some people started raising concerns, OpenAI claimed the cat was already out of the bag and "if we didn't do it, someone else will, so deal with it."
Similar to privacy, just because some people don't care, lack awareness, or don't want the hassle of fighting for it, doesn't justify taking it away from others.
This is repackaging content, laundering it, and reselling it.
As others have noted, IP law has lots of problems; Sam Altman et al are exploiting the gap left between the speed of technology and law and using their own version of social good without waiting for the consent of those they're exploiting.
The anti-AI stance is what is baffling to me.
I think it’s unfair to paint any legal controls over this incredibly important, high-stakes technology as being “anti”. They’re not trying to prevent innovation because they’re cruel, they’re just trying to somewhat slow down innovation so that we can ensure it’s done with minimal harm (eg making sure content creators are compensated in a time of intense automation). Like we do for all sorts of other fields of research, already!And isn’t this what basically every single scholar in the field says they want, anyway - safe, intentional, controlled deployment?
As you can tell from the above, I’m as far from being “anti-AI” or technically pessimistic as one can be — I plan to dedicate my life to its safe development. So there’s at least one counterexample for you to consider :)
OpenAI's case is especially egregious, with the entire starting as 'open' and reaping the benefits, then doing its best in every way to shut the door after itself by scaring people over AI apocalypses. If your argument is seriously that it is necessary to shamelessly steal and lie to do new things, I question your ethical standards, especially in the face of all the openly developed models out there.
I do not have confidence in the Supreme Court in general, and I think there's a real risk that in deciding on AI training they upend copyright of digital materials in a way that makes it worse for everyone.
It's completely unprecedented.
We allowed scraping images and text en masse when search engines used the data to let us find stuff.
We allow copying of style, and don't allow writing styles and aesthetics to be copyrighted or trademarked.
Then AI shows up, and people change lanes because they don't like the results.
One of the things that made me tilt towards the side of fair use was a breakdown of the Stable Diffusion model. The SD2.1 base model was trained on 5.85 billion images, all normalized to 512x512 BMP. That's 1MB per images, for a total of 5.85PB of BMP files. The resulting model is only 5.2GB. That's more than 99.999999% data loss from the source data to the trained set.
For every 1MB BMP file in the training dataset, less than 1byte makes it into the model.
I find it extremely difficult to call this redistribution of copyrighted data. It falls cleanly into fair use.
Their arguments against this amounts to "we're not using it like they intend it to be used, so it's fine if we obtain it illegally", and that's a bs standard, totally divorced from any legal reality.
Fair Use covers certain transformative uses, certainly, but it doesn't cover illegal obtaining of the content.
You can't pirate a book just because you want to use it transformatively (which is exactly what they've done), and that argument would never hold up for us as individuals, so we sure as hell shouldn't let tech companies get a special carve-out for it.
I'm surprised people are surprised.
>> That entity will scrape the internet and train the models and claim that "it's just research" to be able to claim that all is fair-use.
a lot of people and entities do this though... openAI is in the spotlight, but scraping everything and selling it is the business model for a lot of companies...
In my eyes, all genAI companies/tools are the same. I dislike all equally, and I use none of them.
That's the business model of lots of companies. Take, collect and collate data, put it in a new format more useful for your field/customers, resell.
Low level employees are there for the money, not for the drama.
As a moral fig leaf. They can always point to it when the press calls -- "see it is a non-profit".
I'm ignorant on this topic so please excuse me. Why did `AI` happen now? What was the secret sauce that OpenAI did that seemed to make this explode into being all of a sudden?
My general impression was that the concept of 'how it works' existed for a long time, it was only recently that video cards had enough VRAM to hold the matrix(?) within memory to do the necessary calculations.
If anybody knows, not just the person I replied to.w.r.t. Branding.
AI has been happening "forever". While "machine learning" or "genetic algorithms" were more of the rage pre-LLMs that doesn't mean people weren't using them. It's just Google Search didn't brand their search engine as "powered by ML". AI is everywhere now because everything already used AI and now the products as "Spellcheck With AI" instead of just "Spellcheck".
w.r.t. Willingness
Chatbots aren't new. You might remember Tay (2016) [1], Microsoft's twitter chat bot. It should seem really strange as well that right after OpenAI releases ChatGPT, Google releases Gemini. The transformers architecture for LLMs is from 2014, nobody was willing to be the first chatbot again until OpenAI did it but they all internally were working on them. ChatGPT is Nov 2022 [2], Blake Lemoine's firing was June 2022 [3].
[1]: https://en.wikipedia.org/wiki/Tay_(chatbot)
[2]: https://en.wikipedia.org/wiki/ChatGPT
[3]: https://www.npr.org/2022/06/16/1105552435/google-ai-sentient
Thanks for the links too!
1986: Geoffrey Hinton publishes the backpropagation algorithm as applied to neural networks, allowing more efficient training.
2011: Jeff Dean starts Google Brain.
2012: Ilya Sutskever and Geoffrey Hinton publish AlexNet, which demonstrates that using GPUs yields quicker training on deep networks, surpassing non-neural-network participants by a wide margin on an image categorization competition.
2013: Geoffrey Hinton sells his team to the highest bidder. Google Brain wins the bid.
2015: Ilya Sutskever founds OpenAI.
2017: Google Brain publishes the first Transformer, showing impressive performance on language translation.
2018: OpenAI publishes GPT, showing that next-token prediction can solve many language benchmarks at once using Transformers, hinting at foundation models. They later scale it and show increasing performance.
The reality is that the ideas for this could have been combined earlier than they did (and plausibly future ideas could have been found today), but research takes time, and researchers tend to focus on one approach and assume that another has already been explored and doesn’t scale to SOTA (as many did for neural networks). First mover advantage, when finding a workable solution, is strong, and benefited OpenAI.
Honestly, those are not the missing parts that most matter IMO. The evolution of the concept of attention across many academic papers which fed to the Transformer is the big missing element in this timeline.
Not really:
History: https://arxiv.org/abs/2212.11279 (75 pp.)
Survey: https://arxiv.org/abs/1404.7828 (88 pp.)
Conveniently skim-read over the course of the four weekends on one month.
But out of the co-founders, especially if we believe Elon's and Hinton's description of him, he may have been the one that mattered most for their scientific achievements.
We've had upgrades to hardware, mostly led by NVidia, that made it possible.
New LLMs don't even rely that much on that aforementioned older architecture, right now it's mostly about compute and the quality of data.
I remember seeing some graphs that shows that the whole "learning" phenomena that we see with neural nets is mostly about compute and quality of data, the model and optimizations just being the cherry on the cake.
Don’t they all indicate being based on the transformer architecture?
> not entirely because of transformers but because of the hardware
Kaplan et al. 2020[0] (figure 7, §3.2.1) shows that LSTMs, the leading language architecture prior to transformers, scaled worse because they plateau’ed quickly with larger context.
A tale as old as time. Some of us could see it, from afar <says while scratching gray, dusty beard>. Lack of upvotes and excitement does not mean support, but how to account for that in these times? <goes away>
the well known scammer successfully scammed everyone twice. obviously he's keeping it around for the third (and forth...) time