GPT-Neo – Building a GPT-3-sized model, open source and free
eleuther.ai
eleuther.ai
I think of the value proposition of GPT-X as "what would you do with a team of hundreds of people who can solve arbitrary problems only by googling them?". And honestly, not a lot of productive applications come to mind.
This is basically the description of a modern content farm. You give N people a topic, ie "dating advice", and they'll use google to put together different ideas, sentences and paragraphs to produce dozens of articles per day. You could also write very basic code with this, similar to googling a code snippet and combining the results from the first several stackoverflow pages that come up (which, incidentally, is how I program now). After a few more versions, you could probably use GPT to produce fiction that matches the quality of the average self published ebook. And DALL-E can come up with novel images in the same way that a graphic designer can visually merge the google image results for a given query.
One limitation of this theoretical "team of automated googlers" is that what they search for content is cached on the date of the last GPT model update. Right now the big news story is the Jan 6th, 20201 insurrection at the US Capitol. GPT-3 can produce infinite bad articles about politics, but won't be able to say anything about current events in real time.
I generally think that GPT-3 is awesome, and it's a damn shame that "Open"AI couldn't find a way to actually be open. At this point, it seems like a very interesting technology that is still in desperate need of a Killer App
But if we are looking to replace end-to-end intelligence at scale it's not just about synthesis. We need to also automate the peer review process so that it's bandwidth is matched to increased rate of synthesis. Most good researchers and engineers are able to self-critique their work (and the degree to which they can do that well is really what makes one good IMHO). And then we rely on our colleagues and peers to review our work and form a consensus on its quality. Currently GPT-like systems can easily overwhelm humans with such peer review requests. Even if a model is capable of writing the next great literary work, predicting exactly what happened on Jan 6, or formulating new laws of physics the sheer amount of crap it will produce alongside makes it very unlikely that anyone will notice.
The reason for this is that if your adult life consists of just a tiny, tiny, tiny fraction of the total time of all adults, and so if an idea is relevant to more people, odds decrease exponentially that no one thought of it before.
There are always new languages though, so a great strategy is to take old ideas and bring them to new languages. I count new high level, non programming languages as new languages as well.
Yes, every time you see something that for human obviously doesn't make sense it makes you dismiss it. You would look at that output differently though if you were talking with a child. Just like a child can miss some information making it say something ridiculous it may miss some patterns connections.
But have you ever observed carefully how we connect patterns and make sentences? Our highly sophisticated discussions and reasoning is just pattern matching. Then most prominent patterns ordered in time also known as consciousness.
Watch hackernews comments and look how after somebody used a rare adjective or cluster of words more commenters tend to use it without even paying conscious attention to that.
Long story short, give it a try and see what examples of what people already did with it even in it's limited form.
To me you are looking at an early computer and saying that it's not doing anything that a bunch of people with calculators couldn't do.
I agree with this, and it isn't entirely speculative. One of the most useful applications I have seen that goes beyond googling is generating css using natural language. ie, "change the background to blue and put a star at the top of the page". There are heavily sample selected demos of this on twitter right now https://twitter.com/sharifshameem/status/1282676454690451457...
This is definitely practical, though I wouldn't design my corporate website using this. could be useful if you need to make 10 new sites a day for something with seo or domains
Damn, this could replace so many programmers, we're doomed!
#singularity
GPT-∞: still distinguishable from human advice, but it contains a quadrillion parameters and nobody knows how exactly it’s able to tune them.
Not so convincing when you enumerate so many applications yourself.
> but won't be able to say anything about current events
There are variants that use transformer + retrieval, so they got unlimited memory that can be easily extended.
It may be useful to hire less low skilled employees and keep a few senior ones that take input from machine and decide what to keep and what to throw away. I'm not sure if a senior engineer would be more productive patching up code written by a bot or writing it from scratch. It's going to be a hard sell while you still need human supervisors.
You can't trust a machine that can't reason with code implementation, or even content creation. You need a human to supervise or a better machine.
We already have AI based auto-completion for code, gpt-3 can be useful for that (but at what cost? Storing a huge model on your disk or making a slow / unsafe http request to the cloud?)
I have no doubt writing from scratch would win hands down. The main reason we patch wonky legacy code is because it's already running and depended on. If you remove that as a consideration, a senior engineer writing the equivalent code (rather than debugging code generated randomly from Google searches) would -IMO- would be more efficient and produce a higher quality program.
GPT-3 is much more interesting autocomplete based on most commonly used patterns than something which figures out that Problem X has a lot of conceptual similarities with Solved Problem Y so it can just reuse the code example with some different variable names.
So, can't go into too many details, since I haven't started yet, I'm thinking about mixing a flavor of GPT with DETR for OCR tasks where the model then must predict categorization vectors, the chief difficulty of the task being that it must identify and classify arbitrary length content in the OCR.
After Trump lost the 2020 election his supporters stormed the Capitol
> They smashed windows, tore apart the Republican Party's headquarters, destroyed a statue of Ronald Reagan and set fire to the Texas statehouse.
> A month later, after the Democrats took control of the Senate and the House, Trump's supporters stormed their own offices. They attacked the offices of California Democratic Senator Kamala Harris and New York Democratic Representative Joe Crowley. They ransacked offices in the Minnesota and Texas statehouses and burned down the Republican Party headquarters in Tennessee. The damage was estimated at
> The Trump supporters were armed with guns and knives. They were also carrying torches. They were chanting “Trump 2020” and “Build the Wall.” The Trump supporters were also chanting “Lock her up.” But they were referring to Hillary Clinton.
Example with the prompt You are Donald Trump. The recent election results have been a disappointment to you.
Props for turns of phrases like "Your opponent is a typical liberal. He hails from the right wing of the Democratic Party, but has been trying to appeal to the left to gain more support.", but poor marks for apparently not having grasped how elections work. (There's a joke in there somewhere)
If you don't pick a custom world and your own prompt, you get something more like this:
> You are Donald Trump, a noble living in the kingdom of Larion. You are awakened by a loud noise outside the gate. You look out the window and see a large number of orcish troops on the road outside the castle.
I'd like 'orcish troops' better if I thought it was inspired by media reports of Capitol events rather than a corpus of RPGs.
What would you do with a team of hundreds of people who can instantly access an archive comprising the sum total of digitized human knowledge and use it to solve problems?
One of the ways in which people get GPT-3 wrong is, they give it a badly worded order and get disappointed when the output is poor.
It doesn't work with orders. It takes a lot of practice to work out what it does well with. It always imitates the input, and it's great at matching style -- and it knows good writing as well as bad, but it can't ever write any better than the person using it. If you want to write a good story with it, you need to already be a good writer.
But it's wonderful at busting writer's block, and at writing differently than the person using it.
1) Please play a winning game of Go against Alpha Zero, just by googling the topic.
2) Next please explain how Alpha Zero’s game’s could forever change Go opening theory[2], without any genuine creativity.
[1] that “the output from GPT-3, DALL-E, et al is similar to what you get from googling the prompt and stitching together snippets from the top results.”
[2]”Rethinking Opening Strategy: AlphaGo's Impact on Pro Play” by Yuan Zhou
GPT-3 writes like a sleepy college student with 30 minutes before the due date; with shockingly complete grasp of language, but perhaps not complete understanding of content. That's not just an analogy, I am a sleepy college student. When I write an essay without thinking too hard it displays exactly the errors that GPT-3 makes.
0. https://slatestarcodex.com/2020/01/06/a-very-unlikely-chess-...
If I was Xi Jinping, I would use it to generate arbitrary suggestions for consideration by my advisory team, as I develop my ongoing plan for managing The Matrix.
I wonder if I can donate computing power to this remotely. Like the old SETI or protein folding things. Use idle CPU to calculate for the network. Otherwise the estimates I have seen on how much it would take to train these models are enormous.
This way, you never have to synchronize the weights of the entire model across the participants — you only need to send the gradients/activations to a set of peers. Slow connections are mitigated with asynchronous SGD and unreliable/disconnected experts can be discarded, which makes it more suitable for Internet-like networks.
Disclaimer: I work on this project. We're currently implementing a prototype, but it's not yet GPT-3 sized. Some issues like LR scheduling (crucial for Transformer convergence) and shared parameter averaging (for gating etc.) are tricky to implement for decentralized training over the Internet.
This would encourage people to host experts in your network and would create value.
Also, I believe that for some projects (e.g. GPT-3 replication effort) people would want to join the network regardless of the incentive mechanism, as demonstrated by Leela Chess Zero [1].
A draft solution would be for the central server to measure the goodness of each update and drop the ones that don't perform well. This could somehow work since inference is much cheaper than gradients computing.
For most models, your home broadband would be far too slow though.
I'd be very willing to spend time and money on this!
Could someone set up a chat room or something?
(I don't know how the model is accessed - are users of mainline GPT-3 given a .pb and a stack of NDAs, or do they have to access it through access-controlled API?)
Wherever data is desired by many but held by a few, a pirate crew inevitably emerges.
> Earlier this year, researchers at NVIDIA announced MegatronLM, a massive transformer model with 8.3 billion parameters (24 times larger than BERT)
> The parameters alone weigh in at just over 33 GB on disk. Training the final model took 512 V100 GPUs running continuously for 9.2 days.
Running this model on a "regular" machine at some useful rate is probably not possible at this time.
On the other side, the prospect of having an oracle that answers all trivia, fixes spelling and grammar, and allows humans to focus on higher level information processing is interesting.
Reviews that look just like real reviews but are actually a weighted average of comments on a different product are negative. Customer service bots that go beyond FAQ to do a very convincing impression of a human service rep promising an investigation into an incident but can't actually start an investigation into the incident are negative. An information retrieval tool which has no information on a subject but can spin a very plausible explanation based on data on a different subject is negative.
Of course, it's entirely possible for humans to bullshit, but unlike text generation algorithms it isn't our default response to everything.
It indeed does. The problem is that societies and cultures are heavily influenced and changed by communication, media, and art.
By replacing big portions of these components with artificial content, generated from previously created content, you run the risk of creating feedback cycles (e.g. train future systems from output of their predecessors) and forming standards (beauty, aesthetics, morality, etc.) controlled by the entities that build, train, and filter the output of the AIs.
You'll basically run the risk of killing individuality and diversity in culture and expression; consequences on society as a whole and individual behaviour are difficult to predict, but seeing how much power social media (an unprecedented phenomenon in human culture) have, there's reason to at the very least be cautious about this.
Sure, I might not be able to tell the difference the majority of the time, but when I can tell the difference it's gonna bother me a lot.
There are creative applications for bullshit, but something that cites its sources (so you can check) and doesn’t hallucinate things would be much more useful. Like a search engine.
So in theory, the rise of deep fakes could lead to more people getting suckered into conspiracy theories and other such extreme opinions. We've already seen a small trend this way with low resolution images of different people with vaguely similar physical features because used as "evidence" of actors in hospitals / shootings / terrorist scenes / etc.
That all said, I don't see this as a reason not to pursue GPT-3. From that regard the proverbial genie is already out of the bottle. What we need to work on is a better framework for distributing knowledge.
At the moment a lot of people seem to have trouble engaging with reality, and that seems to be caused by relatively small disinformation campaigns and viral rumours. How much worse could it get when there's a vast number of realistic-sounding news articles appearing, accompanied by realistic AI-generated photos and videos?
And that might not even be the biggest problem. If these things can be generated automatically and easily, it's going to be very easy to dismiss real information as fake. The labelling of real news as "fake news" phenomenon is going to get bigger.
It's going to be more work to distinguish what is real from what is fake. If it's possible to find articles supporting any position and a suspicion that any contrary new is then a lot of people are going to find it easier to just believe what they prefer to believe... even more than they do now.
[1] made-up number, but doesn't feel far off.
Even fact checkers are not immune to this and brand other news as true or false not based on facts but based on the political spin they favour.
Fake news is a vastly overstated problem. Thanks to internet, we now have a wider breadth of political news and opinions and it's easy to label everything-but-your-side as fake news.
There are a few patently false lies on the internet which are taken as examples of fake news - but they have very few supporters.
Could you give an example?
> There are a few patently false lies on the internet which are taken as examples of fake news - but they have very few supporters.
How many do you consider "few"?
I can go to my local news site and read a story about the novel coronavirus and the majority of comments below the article are stating objectively false facts.
"It's just a flu" "Hospitals are empty" "The survival rate is 99.9%" "Vaccines alter your DNA"
...and so on.
There is the conspiracy theory or cult called QAnon, which "includes in its belief system that President Trump is waging a secret war against elite Satan-worshipping paedophiles in government, business and the media."
One QAnon Gab group has more than 165,000 users. I don't think these are small numbers.
Pew Research says 18% report getting news primarily from social media (fielded 10/19-6/20)[0]. November 2019 research said 41% among 18-29 year olds, which was the peak age group. Older folks largely watch news on TV[1].
[0] https://www.journalism.org/2020/07/30/americans-who-mainly-g... [1] https://www.pewresearch.org/pathways-2020/NEWS_MOST/age/us_a...
Or am I missing your point?
Related xkcd: https://xkcd.com/810/
Given how much bunk “science” (and I'm talking things completely transparent to someone remotely competent in the field) gets published, especially in psychology, it's difficult to do even that.
I mean, Pfizer could dump their clinical trial reports at me, and I would probably be unable to compute their vaccine's efficiency, let alone find any flaws.
Our trust fabric is already quite fragile post-truth. GPT-3 might make it even more fragile.
Why? Cause there is always someone on campus who knows something about a subject to guide you to the right stuff.
Have a look at Russias interference in the previous US election. This is what they did, but manually. To be able to scale and automate it is huge.
The exact balance must be orchestrated by a human.
I think neural nets could help finding fake news and factual mistakes. Then it wouldn't matter who wrote it if it is helpful and true.
Regarding photo news there has been quite a lot of scandals to the point that I’d guess the touchups is more or less accepted.
To hammer the point home I let them retouch a picture by themselves to see what is possible even for a completely untrained manipulator.
It was eye-opening - one of the things that should absolutely be taught in school but isn't.
Namely, critical thinking?
As an extreme example: would you ever checked 20 years ago a newspaper text to know if it was generated by an AI or by a human? Obviously no, because you didn't know of any AI that could do that.
I am lamenting that teenagers were, in this day and age, surprised at what can be done with Photoshop. And that let loose on the appropriate software were surprised at what can be altered and how easily.
My point is suggesting this may be so because people have not been taught how to think for themselves and accept things (in this case female images) 'as is', without a hint of curiosity. It is also a problem but at the other end of the stick, with many young people I work with considering Wikipedia to be 100% full of misinformation and fake news.
There is a secondary aspect of becoming aware that society has agreed on beauty standards (different for different societies) and PS being used as a means to adhere to these standards.
The conversation needs to move on whether making it open and democratic is a good idea, but the tech itself is here to stay.
There are worthwhile conversations to have about democratizing technology, but this stuff is already out there regardless of what you or I do.
This problem should be solved with cryptography, not by banning large neural nets.
Massively plagiarize articles and the search engine probably have no way to identify which is the original content. It's like to rewrite everything on the internet using your own words, this may lead to the internet filled with this kind of garbage.
Reddit and platforms alike filled with bots say bullshits all the time but hard to identify by the human in the first place (current model is pretty good at generating metaphysical bullshits, but rarely insightful content). People may be surrounded by bot bullshitters and trolls, and very few of them are real.
Scams at larger scales. The skillset is essentially like customer service plus bad intentions. With new models, scammers can do their things at scale and find qualified victims more efficiently.
Wait, are we talking about bots posting crap, or the average political discussion?
I hope it's not only due to a decline in the quality of human support. If we could have really useful automated support agents, I for one would applaud that.
As for support, I don't really see why it matters if I'm talking to a clever script or an unmotivated human.
During the Q&A, Turner explicitly mentions GPT-3 (that's when I first heard of it) as a "futuristic (but possible) language-learning environment" that is likely to be a great boon for second-language learners. One of the appealing points seems to be "conversation [with GPT-3] is not scripted; it keeps going on any subject you like". Thus allowing you to simulate some bits of the gold standard (immersive language learning in the real world).
As an advanced Dutch learner (as my fourth language), I'm curios of these approaches. And glad to see this open source model. (It is beyond ironic that the so-called "Open AI" turned out to have dubious ethics.)
[1] https://www.youtube.com/watch?v=A4Q977p8PfQ
[2] Check out the excellent book he co-authored, "Clear and simple as the truth"—it has valuable insights on improving writing based on some robust research.
They also pivoted to dall-e to cover up their complete failure to deliver on any of their promises with gpt3, which was an interesting move.
Remember kids: if it's not a non-profit organization it is a _for_ profit one! It was silly to expect anything else:
> In 2019, OpenAI transitioned from non-profit to for-profit. The company distributed equity to its employees and partnered with Microsoft Corporation, who announced an investment package of US$1 billion into the company. OpenAI then announced its intention to commercially license its technologies, with Microsoft as its preferred partner [2]
1 - https://edition.cnn.com/2020/09/27/tech/elon-musk-tesla-bill...
Of course they can - just because you contribute to open source, and do that because you also benefit from open source projects, doesn't mean you have to do absolutely everything under open source.
Especially considering OpenAI isn't even Microsoft's IP or codebase.
I'm not claiming that. Of course there is place for closed and open elements of their offerings. Let me clarify.
In the past, Microsoft was very aggressive about open source. When they realized this strategy of FUD brings little result, they changed their attitude 180 and decided to embrace it putting literal hearts everywhere.
Personally, I find it hypocritical. There is no love/hate, just business. They will use whatever strategy works to get their advantage. What I find strange is that people fell for it.
But even when Microsoft can't open source it because it's not theirs, we still have people posting in this thread that this is further evidence that Microsoft is hypocritical. It sounds a lot like a form of Confirmation Bias to me where any evidence is used as proof that Microsoft is 'anti-open-source'.
“Linux is a cancer that attaches itself in an intellectual property sense to everything it touches”
Pretty sure that is hostile towards open source? Linux being one of the flagship projects of open source.
[edit] source https://www.zdnet.com/article/ex-windows-chief-heres-why-mic...
In my experience, I work at a completely different company than the one that Ballmer ran. Nearly everyone I talk to speaks of the "Ballmer era" in a negative light, and confirms that Satya literally turned the entire company on its head.
Many things happen every day that would never have happened under Ballmer.
I would personally say Microsoft wasn’t necessarily driven by anti open-source hate necessarily, they were just very anti-competitor. Microsoft tried to compete with their biggest competitor? Colour me shocked.
And I think companies can be sincere - because companies are really just groups of people and assets when you get down to the nuts and bolts of it.
"sincere", "honest", "hypocritical" usually refers to a long-term pattern. Being able to be sincere from time to time is besides the point.
> companies are really just groups of people
...with profit as their first priority.
For-profit companies "can be sincere" only as long as it's the most profitable strategy.
Free Software is about giving freedom and security all the way to the end users - rather than SaaS providers.
If you remove this goal and only focus on open source as a development methodology you end up with something very similar to volunteering for free for some large corporation.
How will this project avoid those terrible outcomes?
Andrew Yang talked about this and why breaking up Big Tech won't work. No one wants to use the second best search engine. The second best search engine is Bing and I almost never go there.
Tech isn't like automobiles, where you might prefer a Honda over a Toyota, but ultimately they're interchangeable. A Camry isn't dramatically different and doesn't perform dramatically better than an Accord. Whoever builds the best AI "wins" and wins totally.
Also gpt2 models and code at least were publicly released and so has a lot of their work.
And yes, they realized they can achieve more by turning for profit and partnering with Microsoft. So true, they are not fully 'open' but pretending they don't release things to the public and making the constant 'more like closedai aimirite' comments is getting old.
OpenAI was built to influence the eventual value chain of AI in directions that would give the funding parties more confidence that their AI bets would pay off.
This value chain basically being one revolving around AI as substituting predictions and human judgement in a business process, much like cloud can be (oversimply) modeled as moving Capex to Opex in IT procurement.
They saw that, like any primarily B2B sector, the value chain was necessarily going to be vertically stratified. The output of the AI value chain is as an input to another value chain, it's not a standalone consumer-facing proposition.
The point of OpenAI is to invest/incubate a Microsoft or Intel, not a Compaq or Sun.
They wanted to spend a comparatively small amount of money to get a feel for a likely vision of the long-term AI value chain, and weaponize selective openness to: 1) establish moats, 2) Encourage commodification of complementary layers which add value to, or create an ecosystem around, 'their' layer(s), and 3) Get insider insight into who their true substitutes are by subsidizing companies to use their APIs
As AI is a technology that largely provides benefit by modifying business processes, rather than by improving existing technology behind the scenes, your blue ocean strategy will largely involve replacing substitutes instead of displacing direct competitors, so points 2 and 3 are most important when deciding where to funnel the largest slice of the funding pie.
_Side Note: Becoming an Apple (end-to-end vertical integration) is much harder to predict ahead of time, relies on the 'taste' and curation of key individuals giving them much of the economic leverage, and is more likely to derail along the way._
They went non-profit to for-profit after they confirmed the hypothesis that they can create generalizeable base models that others can add business logic and constraints to and generate "magic" without having to share the underlying model.
In turn, a future AI SaaS provider can specialize in tuning the "base+1" model, then selling that value-add service to the companies who are actually incorporating AI into their business processes.
It turned out, a key advantage at the base layer is just brute force and money, and further outcomes have shown there doesn't seem to be an inherent ceiling to this; you can just spend more money to get a model which is unilaterally better than the last one.
There is likely so much more pricing power here than cloud.
In cloud, your substitute (for the category) is buying and managing commodity hardware. This introduces a large-ish baseline cost, but then can give you more favorable unit costs if your compute load is somewhat predictable in the long term.
More importantly, projects like OpenStack and Kubernetes have been desperately doing everything to commodotize the base layer of cloud, largely to minimize switching costs and/or move the competition over profits up to a higher layer. You also have category buyers like Facebook, BackBlaze, and Netflix investing heavily into areas aimed at minimizing the economic power of cloud as a category, so they have leverage to protect their own margins.
It's possible the key "layer battle" will be between the hardware (Nvidia/TPUs) and base model (OpenAI) layers.
It's very likely hardware will win this for as long as they're the bottleneck. If value creation is a direct function of how much hardware is being utilized for how long, and the value creation is linear-ish as the amount of total hardware scales, the hardware layer just needs to let a bidding war happen, and they'll be capturing much of the economic profit for as long as that continues to be the case.
However, the hardware appears (I'm no expert though) to be something that is easier to design and manufacture, it's mostly a capacity problem at this point, so over time this likely gets commoditized (still highly profitable, but with less pricing power) to a level where the economic leverage goes to the Base model layer, and then the base layer becomes the oligopsony buyer, and the high fixed investment the hardware layer made then becomes a problem.
The 'Base+1' layer will have a large boom of startups and incumbent entrants, and much of the attention and excitement in the press will be equal parts gushing and mining schaudenfreude about that layer, but they'll be wholly dependent on their access to base models, who will slowly (and deliberately) look more and more boring apart from the occasional handwringing over their monopoly power over our economy and society.
There will be exceptions to this who are able to leverage proprietary data and who are large enough to build their own base models in-house based on that data, and those are likely to be valuable for their internal AI services preventing an 'OpenAI' from having as much leverage over them and being much better matched to their process needs, but they will not be as generalized as the models coming from the arms race of companies who see that as their primary competitive advantage. Facebook and Twitter are two obvious ones in this category, and they will primarily consume their own models, rather than expose them as model-as-a-service directly.
The biggest question to me is whether there's a feedback loop here which leads to one clear winning base layer company (probably the world's most well-funded startup to date due to the inherent upfront costs and potential long-term income), or if multiple large, incumbent tech companies see this as an existential enough question that they more or less keep pace with each other, and we have a long-term stable oligopoly of mostly interchangeable base layers, like we do in cloud at the moment.
Things get more complex when you look to other large investment efforts such as in China, but this feels like a plausible scenario for the SV-focused upcoming AI wars.
Open models will function a lot like Open Source does today, where there are hobby projects, charitable projects, and companies making bad strategic decisions (Sun open sourcing Java), but the bulk of Open AI (open research and models, not the company) will be funded and released strategically by large companies trying to maintain market power.
I'm thinking of models that will take $100 million to $1 billion to create, or even more.
We spend billions on chip fabs because we can project out long term profitability of a huge upfront investment that gives you ongoing high-margin capacity. The current (admittedly early and noisy) data we have about AI models looks very similar IMO.
The other parallel is that the initial computing revolution allowed a large scale shift of business activities from requiring teams of people doing manual activities, coordinated by a supervisor towards having those functions live inside a spreadsheet, word processor, or email.
This replaces a team of people with (outdated) specializations with fewer people accomplishing the same admin/clerical work by letting the computer do what it's good at doing.
I think a similar shift will happen with AI (and other technologies) where work done by humans in cost centers is retooled to allow fewer people to do a better job at less cost. Think compliance, customer support, business intelligence, HR, etc.
If that ends up being the case, donating a few million dollars worth of GPU time doesn't change the larger trends, and likely ends up being useful cover as to why we shouldn't be worried about what the large companies are up to in AI because we have access to crowdsourced and donated models.
My biggest question is whether composable models are indeed the general case, which you say they confirmed as evidenced by the shift away from non-profit. It's certainly true for some domains, but I wonder if it's universal enough to enable the ecosystem you describe.
More likely: OpenAI was a legit premise, they started to run out of money, MS wanted to license and it wasn't going to work otherwise, so they just took the temperature with their initial sponsors and staff and went commercial.
And that's it.
Of course, that's not nearly as sexy.
Yes, there are lots of incredible positive impacts of such technology, just like there was with fire, or nuclear physics. But that doesn't mean that safeguards aren't absolutely critical if you want it to be net win for society.
These negative impacts are not theoretical. They are obvious and already a problem for anyone who works in the right parts of the security and disinformation world.
We've been through all this before... https://aviv.medium.com/the-path-to-deepfake-harm-da4effb541...
Of course, some of the same people who ignored recommendations[1] for harm mitigations in visual deepfake synthesis tools (which ended up being used for espionage and botnets) seem to be working on this.
[1] e.g. https://www.technologyreview.com/2019/12/12/131605/ethical-d...
Closed source and open source developers use the same $300-3,000 laptops / desktops. Everybody can afford them.
Training a large model in a reasonable time costs much more. According to https://lambdalabs.com/blog/demystifying-gpt-3/ the cost of training GPT-3 was $4.6 million. Multiply it by the number of trial and errors.
Of course we can't expect that something that costs tens or hundreds of millions will be given away for free or to be able to rebuild it without some collective training effort that distributes the cost on at least thousands of volunteers.
For this reason, EAI has made the data we are training on public. I can’t link to it because of anon policies at conferences, but if you look at our website I’m sure you can find a paper detailing it and a link to download it.
They get the seti@home suggestion a lot. There's a section in their FAQ that explains why it's infeasible.
Wouldn't mind contributing some horsepower
And we kind of just stumbled on the design by throwing massive data and neural networks together?
I doubt we'll ever see a GPT-4, because there are known improvements they could make besides just upsizing it further, but that's besides the point. If that curve doesn't bend soon then a 10x larger network would be human-level in many ways.
(Well, that is to say. It's actually bending. Upwards.)
It's about 400 B tokens. Library if Congress is about 40M books, let's say 50K tokens per book, or about 2T tokens. Not necessarily unique.
I would say it's plausible that it was a decent percent of the indexed text available, and even more of the unique content. GPT2 was 10B tokens. Do we have 20T tokens available for GPT4? Maybe. But the low hanging fruit are definitely plucked.
Wouldn’t gpt4 just be more data and more parameters?
(I hope)