Meta is inviting researchers to pick apart the flaws in its version of GPT-3
technologyreview.com
technologyreview.com
https://ai.facebook.com/blog/democratizing-access-to-large-s...
This is the true secret sauce -- all the tricks on how to get these things to train properly that aren't really published.
> CSP fat fingered and deleted our entire cluster when trying to replenish our buffer nodes
Yikes
But why do they have so many problems to keep this cluster stable? Network failures? Bad GPU's? Bad drivers? Bad software?
Running fixmycloud and going after all those cryptic errors every day seems like a nightmare to me...
It makes sense to think of these things like Formula 1 cars, they are trying to eke out the absolute maximum performance, and reliability suffers because of that.
"Ordinary" cloud is more like a Toyota where you optimize for fuel economy and low maintenance.
100 Pages of raw notes released with the language model OPT-175 - https://news.ycombinator.com/item?id=31260665 - May 2022 (26 comments)
These kinds of large transformers can be relatively reproduceable in results and benchmarks. However, making them converge to the exact same parameter set might not be a reasonable expectation.
Not much, but it also depends on what you mean by "reproducible".
Do you expect a similar internal representation (weights) or a similar behavior?
Some interesting takeaways:
- Meta aren't using any software for scientific logbooks, just prepending a document
- So many hardware/cluster issues.
- Hot-swapping algorithms is common and likely underreported (in this case activation functions and optimization method)
- A well resourced team didn't solve enough issues to fully utilize compute resources until >50% of the total time into the project
I agree hot-swapping is underreported, but there are some good existing reports. In fact, OpenAI Five paper today is probably more valuable for its details on hot-swapping than its details on main model which used LSTM and not transformer.
"This new piece of technology might be dangerous and we don't fully understand it, so we should not poke at it or allow people to study it."
Maybe there's just something about my personality that is deeply at odds with this sentiment, but it's also about the lack of testable predictions coming from people like this. Their position could be taken about literally anything with the same logical justification. It's a political and emotional stance masquerading as a technical or scientific process.
Another more broad prediction: In a decade, the overall influence of language models on our society will be universally seen as a net negative.
My prediction is that in the next 10y we will really struggle to determine between fake-people and real-human. There will be an explosion of fake-identities posting more and more human-like.
But I'm not Nostradamus so I could be very very off here.
I’d say Neil Stephenson has a pretty good take on what this might look like in his recent book: Fall, where in everyone has a “feed” and those that or more savvy/wealthy have better editors (AI, etc) of their feed.
And even if the powers that be manage to get those future AI bots to post stuff that will very much resemble what we now post in here, it is my belief that the uncanny valley will be, in the end, impossible to pass (in fact that's one of the main motifs of many of Asimov's books when it comes to robots).
You know, the PGP web of trust idea may yet take off all these decades later, not because we need a web of trust to send 100.0000000% safely encrypted messages to each other to protect from governments, but because we need a web of trust just to know who the real humans are.
[1]: https://forum.agoraroad.com/index.php?threads/dead-internet-...
There would be some of the usual web of trust problems, e.g., trying to explain to Joe Q. Public that you only sign for people you know, beyond doubt, are human. Preferably in person. Many other problems, too.
I guess you could say, my thought here isn't that this would solve the problems. The problems at this point are somewhat well known. What has been missing is any benefit significant enough to motivate us to get past those problems. If there is, it's obviously many years away. Wouldn't exactly suggest building a startup around this idea right now, if you get my drift. We still need to go through a phase of the problem getting larger before we even get to the phase where people start to realize this is a problem and start demanding that people online prove they are actually people, and goodness knows "I'm a human" is merely the lowest of low bars itself, not the solution to all trust problems.
A lot of the outputs look incredibly genuine. We live in interesting times.
GPT type models are much better suited to low effort blogspam, and whilst that's not a good thing, they produce better blogspam than existing blogspamming techniques. I think we underestimate how bad the Internet already is, and at worst text generated by AI is simply going to reflect that.
Been there (as an independent press member years ago), simply you cannot beat that.
Real human journalists have a delay of about 1m before making a short tweet. Funny (or not), something similar was in the "live update" article page in less than 10s. Including photo(s). I was on quite a lot of tech-conferences/live events and earned a decent living then as an independent tech journalist (but then I got bored and really it was a 1-man-show).
Another personal observation (from field), that was not happening prior to 2010-2012, the years we all got Siri, Cortana..
You can make the dots and dashes.
I've been hitting these sites in searches accidentally more and more over the past few months. Goodness help you if you don't realize it's totally fake; some of what I've seen is dangerous, like, bad electrical advice being blithely generated by whichever exact transformer variant is spewing that stuff.
The actually bad consequence is SEO spam of high quality. You can now generate a hundred articles a minute in any topic.
Language models can now pass as human in many situations but there are already billions of humans capable of writing fake news, this isn't a new capability.
We have already created mechanisms for deciding which voices to trust and no matter how good language models get they will not be able to prevent you from visiting economist.com
You have to pay them, and most of them are not very good at writing. Even with a big budget, you get a limited number of good articles per day.
If you can make writing fake news 100x cheaper, and then just throw everything at social networks and let people sort out the most viral stuff, that can change the game.
Also, computers can be faster. If something new happens today and hundred articles are written about it, a computer can quickly process them and generate hundred more articles on the same topic, than a group of humans would. (Many humans can do the writing in parallel, but each of them needs to read individually the things they want to react to.)
The value of exercising critical thinking, checking trusted curated sources, information hygiene, recognizing and avoiding manipulation tactics in news and ads will have to go up. The internet is already almost entirely unreliable without requiring any fancy AI. The listed skills will be necessary regardless any increase of politically manipulative content, advertisements or product misrepresentations.
I don't think people, organizations, and government, pushing false narratives, is some new game. I think it's a game that people are dangerously unaware that they're already playing. Destroying the trust of content on the internet, resulting in having people be more diligent about what they believe, is, I think, almost certainly a net positive.
But, as a counter argument to myself, people are lazy and will, instead, just go to a news source that they "trust", and listen without ever questioning.
Either way, I don't see the fake news, itself, as being anything but a fleeting problem for society. The problem will continue to be people, organizations, and governments taking advantage of laziness.
For an example, imagine Dall-E 2 was released for political use. For a few weeks, we would have a flood of fakes of every politician doing every imaginable act. Society would very quickly believe none, rather than believing all.
That didn't seem to be how it worked with the election-theft narrative. People kept believing garbage from discredited sources.
GPT-3 can't create a second IEEE so it's not an issue to be worried about.
Authoritarians benefit from a society-wide lack of trust in there being any sort of consensus/objective truth. When people don't know what to believe, they either turn to conspiracy theories that claim to offer a peek behind the curtain, or more likely just tune out and hope that their favorite strongman can simplify the chaos.
There's been speculation on Twitter that the conflicting narratives from Russia itself during the invasion aren't really a political problem, since confusion makes it more difficult for the public to rally behind a counter-narrative.
There already exists far more content than anybody can read. We have developed mechanisms for filtering out content which isn't worth our time, and language models don't have any special ability to force you to read their creations.
And even if content can be placed somewhere you will read it: no matter what I put into this box you will not suddenly be susceptible to believing the unbelievable no matter whether I wrote it or whether a language model wrote it.
The issue many people have with fake news is that it's a tool that can sway public opinion without any basis on facts. I'm not sure, by your response, if you find that to be problematic or not.
I think we've recently found that people haven't decided which voices to trust and can be led to believe things placed in front of them. Paired with the ability to spread that information - there is significant impact on society.
That's the reason some people have issues with fake news, from my experience.
Also, getting a computer to do something will always scale several orders of magnitude more than having billions of people do it.
> That's the reason some people have issues with fake news, from my experience.
From my experience, most people who believe Fake News is a significant problem are people who dislike WHOM some have chosen to trust.
Few see fake news coming from their preferred media sources, or supporting their preferred narrative, as highly concerning, and instead usually treat it as benign (human errors happen etc etc).
However, when it's coming from their political adversaries, or from sources they dislike, it suddenly is presented as some huge issue.
In reality, it is all the same, and people are decently good at filtering it out when it contradicts what they want to believe, and very bad at filtering it out when it agrees with them.
For an example of fake news from outlets like the NYT, the most egregious recent one has been the dismissal of the Hunter Biden laptop, calling it a Russian hoax, calling it "fake news", listing intelligence agencies (of all people) as proof of this, getting Twitter and Facebook to outright ban the New York Post article on it - when in fact everything in their article was 100% true.
How many people (especially Dem voters) have taken this story as "dangerous spread of Fake News by the NYT, Washington Post etc"?
My point was that we all have blindspots when the news sources we have chosen to trust are feeding us false information. Also, the problem isn't that people don't choose a news source to trust, it is that they choose a bad news source to trust, and this can happen for very different reasons.
Because of large language models, detecting Fake News becomes trivial and cheap. Building and doing inference on language models is too expensive for most attackers, so they give up. Only well financed state actors are capable of disseminating fake news, and they are no better at it than they are today because content generation is not a bottleneck.
That's not at all clear. You're assuming people will continue to give credence to random stuff they read. But once fake AI-generated content is common, people will surely become less trusting. The end result could easily be that fewer people than before believe fake news is real. Presumably, fewer people will believe real news too, but the result could still be net positive.
But still, you have to assume that the cost of creating fake news is the primary limit to it's appearance in front of people to claim AI will seriously have an impact and that's not at all obvious.
News outlets are nothing without a track record. People trust names they recognize. You can spin up as many CMS instances, domains, and social media profiles for fake news as you want, but without a history shared with its core audience, all the language models in the world aren't going to convince anyone but the most credulous when the content is coming from unfamiliar sources.
I don't know about the next generations(s), the only thing in common for the above mentions was that they were made by a real-human.
Once the account has gained enough real followers, they can carefully start to push payload content, the content the operator really cares about.
This drives polarisation. There's this idea that sinister foreign adversaries are trying to "spread chaos", but I don't buy it. I think it's merely a by-product of building an audience. Russia (to pick an example) doesn't care about BLM, anti-BLM, US culture wars or the Assange case. Rather, it cares about exactly the things you'd expect it to care about: sanctions, conflicts it is involved in, allies etc.
(For that matter, ad-driven media does much the same. Gawker before its demise had really gone all-in on "outrage bait" stories.)
They're going to start with the most credulous, maybe, but they'll build confidence in the same way as everyone else who's an unknown at start.
Sure there is: An inevitability.
I am hoping we will increasingly turn attention towards how to handle it (although I am relatively certain that's already going on at rapidly growing scale at fb, google and openai).
I could see updated legislation doing a lot of heavy lifting – strict rules against automated mass-disinformation – but more action is certainly going to be required.
This is like saying "you should classify this as blue" and then a response like "no, it's heavy".
Search for most any topic on google, e.g. recipes, and the first two pages will be chock-full of AI-generated copy.
Probably, we will be in the same situation relatively soon: And there is little reason to expect the AI systems to have the same pity
Sorry I can't set up a double-blind, testable, peer-reviewed study to help convince you of this
(And then we know one solution to the Fermi paradox.)
Extremists exist at both ends of the spectrum and serve to balance each other out: without people positing the worst-case scenarios, the people positing the best-case scenarios would run full steam ahead without any consideration for what could happen.
Perhaps if the proponents of (various flavours of) AI were doing careful experimentation and iteratively working towards a better understanding, then maybe the loud voices against it would be less valuable, but as we’ve seen through the last 20 years, progress in technology is being made without a second thought for the consequences — and what Facebook are doing here is a bare minimum, so it’s reasonable for proponents to be somewhat cynical about the long term consequence.
There are lots of very serious people seriously looking at these issues and to dismiss them as simple luddites is frankly insulting.
For example Timnit's "parrots" paper confused training with inference and GPUs with TPUs, making specific quantitative estimates that were off by orders of magnitude. If she had talked to a single person working on large language models, she would have recognized the error. But these people work in a bubble where facts don't matter and identify politics is everything.
I assume this is just speculation on your part? Do you have any reason to make that claim? I personally know multiple people doing this full time at large tech companies.
I can give you some examples of serious safety oriented criticism of large language models to contrast with what plays out in the press and amongst "ethicists".
It's well understood that one can generate so-called "adversarial examples" for image classifiers. These adversarial examples can be chosen so that to a human they look like thing A, but the model classifies them as thing B with high probability. Methods of finding these adversaries are well understood. Methods of preventing them from being problematic are less developed but rapidly advancing.
For language models, the situation is much worse. I don't know of any effective way to prevent a large language model from being searched for adversarial inputs, trivially. That is, an attacker could find inputs from large curated input spaces that cause the model to output a specific, desired sequence. For example, an attacker with access to the model weights could probably find an innocuous looking input that causes the model to output "kill yourself".
Is this a risk that AI researchers are aware of? Yes, of course. But the difference between AI researchers and "ethicists" is that AI researchers understand the implications of the risk and will work on mitigations. "Ethicists" do not care about mitigating risk, and they don't care that the people who build the models already understand them and are comfortable with them.
To clarify I think the poster above was talking about the AI Alignment/Control Problem and not the specifics failure modes of particular models, LLM, CNNs etc. Very few people at OpenAI or Deepmind for example are seriously engaging with Alignment. Paul Cristiano at least acknowledges the problem but seems to think there will be available solutions in time to avert serious consequences which may or may not be the case. The folks at MIRI certainly don't seem optimistic.
The failure mode of internal "ethical" control at private enterprises is well-known and has already played out (at least) once when we tried to regulate medical experiments in the 2 decades after WW2. I personally consider the current AI safety positions to be just blatant whitewashing. The lemoine fiasco is a specifically hilarious case in point combining both a) a person that is utterly incompetent and biased to work at that position and b) total failure of leadership to adequately engage with an issue (or even admit it's possible in principle). At the current point, AI safety is roughly as useful as tobacco lobbying (exaggerated for effect).
Some people have figured this out and built careers on it. This wouldn't be a problem, except that this opposition eventually becomes their professional identity - they derive prestige from being the person who is fighting against whatever. So even after researchers address their concerns, they have to ignore the progress or move the goalposts so they can keep on opposing it.
I'm old enough to remember the naive optimism around the internet in the 2000s. "The long tail", "cognitive surplus", "one laptop per child", Creative Commons, the Arab "Spring", breathless Youtube videos about how social media is gonna revolutionize society for the better, etc. Hardly anyone forecasted clickbait, Trump tweets, revenge porn, crypto scams, or social media shaming. If we had a few professional critics who were incentivized to pour cold water on the whole deal, or at least scan the horizon for potential problems, maybe things would've turned out better.
Unfortunately, that difference can only be maintained through some kind of gatekeeping.
I like to try out the new algorithms, but I'm mostly just playing, and I don't see how they make it available to me without letting any random troll use it.
Imagine assigning every single living human being a dedicated 24/7 con artist to follow them around and convince them of something. That's what will soon be possible if not already. It will be intimate con artistry at scale driven by big data, a massive DDOS attack on human cognition and our ability to conduct any form of honest discourse.
What hustlers, trolls, and completely amoral companies will do is bad enough. Now throw in state sponsored intelligence agencies, propaganda farms, militaries, special interest groups, and political parties.
Usually I'm anything but a luddite, but with this I can't help but think of many more evil uses than good ones. It doesn't help that the principal business models of the (consumer) Internet seem to center around surveillance, advertising, propaganda, and addictive forms of entertainment (like slot-machine-like mobile games) designed to suck money or time out of people.
Lesser but also very bad concerns include: the end of useful search engines due to a deluge of continuously learning adversarial SEO spam, the collapse of pretty much any open online forum due to same, and addictive "virtual friend" / "virtual relationship partner" hyper-sophisticated chatbots that hook vulnerable lonely people and then empty their bank accounts in various ways.
I really don't fear AI itself. I fear what human beings will do with AI.
This is an unfounded fear. For one thing, if the value in doing this is high then it's already cheap enough to be practical. The Chinese govt can pay millions of people to do this to dozens of people each. They basically do this already, for specific topics and issues. LLMs won't significantly move the needle here.
Second, are you proposing that attempts to stop Facebook from releasing models will somehow slow down or stop the Chinese, US, or Russian governments? What's the goal, to buy us 6 months? I would much rather the technology be out in the open for everyone to research and understand vs accessible only to state actors or huge tech companies.
I'm not necessarily arguing for intervention to stop this release or something like that. The cat is out of the bag. There's no stopping it. This is going to happen, so get ready for it.
Oh, and throw in deepfakes. You'll have automatic con artistry at scale that can incorporate personalized fake audio and video on demand depicting any supporting detail it needs. It'll be like assigning each person a con artist who's also supported by a staff of content producers.
This would be a continuation of what's happened to the Internet in the last 10-15 years. The Internet is amazing and has tons of incredibly positive uses but all the money is in mass surveillance, addictive "engagement maximizing" stuff, and gambling and scams.
This sounds like traditional Christian teaching of the role of demons.
> It's a political and emotional stance masquerading as a technical or scientific process.
I don't think you understand what ethics is.
If you want guidance on being good, you need a saint, not an ethicist.
They are asking to be regulated because they have finished writing their models.
With regulation it will be harder for up and coming models to gain traction.
Its getting so much coverage because its paid press, I read about it in my newspaper BEFORE tech YouTube and here.
In modern terms, luddite wanted a tax on automation to support loss from tradesworkers who were too old to retrain.
It's similar to when bankers wanted a pension buyout at the introduction of the atm.
Luddites were pro technology, not anti. They just wanted wage distribution of the productivity gains.
They're basically early unions.
I'm with you on being dispositionally pro-tech ad anti-luddite but....
>> lack of testable predictions coming from people like this
I think this is a disingenuous line of argument, more often than not. Popperian science is great, but it is not everything. The majority of our opinions, knowledge and conclusions are not based on testable hypotheses and falsifiable statements.
Take the sentence "One person should not have absolute power." It's not a falsifiable statement, or scientific in other ways. It's based on anecdote and folk wisdom, not science.
I agree with you. And, I think we need to argue back. But, don't argue meta. Meet the arguments head on.
Here's why this matters to me, an independent researcher who wants to start publishing again.
In 2008, Ronan Collobert and Jason Weston had published some work that made neural network training of word vector representations really fast. But only ML people read that paper. Yoshua Bengio and Lev-Arie Ratinov and I plugged old-school cluster based as well as fancy-but-uncool-and-icky neural network word representations into a variety of NLP models. It worked awesome. Before "transformers go brrrrrr" our paper told the NLP community, basically, self-supervised learning and neural networks go "brrrrrrr". People finally started paying attention in the language world, ML stopped being treated with suspicion and the field moved rapidly, our paper racked up 2700 cites and an ACL 10 Year "Test Of Time" award, and here we are.
I don't work in a big research lab but I still publish. I pay for my GPUs the old fashioned way. You know, out of pocket.
It took me ages to get access to GPT-3. Ilya was a colleague of mine, so I messaged him fb, but no dice. Why? I know I could pull other strings through my network but, like, really? Is this where we are right now?
All I'm saying is: It's nice to fill out a form asking my intended use and my previously related publications, as a means of gatekeeping. The access process feels more transparent and principled. Or maybe I'm just being grouchy.
Word Representations: A Simple and General Method for Semi-Supervised Learning
https://aclanthology.org/P10-1040/
(PDF link)
[1]: https://aclanthology.org/P10-1040
I clearly remember Christopher Manning explicitly mentioning your paper at the first ever (to the best of my knowledge) deep learning tutorial at a natural language processing conference (EMNLP 2012), at the very end before we left the room, with something along the lines of: “If there is only one thing you take away from this tutorial: Read Turian’s paper and add word representations as features to your classifiers for whatever task you may have – it works!”
For those less familiar with the history of things, this was prior to Mikolov’s word2vec that arrived in 2013, which apart from popularising neural vector representations (“king - man + woman ~= queen”) above all made fast training of these embeddings feasible – Wikipedia overnight really, compared to the days, weeks, or even months of training for both the cluster-based and neural-based alternatives that were used back then.
Outstanding work and very much ahead of its time. If you are ever around London, my group runs one of the largest public natural language processing talk series in the UK [2] and we would be more than happy to host you if you would be up for example to give a talk on a retrospective on the early days of things up until now.
[2]: https://www.meetup.com/UCL-Natural-Language-Processing-Meetu...
https://www.wpafb.af.mil/News/Article-Display/Article/201113...
Their definition of "open source" turned out to be: "the government owns the source IP instead of some defense contractor. No, you can't see it."
In fairness, I'm impressed that they even got that far. How do you think the defense contractor lobbyists responded to that program?
If I had to bet, just as they pivoted from non-profit to capped-profit, they'll pivot from capped-profit to regular-for-profit when the time comes. Not that I have something against for-profits, but it all (name & mission) feels hypocritical and plain branding/marketing.
0: https://openai.com/blog/openai-lp/
> Returns for our first round of investors are capped at 100x their investment (commensurate with the risks in front of us), and we expect this multiple to be lower for future rounds as we make further progress.
* Yandex released everything as full open
* Facebook released open with restrictions
* OpenAI is completely non-transparent, and to add insult to injury, is trying to sell my own code back to me.
It seems like OpenAI has outlived its founding purpose, and is now a get-rich-quick scheme.
What I really want is a way to run these on a normal GPU, not one with 200GB of RAM. I'm okay with sloooow execution.
All I had to do was prepare the weights in the format Accelerate understands, then load the model with Accelerate. After that, all the rest of the model code worked without any changes.
But it is incredibly slow. A 20 billion parameter model took about a half hour to respond to a prompt and generate 100 tokens. A 175 billion parameter model like Facebook's would probably take hours.
The bad guys get it anyway so this gives the good guys a chance.
I agree with you. Ethics doesn't demand that existing private tech be made available. Who's saying that??
OpenAI is just catching shade because their initial founding mission was to democratize access to AI tech and they've gone pretty far the other way.
It is like: you can not talk to your kids about drugs and pretend they don't exist ... or you can.
Maybe it's my personality but I get the impression since AI is rather limited in 2022 that all the paid AI ethicists spending 90% of the time on bullshit problems because there aren't many real threats. And these gets amplified because the news is always looking for a FUD angle with every AI story.
The priority seems to be protecting random peoples feelings from hypothetical scenarios they invent, when IRL they are releasing research tools on a long-term R&D timeline... GPT-3 isn't a consumer product they are releasing. It's a baby step on a long road to something way bigger. Crippling that progress because of some hyper-sensitivty to people who get offended easily seems ridiculous to me.
If OpenAI wants to concern itself with the ethics of machine learning, why not develop tools to fight misuse?
I think we’re about due for an AI-ethics winter.
"so Meta's GTP-3 is open?"
"correct"
"and the original is not?"
"correct"
"and the original is made by 'OpenAI'?"
"correct"
"hmm"
They just seem to loop around a lot i don't know why... i think they trained on datasets with duplicate content or lots of repeating characters.
Check out https://text-generator.io which is orders of magnitude cheaper than GPT-3 and still makes creative writing/code autocomplete without too much looping issues that youd see with OPT models.
For that repeating you can dial up the repetition penalty or N to generate more sequences in a single request (and you are only charged by the request not by characters/tokens which helps), often generating N results is much more creative than generating a long result as the generated output gets fed into the input when generating long text and often causes that kind of repetitive looping.
Some tactics to mitigate that repetitiveness are:
* Dont use OPT...
* repetition_penalty/retries/seed
* Generate N results and combine instead of doing one big generate as you dont know how many results you'll really get and they are less creative/likely contain shared info/repetitiveness.
* Creative input prompts
There's probably other creative ways of getting variety by splitting the generation into different calls with higher repetition penalty as you get longer or looping detection etc. Its easy to detect repetition in the structured case like chat but hard when doing longer text generation/creative writing
GPT-2 code is very simple, I think the official release by openai is only about a couple hundred lines of code. The challenge of GPT-3 is the scale, the GPT-3 paper basically says "This is GPT-2, but we made changes in the model so we can run it on dozens of GPUs". The changes mostly don't matter, because if you had a big enough GPU (~1.5TB VRAM) you could just up the hyperparameters of GPT-2 and you'd end up with the same results.
So what's novel about GPT-3 that warrants a new name? The discovery here is that after the models reach a certain size, it's able to do many tasks without any training. You can literally ask it to translate from one language to another, at it'll do a decent job at it, if you give it a few examples it'll work even better.
Now that doesn't mean that GPT-3 is the final model, it's still not good enough for many tasks, for example copilot is based on GPT-3 but it was specifically fine-tuned for the task of auto-completing code.
So yes, if you can figure out how to scale a GPT-2 model to 175B parameters, you have a GPT-3 clone.
No one really thinks open-source sponsorships are charity, do they?
Not everything in tech is a sinister capitalistic plot. Open Source and Open Research are truly one of the best ways to accelerate technology advancements, in particular software technology advancements.
That company (and all the other parasitic social media companies like them) needs to be taxed way more heavily than it is being right now, based on the amount of content the people of every nation is generating for them.
How is this usually done in practice ?
Currently, the largest LLM that is both free and commercially usable (Apache 2.0) is 100B YaLM from Yandex (russian’s copy of Google). However, they did not publish any details on their training data.
Yes they did. It's in the README.
if there are more details, can you please share a link?
I am worried that “other sources” may contain Yandex.news which is a cesspool of anti-West and anti-Ukraine propaganda
If so, not many machines have that much RAM. Makes it hard to "play" with.
You can infer using DeepSpeed.