Consortium launched to build the largest open LLM
utu.fi
utu.fi
[0] https://digital-strategy.ec.europa.eu/en/policies/european-a...
The only concrete factual detail is they'll train on euro-focused data
either way - say that they have about 10M A100 hours. if they use half their budget for the biggest model as FB did with Llama (rest on testing / smaller models as they scale), then they'll budget 5M A100 hours for the biggest model, or 3x llama-2's 1.7M A100 hours (2k A100 for 21 days).
However, why wouldn't Llama-3 also be 3x bigger than Llama-2? FB's H100's are coming online; they apparently have 4k H100 in a cluster (https://www.nextplatform.com/2023/09/26/meta-platforms-is-de...), and each H100 is about 3x the speed of an A100 (https://www.mosaicml.com/blog/coreweave-nvidia-h100-part-1). So if they provided the same 'calendar time' of 21 days over a single cluster of 4096 H100, that would be about 6.2M A100 hours for a 70B model.
there is an entirely separate question of, LLM's are functions of compute _and_ data, will they also find enough high quality _and_ kosher tokens in their low resource languages?
E.g. for Norwegian the Norwegian national library and state archives sits on at least a couple of magnitudes more of digitized material in Norwegian than OpenAI appears to have found Norwegian material for GPT3.
How much exactly would depend on willingness to give access to still in-Copyright data, but it'd still be far more even if you only give access to what is currently on their website.
And ChatGPT is well versed enough in Norwegian even without that to be able to convincingly use several regional dialects that are rarely used in writing (I've tested several) so it won't take all that much.
But frankly being able to ask it about the contents of these archives would be amazing because they include a couple of centuries of newspapers and almost every book ever published in Norwegian.
The level of digitization will differ, and Norway was early with that, but most countries have depositary requirements for printed materials, and countries with smaller languages like are often more anal about them...
For example, if GPT-5 (or whatever next gen model) was trained on a staggering amount of high quality Norwegian tokens that talk about medicine, city life, romance, maths, etc... would it be able to apply those concepts & learnings to other languages?
I'm trilingual and have ~100% of transfer of knowledge/reading I learn/read in one language into the others, curious to see if LLMs are like that.
It is just a shiny new toy, they may deem there are better uses for their money
As evidence by this 114 page log of how engineers resolved problems that came up during the training of OPT-175, which lasted for several months. https://github.com/facebookresearch/metaseq/blob/main/projec...
Have you tried it recently? It got a lot better, and I like it more for research tasks. (GPT4 + Bing are alright)
I guess that is easy to solve by throwing money
I'm pretty sure and happy that this will probably turn out as a large consumer windfall. Actually comes right in time to counteract the strong inflationary forces that come with de-globalization and declining efficiency due to increasingly basing the economy on gov spending.
https://www.forbes.com/sites/kenrickcai/2023/04/11/how-alexa...
Most people on HN have no idea about that
https://futurism.com/the-byte/ai-magic-spy-on-you-experts-wa...
And especially OpenAI:
https://time.com/6247678/openai-chatgpt-kenya-workers/
That’s the dirty secret of why ChatGPT 4 is better. But they’ll tell you it has to do with chaining ChatGPT 3’s together, more fine tuning etc.
They go to these poor countries and recruit people to work on training the AI, scan their eyeballs for crypto and whatever else. And then privatize the profits from all this.
But it’s nothing new, capitalists have always exploited labor… and people make their choices.
At least it isn’t ivory or blood diamonds
https://www.reuters.com/technology/kenya-panel-urges-shutdow...
A hidden secret that everyone knows is that good LLMs use copyrighted data like books-3. OpenAI itself cites a dataset from books and does not go into detail in its first papers.
Could we instead use, as corpus, a select body of knowledge primarily based on Western sources such as those texts a student might read as part of an excellent liberal arts education?
Further, would it not be possible to limit the corpus text (and consequently later, queries) to grammatically correct language? That is, must we include all the "errors" of common speech in a corpus, such as occur especially in fiction? Indeed, could we not exclude fictional works altogether, esp. if the model is to be used for scientific work?
I'd be curious to see an LLM trained on "The Great Books" plus a college course plan of textbooks. But again, I am not an LLM modeler.
Anyone experimenting with such "Small Language Models(SLM)"?
But as long as there is low quality data might not that data "pollute" the process (whatever that may be)? i.e., "GIGO".
Is "low quality data" possibly necessary for a properly-working LLM?!
[Aside: why the downvote?]
We didn't create an open-source project for the atomic bomb. Quite the contrary: it was one of the best-kept secrets for years.
Sooooo...
if LLMs are as revolutionary as some claim, shouldn't there be significant effort to keep new advances and developments secret and lead competitors astray? That is, if I knew something that would provide a 10- or 100-fold improvement in LLMs, my tendency would be to profit/benefit from that rather than publish it. And to mislead competitors? Has LLM technology development gone "underground" yet?
And what is the role of the 3-letter agencies and the government in this radical development?
Or is LLM technology all unicorns with rainbow farts where the entire world benefits and the creators get only thanks? Somehow I don't believe so.
It must be nerve-wracking being in charge of huge multiple month training cycles. In the 1980s, we built our own hardware to speed up back propagation learning and recall and runs could still take a day or two. It was always so disappointing when a long training run went bad. I can imagine how tense the responsibility must be to manage these huge training cycles.
Is this the beginning of a LLLM? (Lenin LLM)
/s
Also, the focus on "trustworthy" seems problematic. Beyond the obvious who is defining what can be trusted, trustworthy models are significantly harder to work with.
If they can gather a good multilingual dataset including the smaller European languages, and burn enough money on the compute, they could create a useful model for some specifically European use cases.
https://browse.arxiv.org/pdf/2001.08361v1.pdf
Largeness is a valid goal.
Being on Github / HuggingFace but needing to be on a AWS or Nvidia wait list to get the resources to run it is not great.
In an unlimited energy and chip world I would agree just make em bigger.
I guess going bigger has a greater chance of success in being SOTA than looking at architectures. So I get people don’t want to gamble.
What kind of a world we have devolved into, where technological progress makes us less human, on average.
I am saddened by the changes brought about by AI, be it generative art or LLMs, where, because of it, on a global scale, no one knows what truth is, even in pictures and videos.
And we have the perfect tinderbox of siloed wells on social media, with tools to generate authentic looking content that can be disseminated at essentially 0 cost, automated at 0 effort and be used to indoctrinate people at 0 consequences from governments.
It does not matter if people are caught after the fact, the mind that has been indoctrinated, cannot be reverted with the same ease.
And we are generating a bunch of children now, who will be adults later, who will essentially reject reality, reject the need for empathy because what they see may or may not be the truth.
It used to be that 10 years before, if one saw the pictures of carnage or destruction or dead children, a spark of empathy may be generated, enough to ignite change, no matter how small. In the near future, that may not be the case because that photo or video or content has the probability of being AI generated and that spark of empathy will die without a chance to ignite change.
And now we have billions of dollars being invested by the very governments that will feel the effects of these short sighted decisions to unleash technology seldom understood and have no clue as to its long term effects.