OLMo: Accelerating the Science of Language Models [pdf]
allenai.org
allenai.org
...
Weights & Biases logs for our training runs."
That's amazing. I've never seen that before in a paper of this quality. Or, any paper at all.
A pity they didn't release the speed of training, but the software is now there for someone else (not under benchmark embargo) to do that.
Is there any information on how much the computing costs were for renting the clusters?
Is the barrier to entry for a 7B model only a couple $100K?
EDIT: https://news.ycombinator.com/item?id=39223467#39224534
Perhaps only $85K total
1. Consumes a minor amount of electricity (Data centers is only 2% of US electricity use, and currently AI is maybe only 5-10% of that). Its trivial compared to say metal smelting.
2. Consume water for cooling.
That's it, there is 0 direct pollution generated from AI, and even the water use is very minor compared to say farming, and can be improved via more water efficient cooling techs.
The main concern is the scaling speed. As LLMs scale up 10x, 100x, 1000x, those previously very minor electricity costs can quickly become grid impacting in a decade.
Externalities are never a part of capitalist math. Non-trivial consequences can never hurt if one never looks further than their own nose.
I suppose this depends greatly on how you view the utility of LLMs. In a capitalist sense, sure—there's great utility here persuading VCs to part with their coins and jobs to be replaced with correspondingly larger profit margins. But the opportunity cost of not solving major problems most of humanity can agree on seems nearly incalculably large. Not that capitalists give a shit.
Imagine how not obvious the first machines must have seemed at the start of the industrial revolution. You only have to feed a man and he can work, but a machine requires iron, oil, water, fuel, engineers, operators. The up front cost for exploring early digging machines must have been absurd. And im sure some people at the time thought: "Wow we could be spending this money on bread for the poor instead."
Arent you glad we didnt.
You really want to have 10 kids and have 50% or more of them die before 10 years old? You want a world before penecillin and antibiotics? No computers? No travel. Women getting marrried off at 15 immediatly pregnant. Most of the world in absolute poverty. Destroyed by a single bad season. Mass famines, plagues, tribal warfare that sweeps over your village. No clean water and soap. malnutrition.
These are just non problems for huge portions of the planet now.
Maybe we can invest more human hours in speeding up the path to zero emissions and energy abundance, or re-planting deserts, or cleaning up forever chemicals / microplastics, or helping at-risk kids, etc etc.
i hope the bottom 10% rung on a dyson sphere society doesnt just look like hungry homeless people, but on a space station.
Or how to tax labor no greater than capital.
Or view quality education and healthcare for children, and keeping their parents out of survival mode, as a much better investment for everyone, than funding the adventures of overly war happy presidents.
I am enthusiastically agreeing with you. Behavioral changes at the top and bottom of society are most of the problem - not tech.
Being poor like this is not a money problem. Its a behaviour problem.
You cant fix this by giving them money. They just spend it on alcohol and cigarettes.
Having interacted with them I know they arent obviously stupid and they are educated. My wife attended the same strict japanese school. Very high quality compared to an average american school, they made it up through calculus as high schoolers. She still remembers reiman sums 15 years later.
Your current perception of the world isnt quite right. It sounds like youve got this magic fix in your head, but in reality it just wouldnt work. Youre ignoring the thing you profess to actually care about... the people.
(Added emphasis)
I agree with everything you say, lots of irresponsible people and culture. But that isn't the whole story.
The wealthy and asset owners also tilt the economy toward themselves and away from labor and the less wealthy in many ways.
Poor outcomes for young individuals do have strong correlations, with strong causal support, to low income districts with poor health and education resources, poor safety, and poverty level parents. That is a circular problem created by treating the education, health, and safety of children as a "local" issue, instead of what it obviously is, a national issue.
Also, housing is a problem for many working people, while the rich magnify the problem by using the limited availability of real estate as a useful financial instrument to park money, making profitable returns based on exclusivity and productive economic growth elsewhere which increases further investment in land, even if the land is underutilized.
This is due to the perverse incentive of taxation on total land and development value instead of just the land. (Development on land should be encouraged, not taxed. Other developent and property isn't "wealth" taxed. Whereas, the underlying land is limited, so taxing those who make it unavailable for others is a community neutral bargain - and makes the underutilization of land unprofitable.)
This goes on and on ... regulation capture, use of personal loans against personal property give wealthy asset owners liquidity events that fund high lifestyles without any taxes associated with it, taxes on labor that increase beyond tax rates on capital and for corporations, etc.
The rich and asset ownership classes use government policy to actively tilt things there way, on the backs of those who's primary "asset" is their labor value, throughout society.
https://www.washingtonpost.com/graphics/2019/world/climate-e...
I feel like we trying to optimize what we measure. No such measurements happen for other industries. How much does Las Vegas use electricity for the extravagant display of lights, water shows and so on.
It's a pity the Mistral 7B Instruct 0.2 dataset isn't available because I've found that a much higher quality than any of the finetunes around, and I suspect we'll have to rely on the same groups doing finetunes for this.
Of course we don't know how to measure this so respect to them for the benchmark performance.
Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
You don't see a lot of 70B or larger models being released for the same reason; it's expensive.
We should just be grateful for what we're getting right now: basically, people are spending 100s of thousands of dollars on training and giving the results away for free. Hugging face is hosting them for free. ollama is hosting them for free. People are writing free inference engines (eg. llama.cpp) and giving them away.
Don't complain. We've got it pretty damn good right now.
That seems remarkably cheap actually and likely getting cheaper fairly quickly with improvements in training efficiencies I’d imagine.
Claiming it’s fair use to suck up the entire web and pay wall the derived result is absurd argument.
We all created the lifeblood of LLM and we’re entitled to the product.
sounds so nice, yet there are going to be objections, NYT for example doesn't think we all should be entitled to the product
Now, I also think the Google model is proven at this point to be a bad model since the web is 90% ads and SEO dogshit. They strip mined the value, took them a while, but its nearly decimated.
That's completely different to Google a d completely different to anything done before. It's as transformative as a human expert news analyst giving you a new perspective on a story.
Seriously, complaining they haven’t spent enough money or didn’t spend 600k making exactly you the model you wanted is…
Let’s just say, ungracious.
Got some cake for my birthday, but it wasn’t the chocolate deluxe cream cake I wanted.
…just remember, the cake is pretty good, and it’s free. :)
Over time the cost of training models will come down and bigger open models will turn up, eventually.
SD 1.x is a ~1B parameter model, so its interesting that it cost so much more than a 7B LLM.
trying to figure out how thick this layer is