Big tech wants to make AI cost nothing
dublog.net
dublog.net
Big tech wants to make AI cost nothing to end-users maybe, but Google and Microsoft want the cost of hosting AI to make your eyes bleed so you don't compete or trim any profits off their cloud services. As the article points out Facebook does not offer cloud services so its interests in this case align with mom and pop shops that don't want to be dependent on big tech for AI.
But Mistral was way more useful to mom and pop shops when they were trying to eke out performance from self-hostable small models. Microsoft took them out of that game. These enormous models may help out boutique data center companies to compete with big tech's cloud offerings but it's beyond a small dev shop who wants to run a co-pilot on a couple of servers.
Microsoft and Google don't want you to learn that a 7B model can come close to a model 50x-100x its size. We don't know that's even possible, you say? That's right we don't know, but they don't even want you to try and find out if it's possible or not. Such is the threat to their cloud offerings.
If they did Microsoft would have made a much bigger deal of things like their Orca-math model and would have left Mistral well alone.
This is the forest for the trees situation. The correct analogy is Chrome and ad blockers. Google didn't tighten the screws until the bean counters started saying it was starting to bite.
Everyone has a small model to do science now and most open source it because realistically otherwise everyone will just use the leader (gpt4 mini)
The title "Why big companies like Meta want to commoditize open source and open weight models to increase demand for their complimentary services" did not quite have the same ring to it to be quite honest
But only until. Facebook rugpulled even React, a (once in the past) javascript library. I can't wait to see what they will pull out when they become the AI overlord.
Deepseek v2 lite is a damn good model that runs on old hardware already (but its slow).
In 2-3 years we will likely have hardware that runs 70b parameter models with enough speed that you will run it locally.
Only when you have difficult questions will you actually pay.
For example I already use https://cluttr.ai to index my screen shots and it costs me $0.
(I made this tool tho)
https://www.reddit.com/r/selfhosted/comments/t33rx5/just_rel...
Generalizing somewhat but focusing on a single company:
https://www.statista.com/statistics/788540/energy-consumptio...
In 2022, Google consumed 22.3G Watt-Hours of energy.
Total electricity consumption by humanity:
https://www.statista.com/statistics/280704/world-power-consu...
In 2022, it was 26T Watt-Hours.
Now, Google is a single company, and if we extrapolate with some cocktail-napkin math, let's say that similar tech giants put together consume, say, 20x Google? 50x Google? So between 2% and 6% of all human electricity consumption.
I realize that's not broken down for AI, but I'm sure if we do break it down we'll find that's an increasing fraction. In this article:
https://cse.engin.umich.edu/stories/power-hungry-ai-research...
the quoted figure is 2% of US electricity usage.
I'm having trouble finding reliable data quickly, but looks like 35% of energy consumption in the world is electrical.
It's still an increasing fraction as you say, but it seems like a doubling or quadrupling of energy used for AI would probably have much less of an impact than the share of the population using ice-cars changing a few percentage points.
My point is its impossible to evaluate energy usage without considering benefits. For example, heating is one of the world's biggest consumers of energy - what % is due to people not wanting to wear a sweater inside?
The numbers on the Statista graphs are in U.S. notation: 22,000 means 22 thousand. (in some other countries it would mean 22)
So Google consumed 22 TWh. (tera, not giga)
And humanity consumed 26,000 TWh or 26 PWh. (peta)
The ratio remains the same though.
Arguably having the weights and code to execute the model available is the preferred form to modify a LLM.
The only reason folks are only paying a small monthly subscription for gippity is literally because of all the VC money flowing in. Training, running, and scaling this stuff has a huge cost. The extra datacentres, the fresh water, all the air conditioning, energy usage, chips, the exploited unprotected labour, etc. It seems very expensive.
Usually when people selling shovels are giving shovels away for free they're banking on a payoff. Usually a regulatory capture payoff. Or a hedge of some kind?
I'm still trying to wrap my head around this.
Update: .. VC money flowing in and the extremely favourable taxation "exceptions" for tech companies in thirsty economies...
The payoff is different... everyone's banking on someone making a breakthrough regarding AGI/ASI and then being the first one to actually make a mass market product out of it that achieves dominance.
But for that to happen, you need a lot of extremely smart people working for you, and you get these people working for you by giving them something to play around with.
I guess it is Meta's Chrome moment.
It should have Why in the posting title
Is there a reason these articles like to say "nation-state" rather than "countries"? I think the question of whether China is a nation-state is not entirely settled (the state is broader than the nation in its case), but also it seems odd to exclude countries like Belgium that are not nation-states from these kinds of statements.
What they mean, of course, is an extremely well funded- government sponsored, tech agency acting as an aggressive security service.
IE; NSA, GCHQ.
As opposed to a tech agency that is acting inwards to the country.
IE; CIA, MI5 (Security Service).
And opposed to a self-funded hacktivist group.
I would greatly prefer better nomenclature though. Nation-State is almost a non-term, since every nation is a state, effectively. - Normally I have seen "state-sponsored", which denotes the correct meaning.
In computer science we tend to have quite clear names for things (blue team, red team). Maybe someone could come up with something better?
This is not true, especially outside of Europe and the US. Even within Europe, Belgium is not a nation-state for example. Many African countries are not either, with Nigeria being an easy example.
State-sponsored is a much better term, I agree, as the phrases are typically used to describe state activity.
For cloud hosts like Amazon, Google, and Microsoft, the product is their inference hardware (mostly GPUs). The bigger the open model, the better!
For Meta however, its a bit more mysterious. From what I can glean, the stated objective is to ensure that the world standardizes on the same AI platform tooling Meta uses internally (presumably to drive down dev costs?).
The bigger "product" however is content creation itself. Giving users the ability to generate engaging content more easily can keep folks using social media and buying ads.
Unbiased AI is, I believe, an existential threat to the “powers that be” retaining control of the narrative, and must be avoided at all costs.
I think it's obvious that Google went to a ridiculous extreme in the other direction, but there does need to be some amount of work done here. For example, we repeatedly have seen that just changing the name on a resume to something more European sounding can have significant impact on callback rates when applying to a job, and if you trained a model to screen resumes based on your own resume result data, this bias could be picked up by the model. That's the sort of situation these are meant to correct for.
For example, if you ask AI to write a realistic story about an NBA team, and it comes back with a team with stereotypically Asian named players, that would be unrealistic. If it came back with a team with stereotypically Black named players, that would be fine. Does it reflect a real-world pattern? Yes. But not changing the algorithm to generate diverse names isn't inserting bias. It's letting AI reflect the real world, as it exists.
Clear cases like chinese NBA players aren't contested, but ugly social issues with layers of abstraction and contradiction.
Here's a similar example from the DALL-E system prompt: https://simonwillison.net/2023/Oct/26/add-a-walrus/#diversif...
You keep using that word. I don't think it means what you think it means.
But seriously; a "mistake" is usually something that cannot be foreseen by a group of people reasonably talented in the state of the art.
This product release was so far from a "mistake", that it isn't funny. It was spectacularly well tested, found to be operating within design parameters, and was released to great fanfare.
They expressed delight in their product, and actually seemed surprised that there was a backlash by the great benighted unwashed masses of their lessers, who clearly couldn't be expected to understand the elevated insights being produced by their creation!
So: not a "mistake". Institutional Bias, baked into a model. Remember: a system's purpose is what is does, not what you think it is supposed to do.
As someone who works either these models as an engineer, I think it's important to understand that a feature implemented as part of the user-facing UI to a model is irrelevant to the work I do with that model via an API.
This would go a long way to reassuring users of the resultant AI, of the neutrality of the trainer.
It would simply reveal the core beliefs of the trainer. If it becomes evident (for example), that Marxist or Keynesian or MMT (or whatever) texts are given high validity measures, but texts by Hayek or Sowell are given negative validity, one could assume the trainer is a leftist, economically.
What benefit is there to not reveal these facts to the users of the resultant AI, if not to hide the internal bias of the trainer? Yet I am unaware of any large commercial AIs that reveal these training bias indicators...
This just seems like a vague X-Files conspiratorial statement without those details.
It's one of the chief problems of competing on price generally, and particularly so in the case of informational exchange.
I'm relatively confident I'd disagree on at least some of OP's classifications of biased information. I can still agree with their general point all the same.
And in either case, coming up with ways of testing for bias, and eliminating counterfactual biases, in AI outputs and systems, would I sincerely hope be a Good Thing.
(Though in writing that I suddenly have my own set of doubts, we've been fooled before....)
The methods to stay in power tend to evolve, but they match the same patterns throughout history (e.g. Divide and Conquer).
That's it. That's the big conspiracy. Some people like to control others.
I remember when the internet was supposed to be an existential threat to the "powers that be". I'm pretty skeptical of narratives like this because the "powers that be" have a lot of resources to leverage any new technology for their benefit. At best a new technology is gives an asymmetrical advantage to small actors for a short time before everyone else catches on.
Furthermore, unbiased AI isn't likely to be any more usable than the garbage we have today. People care about hallucinations, model latency, token pricing and other practical improvements that can be made. Biases are one of the last things stopping people from using AI for legitimate purposes; the other issues are far too glaring to ignore.
We even have alternative meanings for bias within ML, such as for the bias added before non-linearities in many neural networks.
He obviously means censored LLMs, and I think his view is actually right, although I'm far from sure that these firms are in some kind of scheme to produce LLMs biased in this sense.
Uncensored, tunable LLMs under the full control of their users could scour the internet for propaganda, look for connections between people and organisations and just generally make the work of propagandists who don't have their reader's interests in mind more difficult.
I think we'll end up with that anyway but it's a reasonable fear that there'd be people trying to prevent us from getting there.
> Uncensored, tunable LLMs under the full control of their users could scour the internet for propaganda, look for connections between people and organisations and just generally make the work of propagandists who don't have their reader's interests in mind more difficult.
Even this example, what sources do you trust that is or is not "in the readers best interest", what is propaganda or what is an implicit value in a society, when you tune an LLM does that just mean you're steering it to give answers that you like more?
Creating an unbiased LLM is as much of a fools errand as creating an unbiased news publication
>Even this example, what sources do you trust that is or is not "in the readers best interest", what is propaganda or what is an implicit value in a society, when you tune an LLM does that just mean you're steering it to give answers that you like more?
You tune the model yourself. You tune it to find the things you're looking for and which interest you.
>Creating an unbiased LLM is as much of a fools errand as creating an unbiased news publication
It's what you do before pretraining. You model human-written texts with metadata and context with the intent of actually modeling those texts, rather than excising something which isn't just causing the model to fail to learn other things.
It's like, asking "what's a cake, really". We can argue about lines etc., but everbody knows. An unbiased language model is a reasonable thing to want and it's not complicated to understand what it is.
Can you imagine unbiased courts, as an ideal? Somebody who just doesn't care about anything other than certain things? Just as such a thing can be imagined, so can you imagine someone who doesn't about reality and just wants to understand human texts.
> unbiased courts
You say this as something could ever exist. A court will always have a bias because it is humans with values and morals that make a decision. Think about the classic "would you steal bread to feed your family", or even the trolley problem, or as a very concrete example the recent overturning of Roe v Wade in America (keeping in mind that both sides of that discussion reveal an implicit bias based on your starting set of morals and values). Any question that involves a base set of values and morals will never have an unbiased answer.
But the notion of promoting a viewpoint and distributing it freely is as old as myths and sagas, it's at the heart of propaganda (and is why propagandistic "news" sources are often cheap or free to access, often heavily subsidised elsewhere).
This isn't to say that all subsidised and low-cost information is propaganda, or that all paid-for information isn't. But you should probably squint hard when accessing the freely-available stuff, and perhaps make use of several largely-independent sources in making assessments.