The Inference Cost of Search Disruption – Large Language Model Cost Analysis
semianalysis.com
semianalysis.com
- Google qps is closer to 100k then 320k [1]
- Not every query has to run on LLM. Probably only 10% would benefit from it
- This means 10,000 queries per second, each needing 5 A100s to run, so 50,000 A100s are sufficient. Cost for that is $500MM, quadruple that to $2B with CPU/RAM/storage/network. That is peanuts for Google.
- Let's say AI unlocks new, never seen before queries ('do more people live in madrid or tel aviv') and those are 10x in volume. So capex is now $20B, still peanuts.
- Latency, not cost, is a bigger issue. This should be addressed soon by H100 and newer chips.
- The main issue for the consumer remains the business model. Who will pay for LLMs (user or the advertiser) and will Google and Microsoft stuff ads into LLM responses?
[1] https://www.internetlivestats.com/one-second/#google-band
Also, the main factor in your analysis, 10%, was just completely made up by you.
I use ChatGPT throughout the day as a widely knowledgeable coworker that periodically eats a bag of mushrooms, and have paid the $20 / month. I’d note this description matches all the most useful coworkers I’ve had, except that they cost a ton more per month.
These NVIDIA A40 GPUs mentioned on the page you’ve linked cost $4000 (Amazon) to $10000 (Dell), deliver similar performance to Intel A770 which cost $350. That’s 1 order of magnitude difference in cost efficiency.
(Someone will probably claim ChatGPT would be less likely than Bard to commit the sort of fabrication spotted in the demo, but it's not. I gave ChatGPT the exact same question but replaced JWST with CERN and it told me antimatter was discovered at CERN when according to CERN's own website [1] it was both predicted and observed decades before CERN existed. I didn't have to fish for this. It was the first result for basically the exact same question and was a highly believable fabrication of the exact same kind, attributing a discovery that had been made earlier. I switched JWST for CERN just in case there has been any special training regarding JWST in light of the publicity - no idea if there has been, of course.)
What new discoveries from the CERN can I tell my 9 year old about?
Here are a few recent discoveries from CERN that you can explain to your 9-year-old:
Higgs Boson: CERN scientists discovered the Higgs Boson particle in 2012, which helps explain how other particles have mass. You can explain it to your child as a particle that helps give everything in the universe its size and weight. Antimatter: CERN scientists study antimatter, which is a type of matter that has the opposite charge of normal matter. This helps us understand more about the basic building blocks of the universe. Dark Matter: CERN scientists are searching for evidence of dark matter, which is a mysterious substance that makes up most of the universe but cannot be seen directly. Scientists believe that dark matter could help explain the behavior of galaxies and other celestial objects. Neutrinos: CERN scientists study neutrinos, which are tiny particles that are produced in the sun and other celestial objects. These particles can travel through the universe and can help us understand more about how the universe works. Remember, it's important to keep the explanations simple and use examples that your child can relate to, like using a ball as a representation of a particle.
What CERN discoveries can I tell my 9yo about?
Here are some CERN discoveries that you can explain to your 9-year-old:
1. Higgs boson particle: This particle was discovered at CERN and helps explain why other particles have mass. Think of it like an invisible glue that holds everything together.
2. Antimatter: CERN is also known for discovering antimatter, a type of matter that has the opposite charge of normal matter. When matter and antimatter meet, they annihilate each other and release a huge amount of energy.
3. The Large Hadron Collider (LHC): This is a huge machine that CERN built to study particles and the fundamental nature of the universe. It's like a giant microscope for particles!
4. The internet: CERN played a major role in the development of the internet, which allows people to communicate and share information all over the world.
5. Dark matter and dark energy: CERN scientists are working to understand these mysterious substances that make up most of the universe. Dark matter and dark energy don't interact with light, so they are invisible to us. But scientists can detect their presence by observing their effects on other objects.
I hope this information helps you explain some of the exciting discoveries that CERN has made!
I had a really weird idea the other day. What if the government invested in computer infra like it does for roads. Give each consumer X compute credits that they can use directly at cloud providers or trade in to services to help offset (or completely pay for) their own use.
I would use it to run my open source infra. I imagine my parents would use it to decrease their Netflix bill (for example).
I think it could unlock interesting non-marketing based “free” services. But it’s a hard concept to pitch to the average person who wouldn’t know why they would want 1000 CPU hours (for example).
I think the answer would be for them to host a competing service AND allow people to use their credits to choose that service or a different vendor.
That number is wrong. I have a number from googler, not livestats which cannot have google internal data
- Not every query has to run on LLM. Probably only 10% would benefit from it
Agree, i have something different coming up that looks into this more, 10% may be too low. I know i used 100% which isn't right, and explicitly say that
- This means 10,000 queries per second, each needing 5 A100s to run, so 50,000 A100s are sufficient. Cost for that is $500MM, quadruple that to $2B with CPU/RAM/storage/network. That is peanuts for Google.
50k A100s networking ramp cost way more than $2B HW utilization rate
- Latency, not cost, is a bigger issue. This should be addressed soon by H100 and newer chips.
Thats discussed in the subscriber section. It's both, but yes latency is bigger issue. H100 helps but doesn't solve.
Capex: $20B
Driving Google's margin to zero and feasting on their remains: priceless
https://www.ft.com/content/2d48d982-80b2-49f3-8a83-f5afef98e...
> “From now on, the [gross margin] of search is going to drop forever,” Nadella said in an interview with the Financial Times.
> “There is such margin in search, which for us is incremental. For Google it’s not, they have to defend it all,” he added, referring to the competition against Google as “asymmetric”.
Google should easily spend the $33B if it keeps them on top while finding ways to make things cheaper, more accurate, smarter.
Huh it literally uses real throughput figures
> doesn’t account for ways to make things cheaper over time
It does in the subscriber section and it says it does say that in the free section.
> We already have papers that suggest most big models are undertrained and smaller models can get the same accuracy.
Why assume 2023 model is the same as 2020 GPT-3 175B parameter
Anyway, this is the reason Google invested in TPUs early on. They may have to start spending more on building out TPU clusters, but at least they don't have to pay Nvidia's profit margins to do it.
-We don't know how costs will look like after intensive optimization
-The results of chatGPT are not trustworthy. If someone solves or greatly improves on that, he'll win.
-Even if it would mean reduction in gross margin, Google has no choice. They'll accept that fact sooner or later.
The mistake here is assuming the business models will stay the same (i.e. free usage supported by ads). I think that's a mistake.
The second mistake is assuming people use search engines to answer questions or generate content. They might. Or they might just be looking for a particular web page or article. The former is kind of a new use case and business model. The latter is what Google is already doing quite well.
If I were Google, I would make this a paid feature with an initially generous freemium layer to lure people in and try before they buy. That would allow them to control the pace of rolling this out and allow them to focus on not having to cripple the experience.
MS is doing this indirectly via office365 and teams, which are already paid features and by partnering with OpenAI. Bing is a relatively minor business for them at this point so they can experiment there without risking much. MS is not going to be able to absorb most of Googles users anytime soon. My guess is that even MS does not believe that is actually going to happen. What might happen is that a lot of early adopters will come along and then find their way to their paid products. From a revenue point of view that seems like a good plan to me.
The issue with Google is that right now they are just responding to what MS is doing. They don't seem to yet have a coherent strategy of their own beyond just coasting along on ad revenue.
"We estimate the cost per query to be 0.36 cents."
"<2000ms latency"
"50% hardware utilization rates"
Assuming all queries take 2s this implies that ( 1 * 3600 / ( 2 / 0.5 ) ) * 0.36 = 324 GPUs would be used for each query. This seems implausibly high.
It's 2000ms per token not for the whole query.
Hardware utilization rates and MFU are not the same thing, you forgot the latter.
You are pretending its perfectly parelelized on 1 GPU too. I use 8x GPU box throughput.
And that puts the estimate entirely in the realm of the possible. Whoops!
Great piece.