Today's Cheap AI Services Won't Last
vincentschmalbach.com
vincentschmalbach.com
I disagree. Yes, the VC funding will dry up, but hardware and algorithmic advances will decrease running costs by equal amounts.
Lack of a moat will prevent companies recouping past expenses, since any who try will be outcompeted by new market entrants who don't have those past expenses.
I think we'll see a lot of vendor lock in by introducing service complexity. We are already at a point where the hassle of switching from one big API provider to another is not always worth the hassle for a small reduction in price or a tiny improvement on some benchmarks. AI will go the same route as all SaaS products until we have GAI-level stuff running locally on affordable consumer hardware. And there's a lot of money to be made until then.
As with the cloud, prices increase once competition dries out. Even if hardware becomes cheap, the services will cost more because shrewd and greedy boards run most businesses.
2. We have at least one real-world example suggesting profitability is challenging. Microsoft is reportedly still losing money on GitHub Copilot, even after years of operation and many paying customers.
These companies, yes. But new startups will come, spend only millions, and achieve the same thing. No moat.
We're in the dialup days of AI, where capabilities are in the hands of very few companies because hardware and training costs are prohibitively expensive. Sure, the apps we use are heavily subsidized by investment funding and the competition is very fierce. I'll also concede that 99% of the AI startups today will fail. But that doesn't mean only the 1% will be left: new ones will continuously enter the arena, compete for attention, and to do that they'll need to lower their prices. All the while, hardware costs will decrease, and incumbents like NVIDIA will inevitably grow stagnant and others will come to eat their lunch. It's the circle of (business) life.
A Digital Ocean VPS starts at $4 a month. It's not so much that we've lost the cheap part of the cloud, it's that the big cloud providers figured out how to do enterprise sales and capture the part of the market that was already overpriced and wasn't sophisticated enough to optimize itself.
Two caveats: application-specific patents are still possible (and many torpedo patents are undoubtedly en route right now) and this argument might not hold at the middleman level (where wrapper apps might get their margins squeezed out.)
A general LLM model, that is "jack of all and master of none" may remain the go-to choice for the masses. These systems leave the last bits of intelligence to be filled by humans.
5 joules per token[0] * 200 tokens per query * 10 queries per day * 365 days per year = 1.014kWh
Which, going by the US's energy mix[1], means a year of someone's LLM usage would be about 0.4kg of CO2 - the same as a single cup of coffee[2]. Worth double-checking, since these are just napkin calculations and I could have made a large mistake.
If it is correct (within an order of magnitude or two), that seems to me a relatively small amount of energy. Obviously still expensive to provide for free to hundreds of millions of users, as almost anything would be.
[0]: https://arxiv.org/pdf/2310.03003
Also, there's that pesky training that is hugely expensive on electricity.
Oh dont get me wrong - there are good uses of neural networks, but the stuff that pushed so much is not that. An example of good one was nvidias sw (rtx voice) to clear up voice during calls, dunno if it exists or not.
I can't find that figure, or what's being included in it: major Google products like Search and Translate are AI (since at least ~2018, depending on your definition of AI) and have huge numbers of users - I think it'd be hard to deny their utility.
> Its not just about one query [...]
To be clear, I didn't calculate that "one query" was equivalent to a cup of coffee - it was a year's worth of someone's LLM usage.
> Also, there's that pesky training that is hugely expensive on electricity.
I don't believe the coffee comparison took into consideration the one-off costs either, like manufacturing of your coffee machine, but let's consider the amortized emissions of training anyway:
500 tons of CO2[0] over 200,000,000 users[1] means 0.0025kg of CO2 per user, barely shifting from the figure of inference alone. Again, rough calculation - would be lower per-user if taking into account API/3rd-party users, but then higher taking into account more model versions.
Realistically, I don't think someone's CO2 emissions on the order of 0.4kg, whether that's coffee, video streaming/rendering, LLM usage, or so on, should take so much of the focus while we have, for instance, celebrities emitting 8,000,000kg a year[2] from private jets.
[0]: https://foundation.mozilla.org/en/blog/ai-internet-carbon-fo...
[1]: https://a16z.com/how-are-consumers-using-generative-ai/
Can tweak the numbers, but even if you think - say - that the average response length is 2000 tokens (~8000 characters), a year of usage would still just be on the level of a few days' worth of coffee for me.
Unless I've made some major error in the calculations, ~0.4kg just doesn't seem like all that much when in the same timeframe a celebrity will be emitting 8,000,000kg from their private jet.
They also didn't measure power use vs context window size, and used single prompts that would be much smaller than conversations (because in a conversation with an LLM, every prompt entered sends the entire conversation to the LLM again.)
Or the much larger context windows of RAG, or programmers sending entire source code archives with every prompt.
As for the number of interactions, a professional using the tool forty hours a week to support their work could easily have a hundred conversations per week (hundreds of prompts total).
There are enough factors here with enough uncertainty that the result could be off by a lot.
That said, I agree that it's unlikely to be worse than a celebrity with a private jet. I can't take anyone who discusses climate change seriously if they don't demand banning private jets.
GPT-4 uses a MoE architecture, composed of many smaller models - (speculated) total parameter count is not particuarly meaningful when the vast majority of its parameters will be in submodels that aren't active for a given query.
In terms of capability, LLaMA-3 is about on par with GPT-4 (in LMSYS arena[0] the earliest GPT-4-0613 has 1161 ELO, Llama-3-70b-Instruct has 1207 ELO, and the latest GPT-4o-2024-05-13 has 1287 ELO).
You could expect LLaMA-3 to reach the same capability with greater efficiency than the earlier GPT-4 versions, as it has the benefit of time leading to algorithmic advances. However, not only does that advantage now fall on the side of the more recent GPT-4 versions, but also OpenAI's inference efficiency will presumably benefit massively from scale and large amounts of infrastructure-specific engineering - opposed to researchers spinning up a model on GPUs from 2017.
So I still believe 5 Joules is approximately fair (especially since many smaller models still see use). But really, even tweaking it upwards by an order of magnitude, it just doesn't seem to amount to all that much.
> and used single prompts that would be much smaller than conversations
I only took their Joules per token estimation.
> As for the number of interactions, a professional using the tool forty hours a week to support their work could easily have a hundred conversations per week (hundreds of prompts total).
So, lets say, 500 prompts total in a week (consulting the LLM every few minutes for the full duration of their work). That's about 3kg of CO2 per year, likely less than the coffee they consumed in one week, or one drive into work.
Initial calculation was intended as a rough average user. A professional may well use an LLM more often than average, and hopefully get higher than average utility from it.
> That said, I agree that it's unlikely to be worse than a celebrity with a private jet.
"unlikely to be worse" is a bit of an understatement. You could use it for 10 million years and still not catch up to half of Taylor Swift's annual private jet usage.
Could be because there actually exists a 10x more efficient alternative for the light bulb.
The models are literally free, just run them yourself
Even the $20/month would take a few years to recoup the costs