Maintaining large-scale AI capacity at Meta
engineering.fb.com
engineering.fb.com
Two uses I can think of are i) text and image content moderation on fb and instagram (won't need as many human reviewers if bots are as/more effective), and ii) chatbots for businesses (businesses could provide their business documentation to a meta LLM which could handle customer inquiries via messenger and whatsapp).
Anything else?
The main place I’ve seen content understanding help is coldstart specially for new items by new creators.
EDIT: Also, it's very easy to creat fake profiles on FB, and other people do it all the time. Meta don't need to do it themselves
Or, virtual conversational humans for their boomer userbase could be a hit.
They want to become the AI backend for the Fortune 500
Although even for async training generally I see dataset just sharded and if worker goes down then shard of data may be loss/skipped not some kind of smarter dynamic file assignment factoring when workers go down. Even basic things like job fails continue from last checkpoint with same dataset state for large epoch is messy when major libraries like tensorflow lack a good dataset checkpointing mechanism.
[1] back-of-the-hand-math: 1.7T * 4 bytes = 6.8 TB; 3-4x that for activation + gradients = 27.2 TB; 27.2TB / (80GB / H100) = 349 H100s; 1.5-2x conservative multiplier accounting for not fully using node resources + memory overhead in the machine = ~500-700 H100s.
truly insane numbers.
And of course, it must have changed substantially with GPT-4-Turbo and GPT-4o. It would make sense if the cost reduction was larger than the price reduction, they probably have a higher profit margin now, and the price reduction has been very significant since GPT-4 release.
Now, with the enormous data center growth for AI purposes, companies don’t even bother pretending that any of this is sustainable.
At best, they might delude themselves into believing that a glorified text autocomplete program will magically solve the world’s problems, including the unsustainability of the machines running the program.
One could make a bearish claim on NVidia, that their revenue/valuation is unsustainable unless the AI industry grows 100x over the next few years.
How do we know that AI is going to be the only thing that GPGPUs are the end game?
Tesla's done very expensive EVs -> home/utility PV/storage -> moderately expensive EVs -> FSD (really ADAS) -> Semi/cyber truck so far.
Besides decreasing costs/ramping/improving what they have so far, my guess is they're they're moving to the FSD taxi/model 2/Optimus next.
I suspect Nvidia will try a similar approach to AI hardware/software.
The only place where it's worse than earlier versions is yellow lights. I'm going to dl/try 12.4 tomorrow.
Being able to move from FSD to optimus to whatever comes next is where Tesla can really shine.
Doing all this at +/-200wh/mile is pretty good too. I remember when (2000-2008?) people claimed it wouldn't be viable to get an EV below 500wh/mile.
Other thoughts:
1. Current revenues might be a bit higher than you calculated. E.g. I’m not sure if copilot and azure OpenAI service are fully included in those revenues, and those might be relevant figures at this scale.
2. The 10-100x growth might actually materialize. Corporates are much slower to adopt and scale a technology than people might expect. As a result, many big potential users are only at the very beginning of adopting AI. (I am assuming there are valuable applications for them to use AI for)
There was another bot that popped up in our firm meant to answer questions about corporate policy. I’m guessing that team did something similar but over our policy docs. It vanished about 3 months of being online. It probably gave a bad answer and someone called it out to legal.
I’m curious if anyone else has seen a real application in an enterprise worthy of the hype.
Interesting thing we learned is that agents tend to log during the calls, not afterwards. We now see (qualitative feedback) that they are less busy logging and therefore have more attention for the client, and (quantitatively) we see the calls are getting shorter and people with the tool are doing more calls per day. We do many millions of calls a year, so it sums out to a good number.
Similarly, we have processes with 100s analysts with very high standards to their outputs. The traditional way is to have QA teams review and provide feedback for a few rounds. We’re introducing AI for the first round(s) of feedback to shorten the cycle time reduces context switches) and have the QA teams focus on final reviews.
But I get your point. RAG knowledge bases for experts are in a hard spot. After a few months of employment, the experts tend to know the general knowledge well. As a result, the RAG-bot mostly gets questions about exceptions and niches, where it doesn’t perform very well, and mistakes might be expensive.
Or are we expecting a few companies to be buying 1 million plus H100s each next year
56% of the American economy is also based on intellectual property so it's also a big claim that the existing status quo will have nothing to say about large tech firms trying to displace that, if it is even good enough to do that.
There are tons of chip companies breathing down it's neck. AMD dropped the ball but Huawei Ascend chips are already making a dent and NVIDIA had to drop prices in China. Rest of the industry will start eating into NVIDIA as the market cap grows.
Seen another way, in this example, Nvidia would only need ~7 Alphabets to sustain current revenue.
The same may not apply to audio2audio models, though, and NVIDIA will probably keep a firm grip on training, too.