That's my take, at least.
Nemotron is TERRIBLE, and purposefully so. It must be.
They cannot be THAT BAD at training AI models. I don't believe it.
I can see one of Nvidia's biggest fears is the inference hardware becoming commoditised.
What Nvidia has the market cornered on is flexible GPU architectures. Nobody else has stepped up to the plate on that, and it's how Nvidia will butter their bread with robotics and future model training efforts.
are they actually suppressing the western open models?
china doesn't give a fuck either way
imo, they see the weakness emerging at the intersection of all the labs, everybody knew there was no moat, so they're gonna control its direction and basically tell the Jev guys what they want them to work on
They are just trying to grow the pie because they have nobody else competing for slices.
My guess is they want as many frontier models using their chips as possible. The only threat to their business is companies making their own chips which Google does and the others are working toward. The last thing they want is only 3 frontier labs who are all not buying NVIDIA.
Hugging Face - distribution for model that you can run on your local Nvidia Spark
Neoclouds - Nvidia setup a 500 billion investment fund with Wall Street so Nvidia can sell chip and this news about buying a LLM start up. it seem like another customer for Nvidia.
correct me if i'm wrong but i remember i saw an interview with Jensen where he want more company to have their own model and country to have their own LLM model.
Nvidia doesn't make money from the gold rush. Nvidia made money from selling the shovels.
But we’re getting ahead of ourselves
similar to how eventually AWS started making their own chips for data centers, and Apple did that for their hardware, it's not a ridiculous thing to plan for the AI companies to start making their own chips to optimize for their use cases and cut out the middleman for margins.
But, not really at the current technologies. Kimi and GLM are fucking awesome, but I don’t have 3TB of VRAM to run them, and I don’t expect to even when ram prices drop.
So now you’re back to the scaling issue before talking about power and compute distribution.
in two years, yes I expect us to have much better models. Perhaps many times more efficient. But not 100 times.
you aren't truly in control of your compute in the cloud, and if you came to become reliant on ai, the cloud company could simply turn it off for you and you'd be shit outta luck. Unlike a regulated utility like electricity and water, where you cannot be discriminated against based on how you use said utility.
Therefore, having the ability to run local ai is a good thing. Currently, hardware constraints have kept it out of most hobbyist's hand, but i am hoping that will change soon enough (within 2-3 yrs).
The mega corps are looking to regulate ai such that individuals cannot run their private ai themselves, so that they can become more effective a monopoly. That's why it is crucial that ai is not regulated, but companies providing ai.
I think NVIDIA does want small open models running on-prem to explode as a market! Lots of smaller GPU installations for companies who wisely want on-prem inference.
Of course NVIDIA will also keep making a ton of money selling to hyper scalers, but not forever: Chinese chips are getting better, Google, Microsoft, Amazon, etc. designing their own inference chips.
NVIDIA is handling this brilliantly.
No one wants to pay the Nvidia tax
Once an organization is big enough to have AI-platform or AI-infra teams, you stop being hung-up on details like CUDA; resources will be invested to make whatever abstractions the AI researchers and engineers rely on (e.g. Torch) will be made to work, and work well.
In other words their motivations and your and my motivations are different. I have limited cash, and my primary goal is expanding my ability to get more cash.
By contrast, the Google's of this world have too much cash. Their primary goal is finding things to depend the cash on (billions at a time) that can return, or at least potentially return, either mountains of more cash, or a competitive advantage.
So yes, for them, hard problems, the harder the better, are desirable.
But a software abstraction only helps to reduce lock in. It does not unlock efficiencies on itself. For that new and/or hardware hardware might be needed.
Smaller customers are also less able to develop their own hardware and threaten NVidia's business.
They can't ride the hyperscaler gravy train forever; at some point between Google, AMD, and Apple NVIDIA is going to lose its monopoly on serving large customers.
At that point, it would be useful if a few open models existed which were only a couple of months behind the frontier.
But it's very important that the open models never be TOO good, because the AI companies are buying compute on the assumption that their software will add value. If it becomes a commodity business with frontier open models, NVIDIA won't be able to get away with such a crazy markup.
Of course they would. You need hardware to run the model. nVidia sets the floor.
I wonder if Nvidia is also trying to cover themselves against a market crash the bankrupts the AI labs but leaves the tech standing? (Similar to the dotcom crash or the big railroad crash back in the day.) If Anthropic and OpenAI struggle, Nvidia can always sell local inference hardware. But local inference hardware often has a lower utilization, so you need more GPUs for the same number of tokens.
nVidia sells hardware. They want open models everywhere because it encourages people and companies to buy more of their hardware, which will sit idle most of the time, instead of using data center services which are ruthlessly optimized to maximize token throughput per hardware unit.
Acquiring companies that are good at training models is exactly what you would do if you were bad at training models.
Have you seen who composes most of Nvidia's revenue?
The Nemotron models are just a by-product of this strategy.
NV's long-term strategic incentive in funding a semi-open model provider like Reflection.ai is to ensure competitive frontier models which fully leverage the NV proprietary stack (chips, interconnects, servers, CUDA) continue to be widely available and continue to offer performance worth a higher price to the most profitable market segments.
NV's ~75% margins on hardware(!) at >$100B/yr scale are historically unprecedented and still increasing, creating tectonic pressure on NV's largest customers (hyperscalers and frontier labs) to escape the "NV Tax" by gaining access to competitive frontier chips, servers, and/or middleware at lower margins. Why would NV help create semi-open models that threaten their best customers? Because as Jeff Bezos famously said, "Your margin is my opportunity" and that's turning NV's biggest customers into their largest existential threat.
At the moment, NV's moats blocking significant competition are almost unimaginably deep, but on a decadal time-scale, literal trillions of dollars are at stake. That's enough to get people thinking the unthinkable, making NV the biggest target in modern business history. It's to the point that it's almost "Everyone against NVidia" which is forcing all the big companies into playing 3D strategic chess on multiple time horizons at once, simultaneously working with, investing in and hedging against each other. It's a 'co-opetition' (https://en.wikipedia.org/wiki/Coopetition) race where the smaller players are grouping into tactical alliances and uneasy truces while the biggest players are spending billions to 'commoditize their complements' (https://gwern.net/complement) as NV is doing with Reflection.ai. This can create strange bedfellows overnight. I wouldn't be surprised to see some of NV's biggest customers, who compete fiercely against each other, pooling resources with NV's competitors to create a viable alternative to NV. This is the stuff Jensen has nightmares about, waking up in a cold sweat in his black leather pajamas.
NV doesn't need Reflection to be better than the best models or even be profitable. They just need to ensure a viable alternative to frontier lab's proprietary models: A. Remains widely available at low enough cost for all NV's other customers to buy, B. 'Works best on NVidia', and C. Stays close enough in price/perf to prevent any single proprietary model becoming as dominant in models as NV is in hardware. The Chinese semi-open models have been strategically convenient for NV but it'd be foolish to count on the Chinese govt continuing to subsidize them or Chinese models not getting blocked or limited by some governments. If it only costs a few billion, Reflection.ai being "good enough," especially for a semi-open, (near-)free, US-based model, is a cheap strategic hedge against long-term threats to NV's (near-)monopoly, especially when Jensen has trouble finding room to store the mountains of cash NV is piling up. When your margins are ~75%, there's literally no better place to put money except toward extending your dominance.