Would be nice to be able to work with some of the larger open source LLM models.
Would be nice to be able to work with some of the larger open source LLM models.
So labeling it “consumer” doesn’t really mean much. They’ve tried to enforce the distinction with EULAs before, but that doesn’t work well.
Isn't this good? The production at scale effects will kick in, lowering the price and supply will meet demand after some hiccups.
Demand puts upward pressure on prices.
Supply is already maxed out and growing as fast as possible.
No, it's not anywhere near maxed out anymore. The chip shortage has been over for a while; now there is too much supply.
But that's not the issue. The issue is that Nvidia really wants to force these enterprise customers to pay extremely high prices. This is why they are so afraid to give any more VRAM to the consumer chips, because if they were at all suitable for VRAM-heavy workloads, every HPC company would buy a couple $2,000 consumer cards in place of each $15,000 datacenter card. That'd lose Nvidia something like 70% which they would find entirely unacceptable.
Is that at 10-40nm or at cutting edge 3-5nm?
If there is a surplus supply, why aren't they making more H800, there demand is more than supply.
Not sure if they still have surplus.
You didn't account for the "hiccups", which can vary from 5-20 years until competition catches up, longer than the life of many companies. In spherical cows worlds of economics that would be just a hiccup.
"Consumer ${thing}" to me is somewhere =< $3000.
The most VRAM you could have on a consumer GPU today is 48GB, and that'd be on a 4090 or 7900xtx with clamshell VRAM (which increases cost and makes cooling significantly harder due to putting GDDR6 chips on the back of the GPU where there aren't any fans).
To calculate how much VRAM is possible, you just need to divide the GPU's bus width by 32 (or 64 for Samsung's new weird double capacity but double bit width) and multiply that by the largest GDDR capacity currently available (16Gbit).
As for why GPUs don't increase their bus width, there have been GPUs with 512bit busses in the past, but it makes it quite a bit more expensive (more vram chips, more traces to run, might require more/heavier PCB layers) and increases power draw.
Climate is pretty much my #1 concern about the world, but LLM use of energy is really really far down on the list of important actions for climate.
First and foremost are removing roadblocks for deploying existing technologies for clean energy, and speeding up the necessary supporting infrastructure such as transmission and market policies for choosing cheapest possible solutions (over the objections of dinosaur execs that choose last century's solutions). Then the big hard to decarbonize parts of industry like cement and steel, as well as deploying electrolyzers to get ammonia fertilizer production switched over to carbon neutral production rather than from fossil-generated hydrogen.
Reducing energy consumption is important for advancing AI in general, but ultimately all its energy consumption will be from clean energy sources anyway, and the switch that needs to happen is that switch in energy sources. Reducing energy use by 2x or 10x is not good enough, we must change the sources fundamentally.
Very occasionally I get the feeling HN is entering the /. phase
Just ten years back, squeezing ML models onto microcontrollers sounded completely insane, given their tight memory and power constraints. We've seen NN compilers developed, game-changing techniques like quantization, pruning, and graph-level optimizations pruning. This allowed deployment of ML models in microcontrollers with a newly developed framework like TFLite Micro.
Also, speculative basic research isn't comparable to adding more RAM to a graphics card. You can't substitute one for the other.
So if you had a DIMM with 16 of these chips, you would already be on the same bandwidth as HBM. 96 DIMMs and you get 40TB/s memory bandwidth.
40 CUs, 256 bit LPDDR5X, 16 CPU cores. Or so the rumors say.
What you want is a DL or parallel compute card, not a graphics card.
They are far more expensive though because compute doesn't sell to the average consumer like graphics does.
I'm not sure if you're asking if they're used for a compute application or if you're asking if they can be used from the same execution path as 'compute' code.
Unified memory is honestly a really, really good idea.
Supposedly, if you have professional Apple hardware experts hand-optimize a machine-learning model, then Apple's most powerful M3 chips can compete against, say, a typical/unoptimized PyTorch implementation of the same model on a 4090.
Except the M-series chips can currently have up to 8 times the memory of the 4090, and they don't have to worry about shuffling things in and out of VRAM. The GPU can just work with any of the data that the CPU already has.
In other words, the Mac is being given a huge optimization advantage compared to the 4090. And it's roughly competitive.
Even that is no mean feat.
Source: https://appleinsider.com/articles/23/12/13/apple-silicon-m3-...
I can produce a featureful webapp in Python-Django solo because I do not have to worry about optimally filling registers and CPU cache. As you cut deeper to the hardware limitations, you have to be significantly more cognizant of writing to the hardware constraints than solving the business problem.
Taken to the extreme, we have Electron applications consuming multi-GB of RAM, but it does expand the universe of possibilities.
There's a joke "20 years of software progress has completely undone the 20 years of hardware progress."
Essentially the joke is about how people don't optimize software like they used to. There is some tongue in cheek though because there are more platforms and variations that we have to account for these days but there is also truth that there is some laziness. I mean I have a laundry app on my phone that takes 5s to load each time because it redraws the home screen while it is trying to connect to the network and find the machines but won't use the bluetooth or nfc chips that both my phone and the machines have.
I do think there is the case to move fast and break things, but at some point you got to slow down and fix things. There should be cycles. Don't forget that debt has interest and that interest compounds. If you have enough debt you'll be "moving fast" but being fast in molasses isn't fast. It's kinda why a lot of software have just declared bankruptcy and started over (e.g. "rewriting in rust"). We should definitely encourage people to hire a mixture of those that move fast and those that move slow. There's harmony there but I think we've forgotten that.
What are you talking about? Sure you can appear to be more productive if you are reckless and wasteful of resources, and adding more resources means there is more that you can afford to waste.
But for the past decade or so we have been seeing what happens if you teach every software developer that it's okay to waste resources: resources get wasted.
Back when you couldn't afford to waste resources, you'd see extremely skilled optimizations and extremely clever hacks to get advanced software to run on extremely slow and primitive CPUs. But ever since computing power increased by a couple orders of magnitude, what happens? The exact same software implemented by today's developers would more than compensate for the increased computing power by being a couple orders of magnitude less efficient.
Sure it expands the universe of possibilities. I do want more resources, I really do. But "reducing the cognitive load on bookkeeping" is making software get slower faster than hardware gets more powerful. Which is absolute insanity to me.
Do I bemoan that my 2000s era super computer squanders resources to load 20 kb of text? Sure, but it is also amazing the things that are possible because tools are not restricted to the priesthood of assembly programmers.
There is an entire spectrum between "640k ought to be enough for anybody" and "this app needs its own fully independent copy of an entire browser engine just to display a couple screens of text".
About Electron: Operating systems provide their own system-level webviews. Use those. Tauri uses those; Tauri apps are absolutely tiny and use very little RAM, because the system-level webview doesn't have to duplicate the overhead of an entire browser engine for every single app that uses it. It typically only has to duplicate the overhead of a single page.
But that is for apps that even need a webview. For apps that don't, you don't have to drop all the way down to assembly to pick something like WinForms. Even WinUI 3 isn't always terrible, at least compared to webviews.
Hell, I wonder if even Python + wxPython would be more efficient than using a webview. Honestly, none of the software you mentioned is "too wasteful", or at least those aren't what I was referring to with my original comment.
Have you heard about Microsoft Teams, and how it can take up to 10 seconds just to display the splash screen, even on a decently fast computer? I'd like to see you try to tell me that it would have literally killed them to optimize it a bit more. Just think back to, say, the IRC clients from back in the Windows 7 days where you wouldn't even have to blink before they finished opening up.
`while i < someFunction():`
Instead of ```
criteria = someFunction()
while i < criteria:
```
I've seen simple lines like this compound in complexity as someFunction() gets more complex or as the loop increases. There's a billion variations on the idea but I think it is because people aren't thinking about how someFunction() is getting called every iteration. I do think every programmer should have some basic optimization skills just because these types of things would be nonobvious otherwise. Because frankly, this kind of optimization is going to have no real difference in a toy program or even in many isolated test cases but it can easily add meaningful latency to a program. Definitely more egregious examples but just trying to think of something that is extremely low hanging fruit.The real issue comes when the user is doing things more than just your program...
That may certainly change but it doesn't help compute much, today.
They know that if you need a 40GB card, you probably can't use a 24GB card at all because it'd run out of VRAM. They charge such insane prices for those higher capacity cards, because they know that they can make you pay thousands upon thousands for every tiny scrap of extra VRAM, and you will have literally no choice but to pay that price because you need the VRAM.
Demand isn't low, demand is actually so high that they don't care if lowly consumers want VRAM, they already have enterprise customers that need it so badly that they'll pay prices higher than most consumers would ever imagine.
And that's why they're so stingy with VRAM on the consumer cards. They need to be careful not to accidentally make them useful in the datacenter, so they can ensure that their enterprise customers continue to be forced to pay those extortionary prices.
You see, it's all about the money, and it always will be. Capitalism, baby!