All GB/s without FLOPS – Nvidia CMP 170HX Review
niconiconi.neocities.org
niconiconi.neocities.org
The next best thing is a 3090 or the like with a broken PCIe power connector or some other minor defect. My 3090 is simply missing the bit that holds the clip of the power connector in, however it’s a snug fit anyway and with the cables crammed into my case as they are, I don’t think it’s going anywhere. I paid $200 less than market for that 3090 as a result. Less than a gram of plastic. $200 off.
Meanwhile, as the article points out, AMD is nowhere near as hostile towards its customer base, and modified Radeon cards can apparently be had for $100 or so (from China). The caveat of course is no CUDA support, so it’s kind of moot.
Cupy on Python is mostly what I use.
Anyway, do you know if this is just for pytorch? Is performance roughly 100% of what you’d get on an “equivalent” Nvidia GPU?
I haven't heard about the "bringing native CUDA to AMD GPUs", sounds really interesting. I did come across a picture of Geohot with a ton of AMD GPUs though, wasn't that enough for him or what?
I think this is it: https://news.ycombinator.com/item?id=36189705
If I’m on a budget, my first choice is probably going to be an older, used Nvidia GPU. Maybe a 1050 or one of the older Quadros. The Quadros will get you the most VRAM per dollar, but the GPU will be rather weak, in the realm of a 1050 perhaps, while the 1050 will have better/longer driver support and be physically smaller, quieter, and use less energy.
If those aren’t an option, I would take a closer look at ROCm and HIP. It seems like AMD is prioritizing support for newer GPUs like the Radeon VII, so a $50 RX480 is probably not a good investment, despite the low price.
You will waste time.
But all I'm doing is warning. The consensus viewpoint is not this so you can listen to HN consensus or you can listen to me.
Can you corroborate your points? None of it really aligns with my experience. Nvidia hardware seems quite popular and effective for raster solutions, accelerated RT, dedicated AI and even low-power handheld gaming. I'm typing this out on a Linux box with an Nvidia GPU right now :P
It's worth noting that Nvidia isn't a saint, sure. They play for keeps, and CUDA is limited to paying customers only. CUDA doesn't have open source alternatives, though. Some things do part of what CUDA does really well (or better), but nobody is making a full-stack replacement. Apple is investing in the Accelerate framework which has almost no industry/datacenter application; AMD is doubling down on OpenBLAS and community support. Intel is half-assing some proprietary frameworks and pushing it into demos for a good look.
It would be great if these incumbent companies would pool their vast resource advantage to write, deliver, test and maintain a cross-platform GPGPU library. But that's a lot to ask, and it's easier to just disrupt the entire market with a single integrated package.
My understanding is the likes of oneAPI is supposed to be enable non nvidia gpus to work on CUDA workloads? Is oneAPI one of your these proprietary frameworks?
To be clear, “hostility” is not the word I would’ve chosen, as it attributes emotions to entities that don’t really experience them. Perhaps it’s more useful to talk about whether the company cares if the customers feel exploited or not; and I don’t think Nvidia does (think this will hurt their sales).
Which... sorry to inject my personal opinion here, but it's not. Software is a finite intellectual product designed by motivated human laborers. The hardware can be a commodity, and the design can be a competitive advantage, but software layer is specifically what people consider "monopolized".
Nvidia is not the only company designing GPGPU hardware, and they're not the only company capable of affording commodity silicon from TSMC. The only high-demand thing they entirely control seems to be CUDA, a software feature other companies are too lazy to reproduce. Maybe it's the rest of the market that's being anticompetitive?
On linux they are simply not true.
So are we talking about Windows? Are we talking games?
I also play a few dozen hours of games a month, some new, some old, some AAA, some indie. All through Steam's Proton with no driver issues whatsoever.
I just bought a 4090 and the desktop experience I get is much worse than what I had with the gpu embeded in the Ryzen 7950x: Wayland doesn't work, in Xorg there is tearing in mpv, alt-tab sometimes breaks in gnome. When I launch memory intensive cuda kernels the whole desktop becomes unresponsive. The drivers spews Xid errors in dmesg and breaks for certain applications, such as embergen.
Still, my main complaint is that moving windows around and so on is not smooth. I forget what the term for it is… the window gets jagged and it’s like parts of it are moving at different speeds.
Since then I am on AMD. I game, I build games, work on GPU related stuff.
I refuse to buy Nvidia until they open source their drivers.
I don't care about windows, I am on linux solely and there, from my experience, AMD is doing an excellent job
In 2011 I was using R600, without any problems. Since then the situation improved steadily, especially when Steam got native support.
so you don't use linux
The ROCM drivers are shit. They somehow manage to get an enormous edifice of effort 99% working, then they bungle their package repository. Repeatedly. AMD have a tremendous ability to shoot themselves in the foot five feet from the end of the race, and the thing is, at this point you have to anticipate it. They have the capability to succeed, but not the temperament.
You do realise llama.cpp works on some and cards right?
1. A small group of transient users were buying cards in bulk preventing the long-term users from getting it.
2. They will dump these cards back in market in few years, creating a glut creating fluctuation in market.
170HX and its limitations make perfect sense when you look from that perspective. Had LLM boom and 170HX wouldn't have happened, NVidia would have been struggling with a saturated market and tanked stock price. So yeah, it sucks that you can't use the perfect piece of hardware but than that was its purpose.This is not necessarily true. As seen with android devices you can force digital signature checking mechanisms by varying voltage levels in order to get the device to completely skip the checks as if they were never there.
https://research.nccgroup.com/2020/10/15/theres-a-hole-in-yo...
I'm sure a similar strategy could be developed here.
Slow down there. Glitching is almost never a practical long term strategy. It can take hours (or even days, depending on the target) to successfully bypass a check just once without other follow on effects. Glitching is useful if you need to bypass some mitigation once, such as to extract cryptographic keys, but it's not something you want to do every time you turn on your PC. Glitching gets substantially less reliable with every passing generation due to scaling (increased density/lower Vth increases the odds that you'll corrupt something else, particularly with EM fault injection) and design complexity (glitching out-of-order cores is a HUGE pain).
Of course, the only thing that'll POST are going to be just other vendor images from the same card model usually. (For different power limits, usually)
So, sure, you can skip the verification checks, but your GPU won’t be stable enough to be useful for anything
Also, as the chip and board here seems like an A100 reference design, using an A100 VBIOS image shouldn't fail any signature checks.
The lack of memory would be the most obvious difference. The A100 has 80GB, this has 8GB.
And I really suspect Nvidia probably some way of explicitly locking a chip to a given product ID, like efuses that the BIOS firmware can check on boot.
More sensible way to stopping that would be writing eeprom with encrypted key burned into the GPU itself but I doubt NVIDIA bothered, money loss for few people willing enough to take their GPU apart to replace a chip is insignificiant.
Am I missing something key either conceptually or by failing to read all the stats closely?
My operation was majority on super old model AMD gpu's (RX470 - 580), so there really isn't much use for them now anyway.
That said, if you have a few big container trucks and want to pick them all up, I can put you in touch with the right people... heh.
They run about 4K on eBay, and have 303TFlops of tensor core perf in sparse Bfloat 16. So 150 Dense. They do have 48GB of memory which is great, but at 768GB/s. Source: https://www.nvidia.com/content/dam/en-zz/Solutions/design-vi...
The 4090s run 1599 for a founders edition or a Gigabyte Windforce V2 (my choice). They have at least 165 Tflops of Bfloat16, 330 TFlops of FP16 with FP16 accumulate, and 660 TFlops if you use sparsity. They also support FP8 at 660 TFlops Dense, 1321 TFlops Sparse(!!).
Unless you need the 2x 25% slower memory, the 4090s are much better choices. You get the same scaling over 2 cards anyway.
There is also the RTX A6000 Ada, which is 8K, and based on the same Ada chip as the 4090, except with 48GB of memory. Lower power and clock speeds result in slightly lower peak TFlops numbers. You really really pay for memory.
At FP16? I think you need much less (~2 48GB cards) to finetune 70B with increasing levels of optimization.
I think you can even do it on a single card with QLORA.