Nvidia's Project Digits is a 'personal AI supercomputer'
techcrunch.com
techcrunch.com
Nvidia Jetson Nano, A SBC for "AI" debuted with already aging custom Ubuntu 18.04 and when 18.04 went EOL, Nvidia abandoned it completely without any further updates to its proprietary jet-pack or drivers and without them all of Machine Learning stack like CUDA, Pytorch etc. became useless.
I'll never buy a SBC from Nvidia unless all the SW support is up-streamed to Linux kernel.
In general, Nvidia's relationship with Linux has been... complicated. On the one hand, at least they offer drivers for it. On the other, I have found few more reliable ways to irreparably break a Linux installation than trying to install or upgrade those drivers. They don't seem to prioritize it as a first class citizen, more just tolerate it the bare minimum required to claim it works.
> Nvidia's relationship with Linux has been... complicated.
For those unfamiliar with Linus Torvalds' two-word opinion of Nvidia:But the impression I get from this device is that it's closer in spirit to the Grace Hopper/datacenter designs than it is the Tegra designs, due to both the naming, design (DGX style) and the software (DGX OS?) which goes on their workstation/server designs. They are also UEFI, and in those scenarios, you can (I believe?) use the upstream Linux kernel with the open source nvidia driver using whatever distro you like. In that case, this would be a much more "familiar" machine with a much more ordinary Linux experience. But who knows. Maybe GH200/GB200 need custom patches, too.
Time will tell, but if this is a good GPU paired with a good ARM Cortex design, and it works more like a traditional Linux box than the Jeton series, it may be a great local AI inference machine.
This is more like a micro-DGX then, for $3k.
I can only think of raspberry pi...
Compute is evolving way too rapidly to be setting-and-forgetting anything at the moment.
In 4 years, you'll be able to combine 2 of these to get 256gb unified memory. I expect that to have many uses and still be in a favorable form factor and price.
This isn't the 80s when compute doubled every 9 months, mostly on clock scaling.
Xeon Phi failed for a number of reasons, but one where it didn't need to fail was availability of software optimised for it. Now we have Xeons and EPYCs, and MI300C's with lots of efficient cores, but we could have been writing software tailored for those for 10 years now. Extracting performance from them would be a solved problem at this point. The same applies for Itanium - the very first thing Intel should have made sure it had was good Linux support. They could have it before the first silicon was released. Itaium was well supported for a while, but it's long dead by now.
Similarly, Sun has failed with SPARC, which also didn't have an easy onboarding path after they gave up on workstations. They did some things right: OpenSolaris ensured the OS remained relevant (still is, even if a bit niche), and looking the other way for x86 Solaris helps people to learn and train on it. Oracle cloud could, at least, offer it on cloud instances. Would be nice.
Now we see IBM doing the same - there is no reasonable entry level POWER machine that can compete in performance with a workstation-class x86. There is a small half-rack machine that can be mounted on a deskside case, and that's it. I don't know of any company that's planning to deploy new systems on AIX (much less IBMi, which is also POWER), or even for Linux on POWER, because it's just too easy to build it on other, competing platforms. You can get AIX, IBMi and even IBMz cloud instances from IBM cloud, but it's not easy (and I never found a "from-zero-to-ssh-or-5250-or-3270" tutorial for them). I wonder if it's even possible. You can get Linux on Z instances, but there doesn't seem to be a way to get Linux on POWER. At least not from them (several HPC research labs still offer those).
Sad to see big companies like intel and amd don't understand this but they've never come to terms with the fact that software killed the hardware star
Windows has always been a barrier to hardware feature adoption to Intel. You had to wait 2 to 3 years, sometimes longer, for Windows to get around us providing hardware support.
Any OS optimizations in Windows you had to go through Microsoft. So say you added some instructions custom silicon or whatever to speed up Enterprise databases, provide high-speed networking that needed some special kernel features, etc, there was always Microsoft being in the way.
Not just in the drag the feet communication. Getting the tech people a line problem.
Microsoft will look at every single change. It did as to whether or not it would challenge their Monopoly whether or not it was in their business interest whether or not it kept you as the hardware and a subservient role.
This is a genius move. I am more baffled by the insane form factor that can pack this much power inside a Mac Mini-esque body. For just $6000, two of these can run 400B+ models locally. That is absolutely bonkers. Imagine running ChatGPT on your desktop. You couldn’t dream about this stuff even 1 year ago. What a time to be alive!
That said, enthusiasts do help drive a lot of the improvements to the tech stack so if they start using this, it’ll entrench NVIDIA even more.
Surely a smaller market than gamers or datacenters for sure.
I have a bit of an interest in games too.
If I could get one platform for both, I could justify 2k maybe a bit more.
I can't justify that for just one half: running games on Mac, right now via Linux: no thanks.
And on the PC side, nvidia consumer cards only go to 24gb which is a bit limiting for LLMs, while being very expensive - I only play games every few months.
Maybe (LP)CAMM2 memory will make model usage just cheap enough that I can have a hosting server for it and do my usual midrange gaming GPU thing before then.
I do hope that a AMD Strix Halo ships with 2 LPCAMM2 slots for a total width of 256 bits.
It’s purely an ecosystem play imho. It benefits the kind of people who will go on to make potentially cool things and will stay loyal.
100%
The people who prototype on a 3k workstation will also be the people who decide how to architect for a 3k GPU buildout for model training.
It will be massive for research labs. Most academics have to jump through a lot of hoops to get to play with not just CUDA, but also GPUDirect/RDMA/Infiniband etc. If you get older/donated hardware, you may have a large cluster but not newer features.
Also why aws is giving trainium credits for free
No one goes to an Apple store thinking "I'll get a laptop to do AI inference".
Performance is not amazing (roughly 4060 level, I think?) but in many ways it was the only game in town unless you were willing and able to build a multi-3090/4090 rig.
Since the current MacOS comes built in with small LLMs, that number might be closer to 50% not 0.1%.
If what you say it's true you were among the first 100 people on the planet who were doing this; which btw, further supports my argument on how extremely rare is that use case for Mac users.
Incredible fumble for me personally as an investor
And if you truly did predict that Nvidia would own those markets and those markets would be massive, you could have also bought Amazon, Google or heck even Bitcoin. Anything you touched in tech really would have made you a millionaire really.
Plus, YouTube and the Google images is already full of AI generated slop and people are already tired of it. "AI fatigue" amongst majority of general consumers is a documented thing. Gaming fatigues is not.
Suppose you're a content creator and you need an image of a real person or something copyrighted like a lot of sports logos for your latest YouTube video's thumbnail. That kind of thing.
I'm not getting into how good or bad that is; I'm just saying I think it's a pretty common use case.
Do I buy a Macbook with silly amount of RAM when I only want to mess with images occasionally.
Do I get a big Nvidia card, topping out at 24gb - still small for some LLMs, but I could occasionally play games using it at least.
Titanic - so about to hit an iceberg and sink?
No. There's already too much porn on the internet, and AI porn is cringe and will get old very fast.
How so?
Only 40% of gamers use a PC, a portion of those use AI in any meaningful way, and a fraction of those want to set up a local AI instance.
Then someone releases an uncensored, cloud based AI and takes your market?
Do we need more of those? We need plumbers and people that know how to build houses. We are completely full on founders and executives.
True passion for one's career is rare, despite the clichéd platitudes ecouraging otherwise. That's something we should encourage and invest in regardless of the field.
I mean, this is awfully close to being "Her" in a box, right?
Also, it’s $3000. For that you could buy subscriptions to OpenAI etc and have the dystopian partner everywhere you go.
Also, I don't particularly want my data to be processed by anyone else.
Or efficiency gains in hardware and software catchup making current price point profitable.
We still schedule "bi-weekly" meetings.
We can't agree on which way charge goes in a wire.
Have you seen the y-axis on an economists chart?
Those Macs with unified memory is a threat he is immediately addressing. Jensen is a wartime ceo from the looks of it, he’s not joking.
No wonder AMD is staying out of the high end space, since NVIDIA is going head on with Apple (and AMD is not in the business of competing with Apple).
The fire-breathing 120W Zen 5-powered flagship Ryzen AI Max+ 395 comes packing 16 CPU cores and 32 threads paired with 40 RDNA 3.5 (Radeon 8060S) integrated graphics cores (CUs), but perhaps more importantly, it supports up to 128GB of memory that is shared among the CPU, GPU, and XDNA 2 NPU AI engines. The memory can also be carved up to a distinct pool dedicated to the GPU only, thus delivering an astounding 256 GB/s of memory throughput that unlocks incredible performance in memory capacity-constrained AI workloads (details below). AMD says this delivers groundbreaking capabilities for thin-and-light laptops and mini workstations, particularly in AI workloads. The company also shared plenty of gaming and content creation benchmarks.
[...]
AMD also shared some rather impressive results showing a Llama 70B Nemotron LLM AI model running on both the Ryzen AI Max+ 395 with 128GB of total system RAM (32GB for the CPU, 96GB allocated to the GPU) and a desktop Nvidia GeForce RTX 4090 with 24GB of VRAM (details of the setups in the slide below). AMD says the AI Max+ 395 delivers up to 2.2X the tokens/second performance of the desktop RTX 4090 card, but the company didn’t share time-to-first-token benchmarks.
Perhaps more importantly, AMD claims to do this at an 87% lower TDP than the 450W RTX 4090, with the AI Max+ running at a mere 55W. That implies that systems built on this platform will have exceptional power efficiency metrics in AI workloads.
I think this is a race that Apple doesn't know it's part of. Apple has something that happens to work well for AI, as a side effect of having a nice GPU with lots of fast shared memory. It's not marketed for inference.
They propelled on unexpected LLM boom. But plan 'A' was robotics in which NVidia invested a lot for decades. I think their time is about to come, with Tesla's humanoids for 20-30k and Chinese already selling for $16k.
i think it isn't about enthusiast. To me it looks like Huang/NVDA is pushing further a small revolution using the opening provided by the AI wave - up until now the GPU was add-on to the general computing core onto which that computing core offloaded some computing. With AI that offloaded computing becomes de-facto the main computing and Huang/NVDA is turning tables by making the CPU is just a small add-on on the GPU, with some general computing offloaded to that CPU.
The CPU being located that "close" and with unified memory - that would stimulate development of parallelization for a lot of general computing so that it would be executed on GPU, very fast that way, instead of on the CPU. For example classic of enterprise computing - databases, the SQL ones - a lot, if not, with some work, everything, in these databases can be executed on GPU with a significant performance gain vs. CPU. Why it isn't happening today? Load/unload onto GPU eats into performance, complexity of having only some operations offloaded to GPU is very high in dev effort, etc. Streamlined development on a platform with unified memory will change it. That way Huang/NVDA may pull out rug from under the CPU-first platforms like AMD/INTC and would own both - new AI computing as well as significant share of the classic enterprise one.
No, they can’t. GPU databases are niche products with severe limitations.
GPUs are fast at massively parallel math problems, they anren’t useful for all tasks.
I’m so tired of this recent obsession with the stock market. Now that retail is deeply invested it is tainting everything, like here on a technology forum. I don’t remember people mentioning Apple stock every time Steve Jobs made an announcement in the past decades. Nowadays it seems everyone is invested in Nvidia and just want the stock to go up, and every product announcement is a mean to that end. I really hope we get a crash so that we can get back to a more sane relation with companies and their products.
That's the best time to buy. ;)
I wonder how it would go as a productivity/tinkering/gaming rig? Could a GPU potentially be stacked in the same way an additional Digit can?
About that... Not like there isn't a lot to be desired from the linux drivers: I'm running a K80 and M40 in a workstation at home and the thought of having to ever touch the drivers, now that the system is operational, terrifies me. It is by far the biggest "don't fix it if it ain't broke" thing in my life.
0. https://www.macstadium.com/blog/m4-mac-mini-review
1. https://www.apple.com/mac/compare/?modelList=Mac-mini-M4,Mac...
Apple M chips are pretty efficient.
On the other hand, with a $5000 macbook pro, I can easily load a 70b model and have a "full" macbook pro as a plus. I am not sure I fully understand the value of these cards for someone that want to run personal AI models.
Also I'm unfamiliar with macs is there really a MacBook pro with 256GB of RAM?
Also, macOS devices are not very good inference solutions. They are just believed to be by diehards.
I don't think Digits will perform well either.
If NVIDIA wanted you to have good performance on a budget, it would ship NVLink on the 5090.
And we know why they won't ship NVLink anymore on prosumer GPUs: they control almost the entire segment and why give more away for free? Good for the company and investors, bad for us consumers.
They are good for single batch inference and have very good tok/sec/user. ollama works perfectly in mac.
Did see vague claims of "starting at $3k", max 4TB nvme, and max 128GB ram.
I'd expect AMD Strix Halo (AI Max plus 395) to be reasonably competitive.
[0]: https://newsroom.arm.com/blog/arm-nvidia-project-digits-high...
NVidia works closely with Microsoft to develop their cards, all major features come first in DirectX, before landing on Vulkan and OpenGL as NVidia extensions, and eventually become standard after other vendors follow up with similar extensions.
Here's a link to the part of the keynote where he says this:
Wait, what do you mean exactly? Isn't WSL2 just a VM essentially? Don't you mean it'll run on Linux (which you also can run on WSL2)?
Or will it really only work with WSL2? I was excited as I thought it was just a Linux Workstation, but if WSL2 gets involved/is required somehow, then I need to run the other direction.
?
Yeah starting at $3,000. Surely a cheap desktop computer to buy for someone who just wants to surf the web and send email /s.
There is a reason why it is for "enthusiasts" and not for the general wider consumer or typical PC buyer.
For general desktop use, as you described, nearly any piece of modern hardware, from a RasPI, to most modern smartphones with a dock, could realistically serve most people well.
The thing is, you need to serve both, low-end use cases like browsing, and high-end dev work via workstations, because even for the "average user", there is often one specific program on which they need to rely and which has limited support outside the OS they have grown up with. Course, there will be some programs like Desktop Microsoft Office which will never be ported, but still, Digitis could open the doors to some devs working natively on Linux.
A solid, compact, high-performance, yet low power workstation with a fully supported Linux desktop out of the box could bridge that gap, similar to how I have seen some developers adopt macOS over Linux and Windows since the release of the Studio and Max MacBooks.
Again, we have yet to see independent testing, but I would be surprised if anything of this size, simplicity, efficiency and performance was possible in any hardware configuration currently on the market.
That end of the market is occupied by Chromebooks... AKA a different GNU/Linux.
This isn't competing with cloud, it's competing with Mac Minis and beefy GPUs. And $3000 is a very attractive price point in that market.
I get what you're saying, but there are also regulations (and your own business interest) that expects data redundancy/protection which keeping everything on-site doesnt seem to cover
The owner of the market, Illumina, already ships their own bespoke hardware chips in servers called DRAGEN for faster analysis of thousands of genomes. Their main market for this product is in personalised medicine, as genome sequencing in humans is becoming common.
Other companies like Oxford Nanopore use on-board GPUs to call bases (i.e., from raw electric signal coming off the sequencer to A, T, G, C) but it's not working as well as it could due to size and power constraints. I feel like this could be a huge game changer for someone like ONT, especially with cooler stuff like adaptive sequencing.
Other avenues of bioinformatics, such as most day-to-day analysis software, is still very CPU and RAM heavy.
It is of course possible that these chips enable analyses that are currently not possible/prohibited by cost, but at least for now, this will not be the limiting factor for genomics, but cost of sequencing (which is currently $400-500 per genome)
I've worked in a project some years ago where we were using data from genome sequencing of a bacteria. Every sequenced sample was around 3GB of data and sample size was pretty small with only about 100 samples to study.
I think the real revolution will happen because code generation through LLMs will allow biologists to write 'good enough' code to transform, process and analyze data. Today to do any meaningful work with genome data you need a pretty competent bioinformatician, and they are a rare breed. Removing this bottleneck is what will allow us to move faster in this field.
"DGX OS 6 Features The following are the key features of DGX OS Release 6:
Based on Ubuntu 22.04 with the latest long-term Linux kernel version 5.15 for the recent hardware and security updates and updates to software packages, such as Python and GCC.
Includes the NVIDIA-optimized Linux kernel, which supports GPU Direct Storage (GDS) without additional patches.
Provides access to all NVIDIA GPU driver branches and CUDA toolkit versions.
Uses the Ubuntu OFED by default with the option to install NVIDIA OFED for additional features.
Supports Secure Boot (requires Ubuntu OFED).
Supports DGX H100/H200."
It’s obviously not guaranteed to go this route, but an LLM (or similar) on every desk and in every home is a plausible vision of the future.
It's a garden hermit. Imagine a future where everyone has one of those(not exactly this version but some future version), it lives with you it learns with you and unlike the cloud based SaaS AI you can teach it things immediately and diverge from the average to your advantage.
Maybe it will still make sense to have your personal AI in some data center, but on the other hand, there is the trend of governments and mega corps regulating what you can do with your computer. Try going out of the basics, try to do something fun and edge case - it is very likely that your general availability AI will refuse to help you.
when it is your own property, you get the chance to overcome restrictions and develop the thing beyond the average.
As a result, having something that can do things that no other else can do and not having restrictions on what you can do with this thing can become the ultimate superpower.
In the past, in Europe, some wealthy people used to look after of a scholar living on their premises so they can ask them questions etc.
$100M, 2.35MW, 6000 ft^2
>>Designed for AI researchers, data scientists, and students, Project Digits packs Nvidia’s new GB10 Grace Blackwell Superchip, which delivers up to a petaflop of computing performance for prototyping, fine-tuning, and running AI models.
$3000, 1kW, 0.5 ft^2
Beyond that, the factors seem reasonable for 2 decades?
Isn't it actually FP64?
https://www.okdo.com/wp-content/uploads/2023/03/jetson-agx-o...
I wonder what the specifications are in terms of memory bandwidth and computational capability.
This is more accurately a descendant of the HPC variants like the article talks about - intentionally meant to actually be a useful entry level for those wanting to do or run general AI work better than a random PC would have anyways.
Anyone willing to guess how wide?
>This paper describes how the performance of AI machines tends to improve at the same pace that AI researchers get access to faster hardware. The processing power and memory capacity necessary to match general intellectual performance of the human brain are estimated. Based on extrapolation of past trends and on examination of technologies under development, it is predicted that the required hardware will be available in cheap machines in the 2020s.
and this is about the first personal unit that seems well ahead of his proposed specs. (He estimated 0.1 petaflops. The nvidia thing is "1 petaflop of AI performance at FP4 precision").
Bit bit hard to tell what's on offer on the GPU side, I wouldn't be surprised if it was RTX 4070 to 5070 in that range.
If the price/perf is high enough $3k wouldn't be a bad deal, I suspect a Strix Halo (better CPU cores, 256GB/sec memory interface, likely slower GPU cores) will be better price/perf, same max ram for unified memory, and cheaper.
[0]: https://newsroom.arm.com/blog/arm-nvidia-project-digits-high...
A lot of people have been justifying their Mac Studio or Mac Pro purchases by the potential for running large AI models locally. Project Digits will be much better at that for cheaper. Maybe it won't run compile Chromium as fast, but that's not what it's for.
edit: While the title says "personal", Jensen did say this was aimed at startups and similar, so not your living room necessarily.
The only thing it really competes with is the Mac Studio for LocalLlama-type enthusiasts and devs. It isn't cheap enough to dent the used market, nor powerful enough to stand in for bigger cards.
However we do know that it offers 1/4 the TOPS of the new 5090. It will be less powerful than the $600 5070. Which, of course it will given power limitations.
The only real compelling value is that nvidia memory starves their desktop cards so severely. It's the small opening that Apple found, even though Apple's FP4/FP8 performance is a world below what nvidia is offering. So purely from that perspective this is a winning product, as 128GB opens up a lot of possibilities. But from a raw performance perspective, it's actually going to pale compared to other nvidia products.
Running a 96GB ram model isn't cheap (often with unified memory 25% is reserved for CPUs), so maybe it will win there.
It's basically the successor to the AGX Orin and in line with its pricing (considering it comes with a fast NIC). The AGX Orin had RTX 3050 levels of performance.
https://s3.amazonaws.com/cms.ipressroom.com/219/files/20250/...
Source: https://nvidianews.nvidia.com/news/nvidia-puts-grace-blackwe...
Not sure if that isn't expected though? Likely most people wouldn't even notice, and the company can say they're dogfooding some product I guess.
The 5090 has 1.8TB/s of MBW and is in a whole different class performance-wise.
The real question is how big of a model will you actually want to run based on how slowly tokens generate.
Ideally we can configure things like Apple Intelligence to use this instead of OpenAI and Apple's cloud.
tinybox red and green are for people looking for a quiet home/office machine. tinybox pro is for people looking for a loud compact rack machine.” [0]
For $40,000, a Tinybox pro is advertised as offering 1.36 petaflops processing and 192 GB VRAM.
For about $6,000 a pair of Nvidia Project Digits offer about a combined 2 petaflops processing and 256 GB VRAM.
The market segment for Tinybox always seemed to be people that were somewhat price-insensitive, but unless Nvidia completely fumbles on execution, I struggle to think of any benefits of a Tinygrad Tinybox over an Nvidia Digits. Maybe if you absolutely, positively, need to run your OS on x86.
I'd love to see if AMD or Intel has a response to these. I'm not holding my breath.
>the size of several ATX desktops
I'm mildly skeptical about performance here: they aren't saying what the memory bandwidth is, and that'll have a major impact on tokens-per-second. If it's anywhere close to the 4090, or even the M2 Ultra, 128GB of Nvidia is a steal at $3k. Getting that amount of VRAM on anything non-Apple used to be tens of thousands of dollars.
(They're also mentioning running the large models at Q4, which will definitely hurt the model's intelligence vs FP8 or BF16. But most people running models on Macs runs them at Q4, so I guess it's a valid comparison. You can at least run a 70B at FP8 on one of these even with fairly large context size, which I think will be the sweet spot.)
This is really game changer.
They should make a deal with Valve to turn this into 'superconsole' that can run Half Life 3 (to be announced) :)
Just like Mac OS is free when you buy a Mac, having the latest high-quality LLM for free that just happens to run well on this box is a very interesting value-prop. And Nvidia definitely has the compute to make it happen.
Do I understand that right? It seems way to cheap.
Main issue is the ram they’re using here isn’t the same as is in GPUs
At $3,000, it will be considerably cheaper than alternatives available today (except for SoC boards with extremely poor performance, obviously). I also expect that Nvidia will use its existing distribution channels for this, giving consumers a shot at buying the hardware (without first creating a company and losing consumer protections along the way).
$3000 gets me a 64-core Altra Q64-22 from a major-enough SI today: https://system76.com/desktops/thelio-astra-a1-n1/configure
And of course if you don't care about the SI part, then you can just buy that motherboard & CPU directly for $1400 https://www.newegg.com/asrock-rack-altrad8ud-1l2t-q64-22-amp... with the 128-core variant being $2400 https://www.newegg.com/asrock-rack-altrad8ud-1l2t-q64-22-amp...
Joking aside, personally will buy this workstation in a heartbeat if I have the budget to spare, one in the home and another in the office.
Currently I have an desktop/workstation for AI workloads with similar 128GB RAM that I bought few years back that cost around USD5K without the NVIDIA GPU that I bought earlier for about USD1.5K, with a total of about USD6.5K without a display monitor. This the same price of NeXT workstation (with a monitor) when it's sold back in 1988 without adjusting for inflations (now around USD18K) but it is more than 200 times faster in CPU speed and more than 1000 times RAM capacity than the original 25 MHz CPU and 4 MB RAM, respectively. The later updated version of NeXT has graphic accelerator with 8 MB VRAM, since the workstation has RTX 2080 it is about 1000 times more. I believe the updated NeXT with graphic accelerator is the one that used to develop original Doom software [1].
If NVIDIA can sell the Project Digits Linux desktop at USD3K with similar or more powerful setup configurations, it's going to be a winner and probably can sell by truckloads. It seems to has NeXT workstation vibe to it that used to develop the original WWW and Doom software. Hopefully it will be used to develop many innovative software but now using open source Linux software eco-system not proprietary one.
The latest Linux kernel now has real-time capability for more responsive desktop experience and as saying goes, good things come to those who wait.
[1] NeXT Computer:
Future versions will get more capable and smaller, portable.
Can be used to train new types models (not just LLMs).
I assume the GPU can do 3D graphics.
Several of these in a cluster could run multiple powerful models in real time (vision, llm, OCR, 3D navigation, etc).
If successful, millions of such units will be distributed around the world within 1-2 years.
A p2p network of millions of such devices would be a very powerful thing indeed.
If you think RAM speeds are slow for the transformer or inference, imagine what 100Mbs would be like.
One can only wish for this, but Nvidia would be going against the decades-long trend to emaciate local computing in favor of concentrating all compute on somebody else's linux (aka: cloud).
Also I consider this a dev board. Soon this tech will be everywhere, in our phones, computers...
You could already plug that to your home assistant and have your own Star Trek computer you can ask questions from. And NVIDIA seems to know this is the future and they were the first in the market.
If one can skip buying gaming rig with a 5090 with its likely absurd price then this 3k becomes a lot easier for dual use hobbyists to swallow
Edit 5090 is 2k
The 5090 surprised me with the two slot height design while having a 575W power budget.
But it's clear that everyone's favorite goal is keretsuification. If you're looking for abnormal profits, you can't do better than to add a letter to FAANG. Nvidia already got into the cloud business, and now it's making workstations.
The era of specialists doing specialist things is not really behind us. They're just not making automatic money, nor most of it. Nvidia excelled in that pool, but it too can't wait to leave it. It knows it can always fail as a specialist, but not as a kereitsu.
I'm bracing for a whole new era of unsufferable binary blobs for Linux users, and my condolences if you have a non-ultramainstream distro.
MediaTek, a market leader in Arm-based SoC designs, collaborated on the design of GB10, contributing to its best-in-class power efficiency, performance and connectivity.
I assume that means USB and such peripherals is MediaTek IP, while the Blackwell GPU and Grace CPU is entirely NVIDIA IP.
That said, NVIDIA hasn't been super-great with the Jetson series, so yeah, will be interesting to see what kind of upstream support this gets.
[1]: https://nvidianews.nvidia.com/news/nvidia-puts-grace-blackwe...
While I'm quite the "AI" sceptic I think it might be interesting to have a node in my home network capable of a bit of this and that in this area, some text-to-speech, speech-to-text, object identification, which to be decent needs a bit more than the usual IoT- and ESP-chips can manage.
https://s3.amazonaws.com/cms.ipressroom.com/219/files/20250/...
This goes against every definition of cloud that I know off. Again proving that 'cloud' means whatever you want it to mean.
First product that directly competes on price with Macs for local inferencing of large LLMs (higher RAM). And likely outperforms them substantially.
Definitely will upgrade my home LLM server if specs bear out.
They mention 1 PFLOP for FP4, GB200 is 40 PFLOP.
Specs we’ve seen suggest the GB10 features a 20-core Grace CPU and a GPU that packs manages a 40th the performance of the twin Blackwell GPUs used in Nvidia’s GB200 AI server.Edit: Sorry fucked up my math. I wanted to do 40x52x4, $4/hr being the cloud compute price but that us actually $8300, so it is actually equivalent to about 4.5 months of cloud compute. 40 hours because I presume that this will only be used for prototyping and debugging, i.e during office hours.