Nvidia releases new AI chip with 480GB CPU RAM, 96GB GPU RAM
nvidia.com
nvidia.com
1 exaflop + 144TB memory
https://nvidianews.nvidia.com/news/nvidia-announces-dgx-gh20...
He calls it the worlds largest GPU. It’s just one, giant compute unit.
Unlike super computers, which are highly distributed, Nvidia says this is 140 TERABYTES of UNIFIED MEMORY.
My mind still gets blown just thinking about it. My poor desktop GPU has 4 gigabytes of memory. Heck, it only has 2 terabytes of storage!
It also explains the reason why despite having middling CPU power, mainframes had a reputation for stinkloads of I/O bandwidth so they could process everyone's credit card transactions, airline bookings, and that: the mainframe's CPU was involved very little in I/O, that was all handled by the channel processors!
> EDIT: found answer to my own question in the datasheet: "The NVIDIA Grace CPU combines 72 Neoverse V2 Armv9 cores with up to 480GB of server-class LPDDR5X memory with ECC."
So, this is not stacked RAM like HBM, it's LPDDR5X which a quick search says is 8.5Gbps.
Keep in mind unified does not mean uniform, the ram is distributed across all the GPUs.
Also, with quantum computers, the parallelization/"distribution" of tasks will be done within the same machine, as it can try all solutions and the same time without having to do divide-et-impera algorithms.
Also, in the future, the algorithms will be a lot simpler, and just have FPGA-s like AI chips, where there is no software, the model is directly modelled in the hardware, so each computation is instant (just the time it takes to propagate the electrons or light through the circuit).
Is it just they haven't done the molding of a production installation? Is it possible that their internal instances might not be that presentable?
I still remember how my mind was blown when I first learned that all of the memory in a Cray Y-MP was static RAM. Transistor-based flip-flops: extremely power hungry, but also very fast. Another way of looking at it is that all of its RAM was what we call "cache".
This, finally, looks like a supercomputer.
All of a sudden you don't care so much about the inefficiencies of walking linked lists or trees. When everything is "already in cache", you can worry less about cache efficient algorithms!
1 cycle memory access latency is one of the reasons why tiny embedded MCUs can do things with a fraction of the MHZ of their larger counterparts.
Now days of course it is all about tons of memory, tons of bandwidth, craptons of compute, and planning the flow of data ahead of time.
Cray loaded a ton of static memory into their computers, then liquid-cooled the whole thing. Sure, the power requirements were through the roof, and you had a whole huge chiller system which you had to run and hope it doesn't fail. If it did fail, you really wanted to shut the machine down fast. From what I recall there was also an emergency propeller inside the back case of the Y-MP 2E, and yes propeller is a much better name for this thing than a "fan". It would delay the inevitable, although dumping those tens of kW of heat into your server room was not something you ever wanted to do.
The whole point of all this was that you could do things that you couldn't with "normal" computers. That's why those were called "supercomputers". And I'm so glad that after a hiatus of about 30 years we're getting another wave of exceptional machines, which aren't just bigger PCs.
MCUs are in this category, lots of embedded stuff, including the two areas I'm familiar with: game controllers and lower spec'd wearables.
Very low power usage, CPU speed around 100mhz, so not too slow.
You can do plenty with 100mhz and SRAM!
The Y-MP came out in 1988, sixteen years after CRI was founded, which itself was several years after the CDC6600.
To be clear, this is floating point quarter-precision operations when using the FP8 tensor core arithmetic unit [1].
[1] https://resources.nvidia.com/en-us-grace-cpu/grace-hopper-su...
In comparison, Frontier is 1 exaflops of dense FP64. Try to run this Nvidia system as dense FP64, and it performance will degrade two orders of magnitude.
Don't get me wrong, the machine is really impressive, but the advertisement is quite misleading.
64 bit (or 'double precision') is still king in the HPC world though, as it is what you will find in large numerical solutions in fields like computation fluid dynamics, nuclear physics, etc.
I'm asking because whenever I look at ML training in the cloud, I never see any availability - either for this architecture or the A100s. AWS and GCP have quotas set to 0, lambda labs is usually sold out, paperspace has no capacity, etc. What we need isn't faster or bigger GPUs, it's _more_ GPUs.
Having said that, I don’t think we’re anywhere near some kind of equilibrium for AI compute. If chip supply would magically double tomorrow, then the large companies would buy it for their datacenters and have 100% utilization in a few weeks. They all want to train larger models and scale inference to more users.
Fabs can run multiple complex designs on the same line simultaneously by sharing common tools. For example, photolithography tools can have their reticles swapped out automatically. Obviously, there is a cost to the context switching and most designs cannot be run on the same line as others.
Ultimately, the smallest unit of fabrication capacity is probably best measured along grain of the lot/FOUP (<100 wafers).
We'll likely see more efficiency from bigger GPUs and hopefully more availability as a result.
There is a mode to attach 16 data pins to each GDDR chip, so with some extra effort you could probably double that to 48GB. Or at least 32GB. Maybe this is a valid niche, or maybe there isn't enough demand.
The alternative to this is HBM, which can stack up big amounts, but it's a lot more expensive.
Apple and their unified memory architecture may be the prod needed to get larger levels of RAM available to single cards solutions. We'll see.
The basic of Supply Chain and Supply and Demand, as you should have all witness during COVID for toilet rolls are the same.
Fab capacity is not that different to any other manufacturing. You just need to book those capacity way ahead of time. ( 6 - 9 months ) And that is also why I said 99% of news, or rumours about TSMC Capacity are pure BS.
So to answer your question. Yes, Nvidia will likely go for the higher margin products. One of the reason why you see Nvidia working with Samsung and Intel.
I dont know why you would assume that. Qualcomm has been using TSMC N4 since last year [1]. I'm sure there are other customers as well.
[1] https://www.anandtech.com/show/17395/qualcomm-announces-snap...
Unless all the planet does is make silicon wafers; no.
My back of the napkin basically suggested that silicon production would need to 4x and fab capacity 4x (neither of which are happening) and NVDA with would have to capture all of that to justify their current price. I didn't bother writing it up, just looked at it mostly because I was on the wrong side of that play. It's something worth considering for sure.
(n.b. that's really good work on your end and I agree with your conclusion, just idly musing about the thing that bugs me, what the heck all these non-leading edge fabs are going to do)
It seems to me that yes while a 200 P/E may be high, they certainly could keep increasing the prices on the already high margin datacenter products, of which get quickly gobbled up by companies no matter what price they are because of the immense demand.
For a science project, we need to manufacture magnets. It's not easy to find a company who has the right iron right now, and it's hard to get, with long lead times. The supply crisis is real.
Take a look at a recent GPU and count the auxiliary components. All of them can cause supply chain difficulties.
Apple already have the "neural" cores, is that more or less what they are?
Could there be a theoretical LLM chip for inference that is significantly cheaper to run?
When talking about the current ML industry it's more like nobody wants to invest significant amounts of money in hardware that will be obsolete before it's even taped out.
If you want more efficiency gain than general matrix multiplication hardware, you need to start getting specific about the NN architectures the hardware will support.
That’s why having 96gb (with another 480gb or whatever) available via high speed interconnect is a big deal. It means we can train bigger models faster.
As far as I can tell that gamble isn't work out particularly well for any of the startups but that might be money drying up before they've hit commercial viability. I know the hardware is pretty good for Graphcore, Cerebras and the software proving difficult.
I would say they are more like GPUs or DSPs: programmable but optimised for a specific application domain, ML/AI workloads in this case. Sometimes people call this ASIPs: application specific instruction set processors. While maybe not a very commonly used term, it is technically more correct.
As a rule companies should only do their own chips if they are certain they can solve and overcome the cogs problems that low yield and low volume penalties entail. If not you are almost certainly better off just eating the vendor margin. It is very very unlikely that you will do better.
Problem is, inference costs do not dominate training costs. Models have a very limited lifespan, they are constantly retrained or obsoleted by new generations, so training is always going on.
Training is not just matrix multiplications, given hundreds of experiments in model architecture, its not even obvious what operations will dominate future training. So a more general purpose GPU is just a way safer bet.
Also, LLM talent is in extreme short supply, and you don't want to piss them off by telling them they have to spend their time debugging some crappy FPGA because you wanted to save some hardware bucks.
Just curious what the current bar is here and which of the LLM-related skills might be worth building.
Being very good at fine tuning for a particular goal. Its much easier to learn fine-tuning, so standards are higher to stand out.
Being able to come up with architectural improvements for LLMs, aka the researcher path.
Wages start at $250k for grads at the big AI companies.
1. For BERT scale model, all you need is a good codebase from GitHub (I had some luck with this one [0]) and a few weeks of trial and error. Want to try training T5 or LLaMA, but don't have the resources needed. Of course training models with more than 100B parameters is another level of labyrinth.
2. Finetuning is mostly related to how well you understand the task and the data you are dealing with. Since the BERT paper focuses on the GLUE benchmark, I've become very proficient in fine-tuning GLUE and eventually got sick of it.
3. Made some architectural improvements to BERT, got decent results so I wrote a paper, and got rejected because the reviewers want a head-on evaluation against some well funded papers from Google.
4. Not in my country. Damn, I am envious.
Capital outlays are tied to the derivative of compute capacity, so even if training just flatlines, hardware spend will drop significantly.
Effectively it'd require the entirely memory controller and the cache, and scheduling. At point point you got most of the GPU w/ a stuck, non-programmable interface of a designated compute. Likely you'd never have to compete for advanced nodes as well.
I wonder what they’re doing with that hardware now.
Ethereum is designed to bottleneck on memory bandwidth (while being uncacheable) so at the end of the day the name of the game is how many memory channels can you slap onto a minimum-cost board. You won't drastically win on perf/w - but as mentioned by a sibling, 30-100% over a fully general-purpose gaming GPU is likely possible, because you don't have to have a whole general-purpose GPU sitting there idling (and it's not a coincidence that gaming GPUs were undervolted/etc to try and bring that power down - but you can't turn everything off). "ASIC-resistance" just means an ASIC is only 1-10x more efficient than a general-purpose device, so general-purpose hardware can still stay in the game. It doesn't mean ASIC-proof, you can still make ASICs and they still have at least some perf/w advantage.
However, if your ASIC costs $100 to get the same performance as a 3060 Ti, that's a huge win even if you only beat the perf/w by 50%. Particularly since your ASIC is likely way easier and more stable to deploy at scale, and doesn't require a host rig with at least a couple hundred bucks of computer gear to even turn on.
Only plebs were buying up GPUs from retailers or sniping websites, buying from ebay was for the chumpiest of chumps. Gangsters were buying them from the board partners a truckload at a time, true elites just pay someone to engineer an ASIC and do a small run of them. Eight-figures (as mentioned by a sibling) is plenty, a $50-75m run of ASICs is quite a lot of silicon even on a fairly modern node (and some mining companies were publicly known to be using TSMC 7nm and other very modern nodes). And when you invest that kind of money, you don't flash it around and scare the marks.
Can anyone with more legal knowledge share how they trademarked the name of Grace Hopper?
(And yes, I'd guess the codenames where chosen back in the day with an eye towards combining them in the same device.)
Trademarks are context specific and you can trademark "common terms" IF (and at lest theoretically only if) it's used in a very narrow use-case which by itself isn't confusable with the generic term.
The best example here is Apple which is a generic term but trademarked in context of phones/computer/music manufacturing (and by now a bunch of other things).
Through there had been an Apple music label with a bit of back and force of legal cases (and some IMHO very questionable court rulings) which in the end Ended by Apple buying that Label.
So theoretically it's not too bad.
Practically big companies like Apple, Nvidia and similar can just swamp smaller companies with absurd legal fees to force their win (AFIK this is Metas strategie because I honestly have no idea how they think the term Meta for data processing is trademarkable), to make it worse local curt have often shown to not properly apply the law in such conflicts if the other party is from an other country (one or two US states are infamous for very biased legal decision in this kind of cases).
So yeah at the core this aspect of the trademark system is not a terrible idea, but the execution is sadly often fairly lacking. And even high profile cases of trademark abuse often have no consequences if it's a "favorite big company". (For balance negative EU example do include Lego and it's 3d trademark and absurdly biased curt rulings, or Ferrero and it's Kinder (german. Children) trademark on Chocolate).
EDIT: also not the two TM: Grace™ Hopper™ both Grace and Hopper are generic terms you can under some circumstances trademark and then use together, but while probably legal you would likely want to avoid trademarking (Grace Hopper)™
Practically speaking, trademarks cover whatever can be litigated successfully.
Cue Apple Music, lmao. It's a suit-and-tie'd cult, I swear.
The worst offender is Tesla, because I'm pretty sure he would have hated that company.
1: https://en.wikipedia.org/wiki/Power_Macintosh_7100#Codename_...
For other's reference, nvidia arch's listed on wikipedia as follows in order:
Kelvin, Rankine, Curie, Tesla, Fermi, Kepler, Maxwell, Pascal, Volta, Turing, Ampere, Lovelace + Hopper
Who knew that one of the most profitable companies on Earth would get there by calling Carl Sagan a "butt head astronomer"!
It's similar to creating a work covered by copyright vs. registering it with the copyright office.
Grace is a trade mark. Hopper is a trade mark.
Hence each term having it’s own TM.
Do you plugin in DDR memory somewhere for the 480GB, or is this already on the board?
EDIT: found answer to my own question in the datasheet: "The NVIDIA Grace CPU combines 72 Neoverse V2 Armv9 cores with up to 480GB of server-class LPDDR5X memory with ECC."
> The NVIDIA Grace CPU Superchip uses the NVIDIA® NVLink®-C2C technology to deliver 144 Arm® Neoverse V2 cores and 1 terabyte per second (TB/s) of memory bandwidth.
> High-performance CPU for HPC and cloud computing Superchip design with up to 144 Arm Neoverse V2 CPU cores with Scalable Vector Extensions (SVE2)
> World’s first LPDDR5X with error-correcting code (ECC) memory, 1TB/s total bandwidth
> 900 gigabyte per second (GB/s) coherent interface, 7X faster than PCIe Gen 5
> NVIDIA Scalable Coherency Fabric with 3.2TB/s of aggregate bisectional bandwidth
> 2X the packaging density of DIMM-based solutions
> 2X the performance per watt of today’s leading CPU
https://resources.nvidia.com/en-us-grace-cpu/grace-hopper-su...
Not sure how one interfaces with it, but it presumably runs an approved Linux distro, with a web server at best.
source: I have one
It is many, many thousands.
These are for the 80GB versions, currently priced at $30,000 per GPU. It will likely be many months before this 96GB version is available to prosumers, if it ever is at all.
0: http://www.shopblt.com/cgi-bin/shop/shop.cgi?action=thispage...
1: https://www.cdw.com/product/nvidia-h100-gpu-computing-proces...
Ex - the HGX A100 platforms sold as single servers usually ran around 150k, but could get up above 200k depending on loadout.
Just getting an H100 (just the gpu) right now is ~40k new.
There is a reason nvidia's stock is doing so well...
- Am I wrong in understanding this is a general purpose computer (with massive graphic capabilities)?
- And if so, what CPU is it using (an NVIDIA ARM CPU)?
- And what OS does it run?
For OS, it will run some form of Linux. I'm not sure if the particular recommended build has been (or will be) publicly released.
The real problem is that ROCm is a fucking joke, pathetic, half assed, pretend project. Nobody with power in AMD seems to care that nobody can learn machine learning on their hardware to push it in other places, or that their GPUs that they have recently spent all this time boasting about their higher VRAM which is literally useless unless you want to play poorly optimized AAA titles ported from the PS5.
People say it works but you basically have to be one of the engineers who wrote it to prove that. Good luck getting it to work with Windows, or any hardware that wasn't purpose built for a cluster partner. It's so stupid. Maybe they genuinely intended to make a real CUDA competitor but noticed the ways that nVidia then had to artificially segment their market through dumb decisions (the VRAM) and bios hacks that didn't work and just gave up on that path.
Here is what works for me:
- Nvidia drivers on base linux system (rpmfusion/fedora in my case)
- Install nvidia container toolkit
- Use a cuda base container image and run all your code inside podman or docker
Back then the docs were just awful. Has this really changed that much in recent times?
Hmmm, what's the difference between homage and appropriation for things like this?
What I am also interested in, is why model sharding has to be done manually. It seems like, one should be able to write a framework that will take your forward step and distribute the amount of layers on the available GPUs, automatically. But I haven't come across such a framework yet.
This is part of the value proposition of Mojo, Chris Lattner’s latest project. The compiler infrastructure is still in its infancy, but looks promising: https://www.modular.com/mojo
Mac Studio M2 ULTRA has 192Gb of RAM, potentially 188Gb available for GPU, for 5K.
Wouldn't apple be able to compete with that if they scaled it up?
Spoiler alert: CUDA ecosystem, Linux suport, and most importantly for data centers, Mellanox high speed interfacing with virtually infinite scalability and great virtualization support so they can rent out slices of their HW to customers in exchange for money.
By contrast, M2 Ultra has 800 GB/s memory bandwidth, 31.6 half-precision TFlops in the Neural Engine, and (extrapolating from https://en.wikipedia.org/wiki/Apple_M2), about 27 single-precision TFlops on the GPU.
So 5x memory bandwidth, more than double generic throughput, and at least 32x peak tensor throughput. Sure, the Mac Studio uses much less power, but depending on the application that usually doesn't make up for the speed difference.
And this is just for a consumer GPU, I haven't even touched on the datacenter-grade stuff.
tl;dr the M2 is an underpowered GPU with a lot of RAM close by, while NVIDIA cards are multiple orders of magnitude more powerful but most of the RAM's a bit farther away.
seriously, measuring graphics cards in gigabytes is like measuring battery life in volts. it's the wrong unit.
Either way quite a bit better than 30x slower.
Yes, but, if you need 48GB to run inference on a model, and you only have 24GB available, you don't get to enjoy the tflops difference.
If nVidia released somewhat lower-performance GPUs but with more VRAM, we could talk. But they're not stupid :)
A more appropriate comparison is the fp16 performance.
It seems to be 27 tflops for the 38 core M2, and 330 for the 4090.
The more useful for training fp16 with fp32 accumulate is 165 for the 4090, I don't know about the apple one.
You might be able to compare the number of CUDA cores to the ALU count of the Apple GPUs. I don't know what that is for M2 Ultra yet, but for the 64 core M1 Ultra each core had 16 execution units and each of those had 8 ALUs, for a total of 8,192 ALUs. The M1 Ultra's FP32 performance was in the ballpark of 21tflops - assuming a ~30% improvement in the M2 Ultra that takes us to ~27tflops. Google suggests that for the RTX 4090 it's 83tflops.
• The GPU in the M2 Ultra has 76 GPU cores and corresponds to 2x M2 Max, which has 13.6 tflops; so 27.2 tflops [1,2]
• RTX 4090 has 82.58 tflops [3] (overclocked can reach 100 tflops [4])
While more powerful NVidia cards are not "multiple orders of magnitude more". Rather it seems a 4090 is around 3-4 faster than an M2 Ultra GPU.
Keep in mind that the Apple Silicon chips also have low-precission Neural Engine circuits for inference of neural nets. For the M2 Ultra they claim 31.6 tops [1].
[1] https://www.apple.com/newsroom/2023/06/apple-introduces-m2-u...
[2] https://www.notebookcheck.net/Apple-unveils-M2-Pro-and-M2-Ma...
[3] https://www.techpowerup.com/gpu-specs/geforce-rtx-4090.c3889
[4] https://videocardz.com/newz/overclocked-nvidia-rtx-4090-gpu-...
For the M1 the stated values are:
M1 Ultra [1]: 21 TFLOPs
M1 Max [2]: 10.6 TFLOPS
Actual performance obviously depends on the type of work-load etc. E.g. Geekbench Ultra is only 50% faster [3]: M1 Ultra Geekbench 6: 150260
M1 Max Geekbench 6: 108198
[1] https://videocardz.com/newz/despite-apples-claims-m1-ultra-g...
[2] https://wccftech.com/m1-max-teraflops-performance-higher-tha...
[3] https://browser.geekbench.com/metal-benchmarksPeople complain about the “Nvidia tax” but the hardware is superior (untouchable at datacenter scale) and the “tax” turns into a dividend as soon as your (very expensive) team spends hour after hour (week after week) dealing with issues on other platforms compared to anything based on CUDA often being a Docker pull away with absolutely first class support on any ML framework.
Nvidia gets a lot of shade on HN and elsewhere but if you’ve spent any time in this field you completely understand why they have 80-90% market share of GPGPU. With Willow[0] and the Willow Inference Server[1] I'm often asked by users with no experience in the space why we don't target AMD, Coral TPUs (don't even get me started), etc. It's almost impossible to understand "why CUDA" unless you've fought these battles and spent time with "alternatives".
I’ve been active in the space for roughly half a decade and when I look back to my early days I’m amazed what a beginner like me was able to do because of CUDA. I still routinely am. What you’re able to actually accomplish with a $1000 Nvidia card and a few lines with transformers and/or a Docker container is incredible.
That said I am really looking forward to Apple stepping it up here - I’ve given up on AMD ever getting it together on GPGPU and Intel (with Arc) is even further behind. The space needs some real competition somewhere.
I think the only reason CUDA isn't talked about like the monumentally important human milestone in technological development that it is, is that it is a pretty abstract thing that is difficult for laypeople to visualize.
I've been accused of being an Nvidia "fanboy" when I touch on this. I attempt to explain:
"Nvidia made the very hard, very expensive commitment to developing and supporting CUDA 15 years ago when this space was in it's infancy (non-existent). They SUNK incredible resources into this gamble/vision to universally support CUDA on every chip and every platform - for 15 years. They didn't achieve their position through shady dealings or luck, they earned it with a decade and a half of investment, focus, and execution."
Granted, they do somewhat abuse the position they have now (as often noted on HN and elsewhere). At the risk of whataboutism I ask: "Show me a corporation that wouldn't. Do you think if AMD had their market share they'd be nice, cuddly good guys?" Microsoft, Intel, Nvidia, Standard Oil, the Phoebus cartel, etc - it goes on and on. Always has and always will. Of course I'm not saying it's a good thing, it's just a fact of the real world.
How are Nvidia consumer GPUs a bad deal? When comparing top of the line cards (I'm not going to bother looking elsewhere in the product lines) for 60% more cost (yes, significant) with an RTX 4090 you get:
- Performance that walks all over the 7900 XTX[0].
- The ability to use it to self-host, experiment, learn, etc a never-ending range of ML applications that (as discussed) "just work" as a docker pull away.
With an RTX 4090 you could have stunning gaming performance one minute and then seconds later be running a local LLM, etc. That is tremendously more value than the 60% price difference. If all I wanted to do was gaming and I was price sensitive I'd save myself the $600 and be happy with an AMD GPU. But, looking at market share[1] (at least 80% across desktop gaming and GPGPU) either the value of AMD GPUs is a little known secret (it isn't) or consumers, with the ability for choice, overwhelmingly see the value in Nvidia GPUs.
[0] - https://techguided.com/7900-xtx-vs-rtx-4090/
[1] - https://wccftech.com/q3-2022-discrete-gpu-market-share-repor...
We (who want a real alternative to NVIDIA) either find a way to pool global resources applicable to this level kind of effort on all new accelerator architectures, or we wait for them to stumble, Intel-like. Intel and AMD not pooling their resources on this is self-defeating.
Would you be willing to elaborate ( ie. I would love to hear you get started )? I absolutely agree that some competition is needed in this space. I am absolutely not an expert so it is hard for me to understand why there is no real alternative to CUDA. Are they just too hard to set up? Not popular enough to have any support?
That's compelling enough to justify the position here but you can do further research to explore the challenges with other platforms (like ROCm). Just glance at issue trackers for Pytorch, Tensorflow, and higher level projects that (rarely) support ROCm - you will notice a clear trend. Even though CUDA is more capable and outnumbers AMD use 10:1 the issues reported currently are:
Pytorch
Search for ROCm - 2,557 issues open (~10% market share)
Search for CUDA - 5,430 issues open (at least 80% market share)
Even in the limited cases where it's attempted the capability and experience with AMD/ROCm is significantly worse, to the point of "almost no one even bothers anymore" (see my top paragraph).
The Studio would run large workload many times slower than a high-end server, if it could run them at all.
The memory bandwidth of the nVidia GPUs is also significantly higher than the M2 parts.
The Apple silicone parts are impressive for what they are, but they don’t have a huge efficiency advantage for GPU compute. The full-size GPUs with huge memory buses and a large number of cores are still significantly more powerful.
There’s also a matter of getting data into and out of the GPUs and across the network, which takes a lot more than 10Gbe
The Apple silicon is great for running development workloads locally, but it’s significantly slower than full size GPUs.
I’m kind of surprised at how quickly everyone forgot that Apple’s marketing material greatly exaggerates their GPU performance.
https://ourworldindata.org/grapher/historical-cost-of-comput...
Sure, they’ve gotten a bit faster but it’s still fairly expensive to outfit more and more RAM.
GPUs though are indeed much more than memory. However, Apple has a unique unified memory model that no one else has matched yet where the memory is the heart of the machine. That means you don’t even need as much bandwidth because you can transparently access the same data by any chip at the same speed. That’s a pretty powerful design. I doubt Apple will really go into the training side of things because that’s not germane to their use cases because that happens in the cloud where they don’t have a presence yet. Inference is so expect more LLM / stablediffusion acceleration. Now if fine tuning models with additional training becomes a thing, then you’ll see acceleration of that modality on Apple’s machines. But training won’t be a focus because Apple doesn’t care about fickle nerd points that aren’t relevant to their business.
> https://www.guru3d.com/news-story/tsmc-has-secured-seven-maj...
There will very likely be products from Intel, Qualcomm, Nvidia and AMD stuff using it too this year.
So, not for everyone.
It's the same case as with high-end electronic components since 10-15 years ago. No one can/could produce their own smartphone because the high end components are sold only to the largest incumbents. A medium-size startup "has the money" to buy these, even in volume. But qualcomm won't sell.
Same fate expects us here in general purpose computing. Hard to believe, but so was what happened to the electronics supply chain, at the time.
Granted, we can use them for bad things, too, but nobody wants to get stabbed with their own knife.
I don't think so Tim.
Just wish I could have typed half as well as Hugh Jackman in Swordfish; then I wouldn't feel so sorry for all of the bugs I introduced into production because of her.
(Though we do need to pay attention to evolution cheating by overfitting relative to what we'd consider a clean design. Some of the complexity may be doing double duty.)
Since we have not succeeded in imitating even the most primitive brains, even though computationally we should have enough juice by now, it would seem that complexity can't be discarded at all, no?
looks at the browser tab with GPT-4 in it
looks back again here
... we didn't?
Seriously, get it to successfully play through a text adventure maze game.
I did exactly that the other day in response to a different objection on a HN thread. Or at least similar enough.
https://cloud.typingmind.com/share/c0a68cb2-5f59-4e83-b383-b...
Now, the goal there wasn't to get it to solve a maze, but rather to see how it can come up with a plan of action and adjust it on the fly. But I see no reason a variant of that wouldn't work with a traditional maze game - provided you remember this is a stateless model without volatile memory, so it needs to be fed its memory with every request.
This is also the reason why its output sounds convincing, but is very often factually wrong.
"imitating even the most primitive brains, even though computationally we should have enough juice by now"
Which is kind of weird to claim today. GPT-4 may be the strongest counterexample to date, but it's far from the only one.
Of course, you need to remember not to confuse the brain with attached peripherals. Just because we can't replicate a perfect worm or fly body, complete with bioelectrical and biomechanical components, doesn't mean we can't do better than their brains in silico.
And if it walks like a duck, and quacks like a duck, ...
GPT-4 is a good example because it's pretty clear that the model isn't merely a stochastic parrot (or, if it is in some sense, then in that sense so are we). But it's not the only game in town. Not all generative transformers deal with language. All seem to be powerful association machines, drawing their capabilities from simple algorithms in absurdly high-dimensional spaces. There are many parallels you can draw to brains here, not the least of which is that the overall architecture is simple enough and scalable, that it's exactly the kind of thing evolution could reach and then get railroaded into building on.
Just like with transformers revolutionising text generation and now things like LoRa and other fine tuning methods are helping us find a better solution to that puzzle, the same will happen for the development of AGIs.
We will do it, one day.
We aren't likely ever going to reduce that to a model as simple as the one used in machine learning, because it probably isn't that simple period.
Neurons are not "just" electrical signalling devices. They are complicated processors and systems in their own right.
False, to invent car.
People manage to learn to operate cars in less than 50 hours
The GP is pointing out that training in fine muscle motor skills, self-awareness and ability to project self-awareness to other objects under ones control etc, all took many thousands of years to develop. AI is faster.
However, it's again unfair, as AI only knows what it knows from us, so in that sense any comparison is built on shaky ground.
But for the purposes of comparing a stock human brain as hardware, versus a current high-end GPU specifically in terms of ingesting information and then perform tasks, the GPU beats the human brain "hands-down" in any category.
The only categories it doesn't are simply ones that no one has trained it to yet - so the argument on a pure hardware capability basis stands.
It's just the wrong way to look at the problem. You're not trying to develop the generic system that can learn how to drive a car, you're trying to develop the specific system that can safely drive a car occupied by humans, naturally employing machine learning.
I would argue that we're 95% there, but solving those last 5% is exponentially more expensive, but not commensurately more valuable. There's a "profit ceiling" imposed by the cost of a human driver, which appeats to make solving the problem economically intractable.
Here's a video of Tesla FSD driving through the complicated streets of Los Angeles for an hour straight, with 0 human intervention.
In compare to human's tens of hours?
We have to include our evolutionary process because a lot of our brain is pretrained, especially visual/motor neurons.
We seem to be pretrained to pick up language, for example, and the language(s) we hear after being born fill that space, our brains are plastic for a reason.
On top of typically around 18 years of learning to process and fuse vision, sound, proprioception and other inputs, to navigate the world and reason about it.
2. 16 years of fine-tuning to adapt to the current modern world.
3. 50-100 hours of specific task based fine-tuning for driving, think LORA training.
Power efficiency does not matter for those.
Those are interesting only for the implementation of specialized cognitive functions.
Additionnaly, the connectomes of those chips are 2D and very localized. Human brains are 3D and much less localized. Simulating 3D connectome with those 2D chips slows everything down by a lot.
https://news.mit.edu/2018/study-reveals-how-brain-tracks-obj...