GPUs Go Brrr
hazyresearch.stanford.edu
hazyresearch.stanford.edu
From a philosophical point of view, we think a frame shift is in order. A “register” certainly shouldn’t be a 32-bit word like on the CPUs of old. And a 1024-bit wide vector register, as CUDA uses, is certainly a step in the right direction. But to us a “register” is a 16x16 tile of data. We think AI wants this."
The hardware needs of AI are starting to focus. GPUs, after all, were designed for an entirely different job. They're used for AI because they have good matrix multiply hardware. "AI GPUs" get to leave out some of the stuff in a real GPU (does an H100 even have texture fill units?). Then there's a trend towards much shorter numbers. 16 bit floating point? 8 bit? 2 bit? 1 bit? That will settle out at some point. This paper indicates that hardware that likes 16x16 tiles makes a lot of sense. It's certainly possible to build such hardware. Someone reading this is probably writing it in VHDL right now, or will be soon.
Then we'll see somewhat simpler, less general, and cheaper devices that do "AI" operations with as little excess hardware baggage as possible. Nice.
Some slides to get the gist of it: https://users.encs.concordia.ca/~asim/COEN_6501/Lecture_Note...
Apple has already been doing this for a few years now. The NPU is totally different from the GPU or CPU on the die itself[1]. Nvidia is likely working on this as well, but I think a device that's a gaming/entertainment/crypto/AI bundle (i.e. sticking with the video card) is probably a better business move.
[1] https://github.com/hollance/neural-engine/blob/master/docs/a...
Nvidia escapes this by designing their GPU architecture to incorporate NPU concepts at a fundamental level. It's less redundant silicon and enables you to scale a single architecture instead of flip-flopping to whichever one is most convenient.
I would guess that their strategy is to not include powerful client-side hardware, and supplement that with some kind of "AiCloud" subscription to do the battery-draining, heat-generating stuff on their cloud. They're trading off their branding as a privacy focused company under the (probably correct) belief that people will be more willing to upload their data to iCloud's AI than Microsoft's.
Fwiw, I think they're probably correct. It has always struck me as odd that people want to run AI on their phone. My impression of AI is that it creates very generalized solutions to problems that would be difficult to code, at the cost of being very compute inefficient.
I don't really want code like that running on my phone; it's a poor platform for it. Thermal dissipation and form factor limit the available processing power, and batteries limit how long you can use the processing power you have. I don't really want to waste either trying to do subject identification locally. I'm going to upload the photos to iCloud anyways; let me pay an extra $1/month or whatever to have that identification happen in the cloud, on a server built for it that has data center thermal dissipation and is plugged into the wall.
So... I think users will be stuck. They'll want to run uncensored models on their phone, but Apple will want to keep them in the walled garden at any cost. It feels like the whole "Fortnite" situation all over again, where users can agree they want something but Apple can't decide.
I don't equate AI with coding. I want AI locally for photo sorting and album management, for general questions answering/list making that I use GPT for, and any number of other things.
I try not to upload personal data to sites that aren't E2E encrypted, so iCloud/Google photos is a no-go.
You might not be in area of poor connection and can't connect to the cloud.
One use for AI is speech recognition / transcription for deaf/HoH individuals. Up until now its almost been done exclusively on the cloud and it works fairly well (depending on conditions). Recently there's been an interest in doing it locally without relying on a network connection.
There's also privacy issues with transmitting this data over a network.
I guess we can assume this is going to be what’s used in what’s being called Apple’s first AI phone, iPhone 16.
So the claims that NVIDIA's GPUs are already thoroughly optimized for AI and that there's no low-hanging fruit for further specialization don't seem too plausible, unless you're only talking about the part of the datacenter lineup that has already had nearly all fixed-function graphics hardware excised. And even for Hopper and Blackwell, there's some fat to be trimmed if you can narrow your requirements.
Some fraction of your transistors MUST go unused on average or you melt the silicon. This was already a thing in the 20nm days and I'm sure it has only gotten worse. 100% TDP utilization might correspond to 60% device utilization.
On kernels such as flash attention, TMA and the L2 cache are both fast enough so as to hide these problems reasonably well. But to make the full use of the hardware, memory request must be coalesced and bank conflicts avoided ”
The depth of the competition is also starting to become apparent. There’s no way the documentation error was totally an accident. Diagrams are the easiest to steal / copy and there must have been some utility for nvidia to have left this in place. Remember when Naveen Rao’s Nervana was writing NVidia Maxwell drivers that out-performed NVidia’s own? Not every documentation mishap in a high-growth product is a competition counter-measure, but given that the researchers spent so long reverse-engineering wgmma and given the China-US political situation of the H100 in particular, it seems NVidia is up to its old tricks to protect its moat.
So don’t over-study the H100 peculiarities, as “what hardware does AI want?” really encompasses the commercial situation as well.
Except for special customers who will pay for the genuine item.
[1] https://www.tomshardware.com/tech-industry/artificial-intell...
The original REV A iMac in late 90s had slotted memory for its ATI card, as one example - shipped with 2mb, could be upgraded to 6mb after the fact with a 4MB SGRAM DIMM. There are also a handful of more recent examples floating around.
While I'm sure there are also packaging advantages to be had by directly soldering memory chips instead of slotting them etc, I strongly suspect the desire to keep buyers upgrading the whole card ($$$) every few years trumps this massively if you are a GPU vendor.
Put another way, what's in it for the GPU vendor to offer memory slots? Possibly reduced revenue, if it became industry norm.
The answer to this question almost has to be "because it will be cheaper to buy it tomorrow." However, GPUs bundle together RAM and compute. If RAM is likely to be cheaper tomorrow, isn't compute also probably going to be cheaper?
If both RAM and compute are likely cheaper tomorrow, then the calculus still probably points towards a wholesale replacement. Why not run/train models twice as quickly alongside the RAM upgrades?
> I strongly suspect the desire to keep buyers upgrading the whole card ($$$) every few years trumps this massively if you are a GPU vendor.
Remember as well that expandable RAM doesn't unlock higher-bandwidth interconnects. If you could take the card from five years ago and load it up with 80 GB of VRAM, you'd still not see the memory bandwidth of a newly-bought H100.
If instead you just need the VRAM and don't care much about bandwidth/latency, then it seems like you'd be better off using unified memory and having system RAM be the ultimate expansion.
the thing is, there's not enough competition in the AI-GPU space.
Current only option for no-wasting-time on running some random research project from github? buy some card from nvidia. cuda can run almost anything on github.
AMD gpu cards? that really depends...
and gamers often don't need more than 12?gb of GPU ram for running games on 4k.. so most high-vram customers are on the AI field.
> If you could take the card from five years ago and load it up with 80 GB of VRAM, you'd still not see the memory bandwidth of a newly-bought H100.
this is exactly what nvidia will fight against tooth-and-nail -- if this is possible, its profit margin could be slashed to 1/2 or even 1/8
No, it doesn't. It could just as easily be "because I will have more money tomorrow." If faster compute is $300 and more VRAM is $200 and I have $300 today and will have another $200 two years from now, I might very well like to buy the $300 compute unit and enjoy the faster compute for two years before I buy the extra VRAM, instead of waiting until I have $500 to buy both together.
But for something which is already a modular component like a GPU it's mostly irrelevant. If you have $300 now then you buy the $300 GPU, then in two years when you have another $200 you sell the one you have for $200 and buy the one that costs $400, which is the same one that cost $500 two years ago.
This is a much different situation than fully integrated systems because the latter have components that lose value at different rates, or that make sense to upgrade separately. You buy a $1000 tablet and then the battery goes flat and it doesn't have enough RAM, so you want to replace the battery and upgrade the RAM, but you can't. The battery is proprietary and discontinued and the RAM is soldered. So now even though that machine has a satisfactory CPU, storage, chassis, screen and power supply, which is still $700 worth of components, the machine is only worth $150 because nothing is modular and nobody wants it because it doesn't have enough RAM and the battery dies after 10 minutes.
If you have a CPU and it has however many cores, the amount of memory or memory bandwidth you need to go with that is totally independent, and memory bandwidth is rarely the bottleneck. So you attach a couple memory channels worth of slots on there and people can decide how much memory they want based on whether they intend to have ten thousand browser tabs open or only one thousand. Neither of which will saturate memory bandwidth or depend on how fast the CPU is, so you don't want the amount of memory and the number of CPU cores tied together.
If you have a device for doing matrix multiplications, the amount of RAM you need is going to depend on how big the matrix you want to multiply is, which for AI things is the size of the model. But the bigger the matrix is, the more memory bandwidth and compute units it needs for the same number of tokens/second. So unlike a CPU, there aren't a lot of use cases for matching a small number of compute units with a large amount of memory. It'd be too slow.
Meanwhile the memory isn't all that expensive. For example, right now the spot price for 64GB of GDDR6 is less than $200. Against a $1000 GPU which is fast enough for that much, that's not a big number. Just include it to begin with.
Except that they don't. The high end consumer GPUs are heavy on compute and light on memory. For example, you can get the RTX 4060Ti with 16GB of VRAM. The RTX 4090 has four times as much compute but only 50% more VRAM. There would be plenty of demand for a 4090 that cost $200 more and had four times as much VRAM, only they don't make one because of market segmentation.
Obviously if they don't do that then they're not going to give one you can upgrade. But you don't really want to upgrade just the VRAM anyway, what you want is for the high performance cards to come with that much VRAM to begin with. Which somebody other than Nvidia might soon provide.
the mass press is very impressed by Cuda, but at least if we're talking AI (and this article is, exclusively), it's not the right interface.
and in fact, Nv's lead, if it exists, is because they pushed tensor hardware earlier.
Nv's lead is due to them having Pytorch support.
Everybody cares about VRAM right now yet you can get a P40 with 24GB for 10% of the price of a 24GB RTX 4090. Why? No tensor cores, the things used for matrix multiplication.
if you segregate AI units from the GPU, the thing is both AI and GPUs will continue to need massive amounts of matrix multiplication and as little memory latency as possible
the move to have more of it wrapped in the GPU makes sense but at least in the short and medium term, most devices won't be able to justify the gargantuan silicon wafer space/die growth that this would entail - also currently Nvidia's tech is ahead and they don't make state of the art x86 or ARM CPUs
for the time being I think the current paradigm makes the most sense, with small compute devices making inroads in the consumer markets as non-generalist computers - note that more AI-oriented pseudo-GPUs already exist and are successful since the earlier Nvidia Tesla lineup and then the so-called "Nvidia Data Center GPUs"
Should be "as much memory bandwidth as possible". GPUs are designed to be (relatively) more insensitive to memory latency than CPU.
https://www.amd.com/en/products/accelerators/alveo/v80.html
XDNA Architecture
There was that recent paper titled "The Era of 1-bit LLMs" [0] which was actually suggeting a 1.58 bit LLM (2 bits in practice).
> Someone reading this is probably writing it in VHDL right now, or will be soon.
Yeah, I think I'm in the "will be soon" camp - FPGA board has been ordered. Especially with the 2-bit data types outlined in that paper [0] and more details in [1]. There's really a need for custom hardware to do that 2-bit math efficiently. Customizing one of the simpler open source RISC-V integer implementations seems like something to try here adding in the tiled matrix registers and custom instructions for dealing with them (with the 2 bit data types).
[0] https://arxiv.org/abs/2402.17764 [1] https://github.com/microsoft/unilm/blob/master/bitnet/The-Er...
1 trit, not 2 bits. 3 trits are 27 states, which can be represented with 5 bits.
Wondering if anyone would be surprised that a huge amount of progress in AI is on the engineering side (optimizing matmuls), and that a huge portion of the engineering is about reverse engineering NVIDIA chips
This is heavily used in ResNets (residual networks) for computer vision, and is what allows training much deeper convolutional networks. And transformers use the same trick.
The line between GPU lingo and Star Trek technobabble fades away further and further.
Have had that thought occasionally with some of the other articles. What it must read like to somebody who gets a ref link for an article over here. Wandered into some Trek nerd convention discussing warp cores.
https://en.wikipedia.org/wiki/Metric_tensor_(general_relativ...
Building a single analog chip with 1 billion neurons would cost billions of dollars in a best case scenario. A Nvidia card with 1 billion digital neurons is in the hundreds of dollars of range.
Those costs could come down eventually, but at that point CUDA may be long gone.
Sure accuracy is an issue, but this is not as impossible as you may think it would be. The main question will be if the benefits by going analog outweigh the issues arising from it.
I don't think reprogrammable analog circuits would really be feasible, at least with today's tech. You'd need to modify the resistors etc. to make it work.
* Finding ways to use extant chip fab technology to produce something that can do analog logic. I've heard CMOS flash presented a plausible option.
* Designing something that isn't an antenna.
* You would likely have to finetune your model for each physical chip you're running it on (the manufacturing tolerances aren't going to give exact results)
The big advantage is that instead of using 16 wires to represent a float16, you use the voltage on 1 wire to represent that number (which plausibly has far more precision than a float32). Additionally, you can e.g. wire two values directly together rather than loading numbers into an ALU, so the die space & power savings are potentially many, many orders of magnitude.+/- 1e-45 to 3.4e38. granted, roughly half of that is between -1 and 1.
When we worked with low power silicon, much of the optimization was running with minimal headroom - no point railing the bits 0/1 when .4/.6 will do just fine.
> Additionally, you can e.g. wire two values directly together rather than loading numbers into an ALU
You may want an adder. Wiring two circuit outputs directly together makes them fight, which is usually bad for signals.
If that was true, then a DRAM cell could represent 32 bits instead of one bit. But the analog world is noisy and lossy, so you couldn't get anywhere near 32 bits of precision/accuracy.
Yes, very carefully designed analog circuits can get over 20 bits of precision, say A/D converters, but they are huge (relative to digital circuits), consume a lot of power, have low bandwidth as compared to GHz digital circuits, and require lots of shielding and power supply filtering.
This is spit-balling, but the types of circuits you can create for a neural network type chip is certainly under 8 bits, maybe 6 bits. But it gets worse. Unlike digital circuits where signal can be copied losslessly, a chain of analog circuits compounds the noise and accuracy losses stage by stage. To make it work you'd need frequent requantization to prevent getting nothing but mud out.
And those things are huge which lead to very small network sizes. This is partially due to the fabrication node, but also simply because there is even less well developed tooling for analog circuits compared to digital ones compared to software compilers
[1] https://electronicvisions.github.io/documentation-brainscale... [2] https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8907969/ [3] https://arxiv.org/pdf/2003.11996
[1] Are H100s' maxtrix floating point units actually IEEE 754 compliant? I don't actually know.
Biological neural networks are nowhere near as connected as ANNs, which are typically fully connected. With biological neurons, the ingress / egress factors are < 10. So they are highly local
It is also an entirely different model, as there is no such thing as backpropagation in biology (that we know of).
What they do have is lieu of backpropagation is feedback (cycles)
And maybe there are support cells/processes which are critical to the function of the CNS that we don't know of yet.
There could also be a fair amount of "hard coded" connectedness, even at the higher levels. We already know of some. For instances, it is known that auditory neurons in the ears are connected and something similar to a "convolution" is done in order to localize sound source. It isn't an a emergent phenomena - you don't have to be "trained" to do it.
This is not surprising give life has had billions of years and a comparable number of generations in order to figure it out.
I guess in theory this could all be done in software. However given the trillion+ neurons in primate/human brains, this would be incredibly challenging on even the thousand-core machines we have nowadays. And before you scream "cloud" it would not have the necessary interconnectedness/latency.
It would be cool if you could successful model say, a worm/insect with this approach.
I wonder where the partial data / feedback is stored. Don't want to sound like a creationist, but it seems very improbable that "how good my sound localization is" is inferred exclusively from the # of children I have.
Being able to localize sound source has a lot of benefits including predation avoidance and prey detection.
would love to revisit the material, especially in this new era of specialized processing units and UMA.
I agree with you in the sense that something has "died" in writings the follow academic paper speak these days. Just yesterday I saw an ancient article surfaced by Scientific American and Peter Norvig on System Analysis by Strachey. It uses quite a bit of formal language but is super approachable at the same time. That kind of skill is rarely seen these days.
https://twitter.com/TheRoaringKitty/status/17900418133798504...
How can I put this in your vernacular...
"Most polite genZ meme enjoyer"
Good writing is also entertaining and engaging.
If you want to use enterprise AMD gpus, I'm renting them. That said, I haven't even had a chance to run/play with them myself yet, they have been rented since I got them last month.
Yes, we are getting more.
Also, write your own cuda kernel to do vector-matrix multiplication (if you use pycuda, you can focus on the kernel, and write everything else with python). Just tell chatgpt that you want to write your own implementation that multiplies a 4000-element vector by 4000x12000 matrix, and to guide you through the whole process.
For renting gpus, runpods is great - right now they have everything from lower tier gpus to h100s. You can start with a lesser gpu at the beginning.
I spent 2 months implementing a matmult kernel in Spiral and optimizing it.
However, in the Spiral series, I aim to go beyond just making an ML library for running NN models and break new ground.
Newer GPUs actually support dynamic memory allocation, recursion, and the GPU threads have their own stacks, so you could in fact treat them as sequential devices and write games and simulators directly on them. I think once I finish the NL Holdem game, I'll be able to get over 100x fold improvements by running the whole program on the GPU versus the old approach of writing the sequential part on a CPU and only using the GPU to accelerate a NN model powering the computer agents.
I am not sure if this is a good answer, but this is how GPU programming would be helpful to me. It all comes down to performance.
The problem with programming them is that the program you are trying to speed up needs to be specially structured, so it utilizes the full capacity of the device.
Even so, creating all the abstractions needed to implement even regular matrix multiplication in Spiral in a generic fashion took me two months, so I'd consider that good enough exercise.
You could do it a lot faster by specializing for specific matrix sizes, like in the Cuda examples repo by Nvidia, but then you'd miss the opportunity to do the tensor magic that I did in the playlist.
[1]: https://matplotlib.org/stable/gallery/showcase/xkcd.html#sph...
causal: https://github.com/HazyResearch/ThunderKittens/blob/main/exa...
non-causal: https://github.com/HazyResearch/ThunderKittens/blob/main/exa...
I'm not in this field at all, but it seems to me that using general purpose processors that communicate over (relatively) slow lanes can only get us so far. Rethinking the design at the hardware level, and eventually bringing the price down for the consumer market seems like a better long-term strategy.
I am not sure that is true. A glance/or long stay at the reddit localllama subreddit basically has a bunch of frustrated CPU users trying their absolute best to get anything to work at useful speeds.
When you can get an Nvidia GPU for a few hundred dollars or a full blown gaming laptop with a 4050 6gb vram for $900, its hard to call a CPU based AI capable.
Heck we don't have GPUs at work, and CPU based is just not really reasonable without using tiny models and waiting. We ended up requesting GPU computers.
I think there is a 'this is technically possible', and there is a 'this is really nice'. Nvidia has been really nice to use. CPU has been miserable and frustrating.
But that may be straining the meme somewhat. :)
In most languages the use of a noun as an adjective is marked, by a particle or by an affix or at least by a different stress pattern (like moving the stress to the last syllable), which removes the ambiguities.
So for most non-native speakers "Canadian goose" makes much more sense than "Canada goose" (which may feel like "Canada and a goose" or "a goose that is also Canada" and not like "a goose from Canada").
In the names "Canada Goose", "Long Island Shellfish" and "Dublin Bay Prawns", "Canada", "Long Island" and "Dublin Bay" are adjectives, because geese are not also "Canada", shellfish are not also "Long Island" and prawns are not also "Dublin Bay".
This kind of names is typical for English, but not for most other languages.
For instance, the scientific name of the Canada goose is "canadensis", which means "Canadian", not "Canada".
An adjective (in the broad sense) is a word that describes a subset of the set named by the noun to which it is attached.
While most languages also include distinct words that are adjectives in the narrow sense, i.e. which have degrees of comparison, adjectives in the broad sense (sometimes called relational adjectives) can be derived from any noun by various means, e.g. genitive case markers, prepositions, postpositions, suffixes, prefixes or accentual patterns, except for ambiguous languages like English, where any noun can also be used as an adjective, and sometimes also as a verb.
The only time I've ever even seen or heard of Canada Goose/Geese are people on the internet telling others they are wrong.
I think it's time to just accept it as correct.
Edit: TIL the tower was properly renamed "Elizabeth Tower" in 2012 [0] but I seriously doubt a single person in the last 12 years has ever used that name...
Is it though? Wouldn't we expect to see more advanced packaging technology eventually?
If that happens the increased memory bandwidth could be an enabler for a unified memory architecture like in the Nvidia Jetson line. In turn that would make a lot of what the article says make GPU go Brr today moot.
I found some datasheet that states 80GB of VRAM, and a BAR of 80GiB. All caches are also in power of two. The bandwidth are all power of 10 though.
https://www.nvidia.com/content/dam/en-zz/Solutions/gtcs22/da...
https://twitter.com/bfspector/status/1789749117104894179?t=k...
My experience with the AI boom couldn't be more different - everyone from my colleagues to my mum are using chatgpt as a daily tool.
I really don't think that AI and crypto are comparable in terms of their current practical usage.
At the peak of the crypto boom/hype cycle I took on a little project to look at the top 10 blockchain networks/coins/whatever.
From what I could tell a very, very, very generous estimate is that crypto at best has MAUs in the low tens of millions.
ChatGPT alone got to 100 million MAUs within a year of release and has only grown since.
ChatGPT 10x'd actual real world usage of GPUs (and resulting power and other resources) in a year vs ~15 years for crypto.
> I really don't think that AI and crypto are comparable in terms of their current practical usage.
A massive understatement!
In other words, it has nothing to do with AI.
I wasn’t talking about the applicability of GPUs for crypto, I was talking about adoption and real world utility.
GPUs in AI provide far more value and utility than crypto ever has regardless of how/if it’s mined with them or not.
Two points:
1. GPUs for AI != GPUs for crypto. They are not the same systems. There is a tiny bit of overlap, but most people bought underpowered gpus that were focused on ROI. For example, you didn't need more than 8GB ram. You also didn't need ultrafast networking.
2. Value is subjective.
At work setting in a tech company, there seems to be a handful that are very in love with AI, a bunch that use it here or there, and a large majority that (at least publically) don't even use it. It's be interesting to see what company enforced spyware would say about ai uptake though for real.
If I didn't know any better I'd consider it technobabble
GPU compute is already broken up - there is a supply chain of other cooperating players that work together to deliver GPU compute to end users:
TSMC, SK hynix, Synopsys, cloud providers (Azure/Amazon etcetera), model providers (OpenAI/Anthropic etcetera).
Why single out NVidia in the chain? Plus the different critical parts of the chain are in different jurisdictions. Split up NVidia and somebody else will take over that spot in the ecosystem. This interview with Synopsys is rather enlightening: https://www.acquired.fm/episodes/the-software-behind-silicon...
How does the profit currently get split between the different links? Profit is the forcing variable for market cap and profit is the indicator of advantage. Break up NVidia and where does the profit move?
George Hotz tried to get a consumer card to work. He also refused my public invitations to have free time on my enterprise cards, calling me an AMD shill.
AMD listened and responded to him and gave him even the difficult things that he was demanding. He has the tools to make it work now and if he needs more, AMD already seems willing to give it. That is progress.
To simply throw out George as the be-all and end-all of a $245B company... frankly absurd.
Also, if the consumer GPUs are hopelessly broken but the enterprise GPUs are fine, that greatly limits the number of people that can contribute to making the AMD AI software ecosystem better. How much of the utility of the NVIDIA software ecosystem comes from gaming GPU owners tinkering in their free time? Or grad students doing small scale research?
I think these kinds of things are a big part of why NVIDIA's software is so much better than AMD right now.
I’d say it simply dials it down to zero. No one’s gonna buy an enterprise AMD card for playing with AI, so no one’s gonna contribute to that either. As a local AI enthusiast, this “but he used consumer card” complaint makes no sense to me.
My hypothesis is that the buying mentality stems from the inability to rent. Hence, me opening up a rental business.
Today, you can buy 7900's and they work with ROCm. As George pointed out, there are some low level issues with them, that AMD is working with him to resolve. That doesn't mean they absolutely don't work.
https://rocm.docs.amd.com/projects/install-on-linux/en/lates...
One way to improve the flywheel and make the ecosystem better, is to make their hardware available for rent. Something that previously was not available outside of hyperscalers and HPC.
I didn't do that, and I don't appreciate this misreading of my post. Please don't drag me into whatever drama is/was going on between you two.
The only point I was making was that George's experience with AMD products reflected poorly on AMD software engineering circa 2023. Whether George is ultimately successful in convincing AMD to publicly release what he needs is beside the point. Whether he is ultimately successful convincing their GPUs to perform his expectations is beside the point.
Except that isn't the point you said...
"there's no hope of them becoming serious contenders in AI without some major changes in AMD's priorities"
My point in showing you (not dragging you into) the drama, is to tell you that George is not a credible witness for your beliefs.
My point is as I wrote in both posts. George was able to demonstrate evidence of poor engineering which "reflected poorly on AMD". From this I could form my own conclusion that AMD aren't in an engineering position to become "serious contenders in AI".
The poor software engineering evident on consumer cards is an indictment of AMD engineers, and the theoretical possibility for their enterprise products to have well engineered firmware wouldn't alleviate this indictment. If anything it makes AMD look insidious or incompetent.
Big ships take time to course correct. Look at their hiring for AI related positions and release schedule for ROCm. As well as multiple companies like mine springing up to purchase MI300x and satisfy rental demand.
It is only May. We didn't even receive our AIA's until April. Another company just announced their MI300x hardware server offering today.
That is obvious and the signs are there showing that they are working on it. What more do you expect?
You look at that and want to take a sledgehammer to a golden goose? I don't get these people
They saw there was nascent compute use of GPUs, using programmable shaders. They produced CUDA, made it accessible on every one of their GPUs (not just the high-markup professional products) and they put resources into it year after year after year.
Not just investing in the product, also the support tools (e.g. a full graphical profiler for your kernels) and training materials (e.g. providing free cloud GPU credits for Udacity courses) and libraries and open source contributions.
This is what it looks like when a company has a vision, plans beyond the next quarter, and makes long-term investments.
But yeah Vulkan is famously verbose. It takes about 1000 LoC to draw a triangle
Nvidia provides devs with great tools (Nsight Systems and Nsight Compute), so you know where you have to optimize.
also things related to graph topology in neural networks maybe, but probably not related to artificial NN.
I was given this video, which I found was pretty interesting: https://www.youtube.com/watch?v=nkdZRBFtqSs (How Developers might stop worrying about AI taking software jobs and Learn to Profit from LLMs - YouTube)
I doubt neuroscience will either, but I’m not as sure on that.
The more impressive AI systems we have moved further away from the neuron analogy that came from perceptions.
The whole “intelligence” and “neural” part of AI is a red herring imo. Really poor ambiguous word choice for a specific, technical idea.
The stuff on spiking networks and neuromorphic computing is definitely interesting and inspired by neuroscience, but it currently seems mostly like vaporware
Other than that, I agree, especially since you added “and other fields.” Psychology might eventually give us a useful definition of “intelligence,” so that’d be something.
Obviously all research can influence other areas of research.
https://www.technologyreview.com/2020/01/15/130868/deepmind-...
There are obvious, huge differences between what goes on in a computer and what happens in a a brain. Neurons can't do back propagation is a glaring one. But they do do something that ends up being analogous to back propagation and you can't tell a priori whether some property of AI or neuroscience might be applicable to the other or not.
The best way to learn about AI isn't to learn neuroscience. it's to learn AI. But if I were an AI lab I'd still hire someone to read neuroscience papers and check to see whether they might have something useful in them.
There are many reasons to hate nvidia, but honestly if this UBC policy is even remotely being considered in some circles, I'd join Linus Torvalds and say "nvidia, fuck you".
That’s the whole logic.