The History of CUDA
youtube.com
youtube.com
The main message is about the motivation behind CUDA: People don't want to learn a completely new language, they want to invest as little as possible. So that means, have just C but on GPU. The motivation of CUDA was to make GPU programming as easy as possible for someone who knows already C programming.
This is the reason I am bullish on Mojo.
not sure if it was a slip and corrected in post ... they are Nvidia afterall so they have all compute they want :).
But if it was indeed corrected in post then they did an excellent job of getting the acoustics / ambient noise perfectly right. Can someone do forensics on the audio track to see any editing artifacts?
PS: this bit is at timestamp 0:12 though
PS2: Channel is "Nvidia Tesla" and as someone commented on youtube he looks a bit like Elon Musk :)
From the 2004 article: "It is also possible that future streaming hardware will share the same memory as the CPU, eliminating the need for data transfer altogether." Unified memory foreseen 16 years before Apple Silicon (and the point of this comparison is to indicate how hard is to go from a prediction in a paper to mass manufacturing/popularity, not that Apple invented unified memory).
[1] http://graphics.stanford.edu/~ianbuck/thesis.pdf
[2] https://www.cs.cmu.edu/afs/cs/academic/class/15869-f11/www/r...
More seriously people need to stop with the Apple comparisons. Unified memory has been a thing for a way longer time. Heck around 2014 AMD had integrated GPUs with not just unified memory but fully unified address spaces with the host. Unified memory in itself happened way before that.
Not to mention that mobiles have always been unified archs. It’s just a design decision.
Honestly, it's good to get some more background information before claiming that Apple invented every innovation till sliced bread.
Microcontrollers, SoCs from various vendors, gaming consoles and Intel CPUs with integrated graphics also had unified memory since .. forever(?), or at least nearly 30 years, because it was as efficient back then as it is now for silicone and SW usage.
Apple didn't reinvent the wheel in this regard, it was already there as a low hanging fruit.
In his 2004 PhD, he tells ATI had a much better performance than nvidia... even if it was still the case, it would not even matter as their tools and drivers are terrible.
[1] "The display section contains its own memory, leaving all of RAM for user programs", http://s3data.computerhistory.org/brochures/apple.applei.197...
[2] "Running Apple 1 software on a breadboard computer (Wozmon)", https://www.youtube.com/watch?v=HlLCtjJzHVI
What Apple appear to have with their M2 chips is shared memory meaning that the CPU and GPU are directly accessing the same memory chips. On the just-announced M2 Ultra chip they are claiming 800GB/sec memory bandwidth, which compares well to the 1TB/sec on a recent NVIDIA card.
Unified memory, at least as NVIDIA use the term, only refers to a unified address space such that the GPU and CPU (located on opposite sides of the PCI bus) can use the same address space to access memory. However, the memory being mapped to by this unified address space may be on either side of the PCI bus (i.e be CPU memory or GPU memory) and may migrate from one side to the other to optimize performance. Given how slow PCI bus transfers are compared to GPU memory bandwidth, the use cases for this is not at all the same as true shared memory... It's really just a developer convenience feature to not have to explicitly orchestrate CPU-GPU memory transfers yourself (which you may be better off doing to maximize performance).
NVIDIA seem to be going in the same direction as Apple here, with their latest designs integrating GPU and CPU on a single module.
https://en.wikipedia.org//wiki/SGI_O2
When SGI's viability became questionable, I always thought there might be some value in Apple scooping them up for innovative bits of value like that but that never came to pass. Would be interested to know if it was ever considered/rejected and why.
There are also AMD versions of PyTorch and TensorFlow.
The problem is that all of these efforts are not quite 100% there ... there are bugs and incompatibilities that seem to make most people abandon them. It's a shame since the hardware itself seems great.
The most bang for buck they could do to improve their competitiveness (and share price)
That is a stark contrast to Nvidia where everything works on even the most entry level GPU, encouraging adoption in third-party software.