Intel's Ponte Vecchio Xe-HPC GPU Boasts 100B Transistors
tomshardware.com
tomshardware.com
https://www.nvidia.com/en-us/data-center/a100/
I'd guess, by the time Intel actually ships anything useful, Nvidia will have it made it mostly moot.
https://cdn.mos.cms.futurecdn.net/BUsZ5EdKUcP8mWRKypTNB4-970...
When excluding ML... (because that's what Intel gave actual metrics on for Xe-HP)
41Tflops FP32 with 4 dies. For comparison, an RTX 3090 (arch whitepaper at https://www.nvidia.com/content/dam/en-zz/Solutions/geforce/a...) has 35.6Tflops FP32, with a single die.
The only way Intel reaches those performance targets is to outdo the current crop of GPUs: MI100 (from AMD) and A100 (NVidia).
Not that that is a guarantee that we got a winner here, but we know the goal for what Intel is shooting for at least.
"Over 100 Billion Transistors" "47 Magical Tiles" "Alchemy of Technologies" "Our Most Advanced Packaging" and also "X Ponte Vecchio".
I haven't cringed this hard at marketing material since the iPad Pro ads.
If anyone at Intel is reading this, please consider releasing all Ponte Vecchio drivers under a permissive open-source license; it would facilitate and encourage faster adoption.
https://github.com/intel/compute-runtime
https://github.com/intel/intel-graphics-compiler
https://github.com/torvalds/linux/tree/master/drivers/gpu/dr...
CUDA is polyglot, with very nice graphical debuggers that can even single step shaders.
Something that the anti-CUDA keep forgeting.
NVIDIA understood also early on the importance of first party libraries and commercial partnerships something Intel also understands which is why OneAPI has wider adoption already than ROCm.
.NET, Java, Julia, Python (RAPIDS/cuDF), Haskell don't have a place on OneAPI so far.
And yes, going back to C++, the hardware is based on C++11 memory model (which was based on Java/.NET models).
So plenty of stuff to catch up, besides "we can do C++".
Most non C++ frameworks and implementations tho would simply use wrappers and bindings.
I also am not aware of any high performance lib for CUDA that wasn’t written in C++.
https://developer.nvidia.com/blog/hybridizer-csharp/
"Simplifying GPU Access: A Polyglot Binding for GPUs with GraalVM"
https://developer.nvidia.com/gtc/2020/video/s21269-vid
And then you can browse for products on https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Cent...
And again it’s a commercial product developed by a 3rd party, whilst someone uses it I wouldn’t even put it as a rounding error when accounting for why CUDA has the market share it has.
Or forgetting the days when games written in C were actually full of inline Assembly.
It is still CUDA, regardless if it goes through PTX or CUDA C++ as implementation detail for the high level code.
The market for these secondary implementations is tiny, and that is coming from someone who worked at a company that had CUDA executed from a spreadsheet.
The C#/Java et. al isn’t what made CUDA popular nor what would make OneAPI succeed or fail.
CUDA became popular because of its architecture, using an intermediate assembly to allow backward and forward compatibility, it had exe Elle to support across the entire NVIDIA GPU stack which means that it could run on everything form bargain bin laptops with the cheapest dGPU to HPC cards. It came with a large library of high performant libraries and yes the C++ programming model is why it was adopted so well by the big players.
And even arguably more importantly is that when ML and GPU compute exploded and that wasn’t that long ago NVIDIA from a business perspective was the top dog in town, CUDA could’ve been dog shit but when AMD could barely launch a GPU that could compete with NVIDIA’s mid range for multiple generations it wouldn’t have mattered.
This is really the only point to be made. Intel could release open source GPU drivers and GPGPU frameworks for every language under the sun, personally hold workshops in every city and even give every developer a back massage and everyone would likely still use CUDA.
The performance gap is still so large.
The vast majority of CUDA applications don’t need 100’s of HPC cards to execute, consumers want their favorite video or photo editor to work, they want to be able to apply filters to their Zoom calls, students and researchers want to be able to develop and run POCs on their laptops as long as Adobe and the likes adopt OneAPI and as long as Intel will provide a backend for common ML frameworks like Pytorch and TF (which they already do) performance at that point won’t matter as much as you think.
Performance at this scale is a business question if AMD had a decent ecosystem but lacked performance they could’ve priced their cards accordingly and still captured some market share. Their problem was that they couldn’t actually release hardware in time, their shipments were tiny and they didn’t had the software to back it up.
Intel despite all the doom and gloom still ships more chips than AMD and NVIDIA combined if OneAPI is even remotely technically competent and from my very limited experience with it it is looking rather good Intel can offer developers a huge addressable market overnight with a single framework.
And when Khronos woke up for that fact, alongside SPIR, it was already too late for anyone to care.
Regarding the trees, I guess my point is that regardless of tiny they are, the developers behind those stacks rather bet on CUDA and eventually collaborate with NVidia than going after to the alternatives.
So the alternatives to CUDA aren't even able to significally atract those devs to their platforms, given the tooling around CUDA to support their efforts.
I would really want to be able to find the people at AMD who are responsible for the ROCm roadmap and ask them WTF were they thinking...
"Alternatively you can let the virtual machine (VM) make this decision automatically by setting a system property on the command line. The JIT can also offload certain processing tasks based on performance heuristics."
A lot of what ultimately limits GPUs today is that they are connected over a relatively slow bus (PCIe), this will change in the future, allowing smaller and smaller tasks to be offloaded.
The toolkit ships with compilers from C++ and Fortran to NVVM, and provides you documentation about the PTX virtual machine at https://docs.nvidia.com/cuda/parallel-thread-execution/index... and about the higher-level NVVM (which compiles down to PTX) at https://docs.nvidia.com/cuda/nvvm-ir-spec/index.html.
A CUDA binary from 10 years ago will still run today on modern hardware, ROCm breaks compatibility between minor releases sometimes and it’s often not documented.
But it's one way only, NVIDIA ships a disassembler, but explicitly doesn't ship an assembler.
Just as an example I just bought an M1 laptop. Even PyTorch CPU cross-compiling hasn't been done properly, and TensorFlow support is full of bugs. I hope that with M1X it will be taken more seriously by Apple.
I understand your point (Julia is a great example), but trying to support a rarely used hardware with a rarely supported framework is just not practical. (Stick with CUDA.jl for Julia :) )
What the point of using OneAPI, a yet another compute API wrapper, to make software just for a single platform?
You can just use regular computing libs, and C, or C++.
Serious HPC will still stay with its own serious HPC stuff, superoptimised C, and fortran code, no matter how labour intensive it is.
So, I see very little point in that.
Wether it would be successful or not is up in the air but it’s goals are pretty solid.
The big data ecosystem is Java-centric.
https://www.nextbigfuture.com/2020/11/cerebras-trillion-tran...
Anyone know how this affects power, compute, or communications metrics compared to monolithic designs?
Or am I off in thinking this approach maximizes yields?
It does effect ennormously, but everything is highly design specific.
The die size limits are not only yield related.
Power, and clock have stopped scaling few generations ago.
New chips have more, and more disaggregated, independent blocks separated by asynchronous interfaces to accomodate more clock, and power domains.
If you have to break a chip along such domain boundary, you loose little in terms of speed unlike if you did it right across registers, logic, and synchronous parallel links.
Caches also stopped scaling too, and making them bigger, also makes them slower.
Instead, more elaborate application specific cache hierarchies are getting popular. L1-2 get smaller, and faster, but L3 can be made to ones fantasy: eDRAM, standalone SRAM, stacked memory etc.
>Or am I off in thinking this approach maximizes yields?
This is absolutely the case here especially for HPC or high performance market where they are all previously served by large die size solutions. You can also use the best node for each tiles / chipsets. Example I/O die may use other node for lower power consumption while Compute die uses High power node to push performance. High Density node for SRAM, low cost node for simple function accelerator etc. You can now customise each part of the tiles/chiplet best according to your target power/performance/cost equation.
Now of course their are trade offs, Intel Tiles packaging (EMIB, Embedded Multi-die Interconnect Bridge) is expensive. And it adds tiny latency for not being on die. Although you could argue you saved latency for certain parts with it not being off a PCIE bus.
But this is very exciting. The best part is with this announcement that you can use Tiles and Chiplet from other foundries on Intel packaging! You could theoretically use Global Foundry ( or Samsung ) FD-SOI node for certain applications and still be on the same packaging. You can now truly mix and match to the maximum.
Not so good part is we are still limited by TDP design. And I dont see anything being done in the consumer / prosumer space.
But when not a single AMD processor is available and Intel processors are, that's just logical. Later on AMD processors were available, but way more expensive than before, with Intel selling nice options like the i5-10400F at their lowest price ever. That price difference in essence still continues today. Right now at that price point the Intel offering is unbeatable for AMD.
Price elasticity of demand for AMD is partially dependent upon the price of Intel chips. There will always be pricing pressure, even if there is scarcity of AMD chips, because to some extent AMD and Intel chips are substitute goods.
Very small laptops with Intel Tigerlake are on level with AMD and Apple products. They have all the new IO bits (PCIe 4, LPDDR4x, Wifi6) and low power usage on 10nm.
If you wanted a bit more battery life, performance, or just want to try a fancier display upgrading could be nice.
We are long past the days of buying a new computer every couple of years.
Too many components from too many different sources, with intel doing the "integration".
Doesn't this remind anyone of the engineering philosophy of the Boeing 787 Dreamliner? Have individual manufacturers build component parts and then use just in time integration to put assembly and packaging at the end. If any individual manufacturer runs out of chips or components, or de-prioritize production (for example, if Samsung or TSMC is being ordered by Korea or Taiwan to specifically prioritize chips for their automotive industries) - this could lead to shortages that will cause ripples down the assembly line for these xe-hpc chips.
Especially in today's world, when companies like Apple are constantly moving toward vertical integration, and bringing in all external dependencies inward (or at least have ironclad contracts mandating partners satisfy their contractual duties), this move by intel is in the wrong direction in the post Covid-chip shortage era.
wow