AMD has put an SSD on a graphics card
techradar.com
techradar.com
The struggle comes in reading the data from storage, where several minutes can be spent loading in a single high resolution raster for analysis/display. When I built my own PC this year, I splurged on an M.2 SSD for the OS and my main data store. Best decision I ever made for my workflow - huge 3D scenes that formerly took minutes to load on a spinning platter now pop up in seconds.
This thing would probably be the bees knees for what I do. Shame it starts at $10k (and that it's AMD so no CUDA, so no way to justify it at work for "data science" :-/).
Edit: this is a good step though. AMD should be pushing the envelope and hopefully with Zen, they can actually realize some of the gains of HSA (which they tried to pioneer but it wasn't so useful since Bulldozer isn't that good)
In contrast, OpenCL could have been a wonderful vendor-independent solution for mobile, but both Apple and Google conspired independently to make that impossible (ironic in Apple's case because of OpenCL's origin story and idiotic in the case of Google and its dreadful Renderscript, a glorified reinvention of Ian Buck's Ph.D. thesis work, Brook).
Fortunately, AMD appears to have figured out OpenCL has no desktop traction and they have embarked on building a CUDA compiler for AMD GPUs called R.O.C (Radeon Open Compute). They have also shown dramatically improved performance at targeted deep learning benchmarks. It's early, but so is the deep learning boom.
The wildcard for me is what Intel will decide to do next.
The big win IMO is vendor-unlocking all the OSS CUDA code out there.
https://github.com/RadeonOpenCompute https://techaltar.com/amd-rx-480-gpu-review/2/
Also they accepted the world has moved on and CUDA accepted C, C++ and Fortran since day 1.
Shortly thereafter they made the PTX Assembly format available and any language could easily target it for GPU execution.
OpenCL got stuck in a world of C with the same runtime compilation model as GLSL.
Only after loosing to CUDA they cave in and created SPIR and the SysCL C++ implementation.
And they have done it again with Vulkan.
While Metal and DX are object based, even the shader languages are C++ like, Vulkan is all about C.
Recently they decided to adopt NVidia's C++ wrapper.
And the reason why it's C is so that you can bind from multiple languages. DirectX doesn't have this problem because it's not plain C++, it's COM--which is designed to support bindings from multiple languages. But COM is effectively Windows only, and Vulkan needs to be platform independent. So Khronos made the correct choice here.
As you acknowledged, they also created an idiomatic C++ wrapper, so I don't see what your complaint is at all. Khronos did the correct thing every step of the way.
Besides the big studios already abstract the graphic APIs on their in-house engines anyway, or they outsource to porting studios.
Usually the HN community, which is more focused on web development and FOSS, seems to miss the point that the culture in the game industry is more focused on proprietary tooling and how to take advantage of their IP.
The whole discussion around which API to use isn't that relevant when discussing about game development proposals.
It is more akin to the demoscene culture, where cool programming tricks were shown without sharing how it was done, than the sharing culture of FOSS.
If NVidia hadn't made their C++ wrapper available, I very much doubt Khronos would have bothered to create one of their own.
> The whole discussion around which API to use isn't that relevant when discussing about game development proposals.
OK, so if the graphics API doesn't matter, why did your parent comment participate in the graphics API war?
(BTW, I agree with you that the graphics API doesn't matter too much anymore. But I think if you're going to attack Vulkan, you should do so based on specific technical reasons.)
> If NVidia hadn't made their C++ wrapper available, I very much doubt Khronos would have bothered to create one of their own.
Khronos isn't a company—it's a standards body. As NVIDIA is a member of Khronos, "Khronos" did bother to create a C++ API.
I rather use APIs that embrace OO, offer math, font handling, texture and mesh APIs as part of SDKs instead of forcing developers to play Lego with libraries offered on the wild.
> Khronos isn't a company—it's a standards body. As NVIDIA is a member of Khronos, "Khronos" did bother to create a C++ API.
I guess that is one way of selling the story.
But it's not "pure C". C is just the glue. You can use the C++ API if that's what you want, and you have a completely object-oriented API.
Can Linux never have an "OO interface" because all syscalls trap into a kernel written purely in C? Of course not, that would be silly. The same is true here. If you program against an object-oriented C++ interface, then you have a fully object-oriented API.
I find interesting that a Rust designer thinks otherwise.
You can easily call into something with a C ABI (er, OS ABI designed around C, or whatever is the correct technical name) from any language. Try that with C++ :D
C++ provides better tools to write safer code than C ever will.
I would rather be using Ada, Rust, System C#, D, <whatever safe systems programming language>, but until OS vendors start providing something else, C++17 will have to do.
Since 1994 I only write C code when forced to do so.
AMD is trying to fix it, but also now "too late", their drivers, that frankly, are crap, on all platforms, AMD drivers are very, VERY buggy.
Also, AMD GPUs need excessive amounts of power, at first I didn't considered this a problem, until I bought my AMD GPU and noticed it is constantly throttling and causing stutter even in professional software, due to power limits, it also bit AMD in the ass during the RX 480 launch (where the excessive power usage went beyond the motherboard limits, and their "Fix" was make the driver instead request power beyond the specification of PSU cables instead, or allow users on Windows, to enable a harder power limit, making it throttle even more).
I had hope AMD would "scare" nVidia into improving, into stopping their shady business practices and just improve their business, but after actaully buying AMD product, and interacting with their crappy support, crappy community, crappy distribution network (it was very hard to get the card!), I concluded that AMD has a loooong way before they make nVidia react, AMD is too far behind in all aspects, and the only reason they are still competitive, is because they sell very power-hungry beefy GPUs for cheap prices, achieving a reasonable performance per dollar, but if you compare their products ignoring that, they are just junk (both in the hardware and software sense).
The demo claims that without the SSDs they were rendering raw 8k video @17 fps and using the SSDs improved rendering to > 90fps. How can this be such a significant improvement over accessing the same SSDs connected directly to a motherboard? The graphics card would have a PCIE 3.0 x16 connection...plenty of bandwidth and very low latency.
Maybe I'm missing something?
[1] - http://www.anandtech.com/show/10518/amd-announces-radeon-pro...
"The performance differential was actually more than I expected; reading a file from the SSG SSD array was over 4GB/sec, while reading that same file from the system SSD was only averaging under 900MB/sec, which is lower than what we know 950 Pro can do in sequential reads. After putting some thought into it, I think AMD has hit upon the fact that most M.2 slots on motherboards are routed through the system chipset rather than being directly attached to the CPU. This not only adds another hop of latency, but it means crossing the relatively narrow DMI 3.0 (~PCIe 3.0 x4) link that is shared with everything else attached to the chipset."
I should clarify that I do think this onboard SSD concept could be really compelling for certain use cases, such as needing to store several hundred gigs of data which needs to be randomly accessed.
In contrast, the route with the on-board SSDs is that the GPU makes the request, the ASIC that handles NVMe requesting from the memory chips, the memory chips sending to the ASIC, which goes into memory. There are a lot fewer steps there, and a lot fewer places where delayed interrupts, etc can introduce lag.
(CPU) WHILE video not done
1 - Initiate DMA from SSD -> main memory with some large (say 128MB) chunk of 8k video
2 - Signal GPU to begin DMA transfer
3 - Wait for interrupt from GPU
(GPU)
WHILE video not done 1 - Wait for CPU to indicate data is ready
2 - Initiate DMA from main memory -> local GPU memory over PCIE 3.0 x16 link
3 - Issue interrupt to CPU
4 - Render chunk of videoIt would have been really cool to have been able to swap it out as needed, but I don't think that will be doable :(
http://www.anandtech.com/show/10518/amd-announces-radeon-pro...
And on the back you can see the standard tiered screw holes for fitting different M.2 card lengths:
"Instead of adding more expensive graphics memory, why not let users add their own...The Radeon Pro SSG features two PCIe 3.0 M.2 slots for adding up to 1TB of NAND flash, "
http://arstechnica.co.uk/gadgets/2016/07/amd-radeon-pro-ssg-...
This would solve that problem nicely.
i.e. the transaction goes from this: GPU -> CPU -> HDD -> GPU
to this: GPU -> HDD -> GPU
Oh and the fact that nobody write a specialised enough driver for a GPU to use this.
That said, onboard SSD is cool if it can read without pre-empting the PCIE bus.
1. GPU driven command buffer workloads will become more significant (now that literally all your textures can exist in VRAM, it's worth going the extra mile to elide a GPU-CPU round trip).
2. Voxel based techniques (which were generally memory intensive) open up again. That's modeling destructible environments, atmospheric effects, transparent materials, etc in a more accurate and performant way.
3. Entire scene graphs can live on the GPU which opens up a lot of design space for new volume hierarchies and data structures.
3D-NAND uses different physics, but as of now hasn't yet matched legacy NAND's endurance (lifespan)
What makes you say that?
3D-NAND uses charge-trap flash while planar NAND uses floating-gate-transistors based flash.
Working in a NAND company also makes me say that.
Lets assume a relatively modest 8bit per component RGB, a 7680x4320 frame is 132MB, at 24fps is 3.2GB/s.
So in 1TB you can store 312 seconds aka 5 mins of video.
You're going to have to top this up to play any more than 5 mins, which means writing to the SSD at 3.2 GB/s whilst you read from it. I do not believe even any NVMe SSDs are full duplex, so that's 6.4GB/sec, which the SSDs cannot provide.
You can DMA to GPU memory at 10GB/sec, so the SSDs are of no benefit for video.
No-one stores 8K data uncompressed anyway, so the CPU will need to read and decompress the data.
Perhaps it's useful for faulting in megatextures for 3D scenes though.
People have already moved data-base computation to the GPU, providing low-latency NVM access to such a set-up would really accelerate things.
I know that PCIe bandwidth is not really an issue and SSD latency is significantly higher than RAM (but with NVMe pushing below 3 µs it makes you wonder how low a custom interface could go..)
I doubt that the same GPU was used. If it is, we're in trouble (or have unoptimized software). I thought that someone would have already solved the "let's stream data in and out and not interrupt whatever the GPU is doing" problem, but meanwhile many people are fretting over what PCIe version (thus, attained speeds) they can use.
A GPU's 16 lanes of PCIe 3.0 provide almost 16 GB/s of bandwidth. An M.2 SSD (used in this card) is at best 4 GB/s (PCIe 3.0 x4 lanes). This card has two M.2 slots.
Even at those speeds, accessing data from system RAM is faster. (depends on specifics, but system RAM is easily 40+ GB/s) How does increasing data bandwidth 50% increase performance by 4? Are inefficiencies (and latencies) really eating up all that?
You're quoting this but I'm not sure I understand what problem you're referencing. Streaming data to the GPU isn't a problem of interrupting it. GPUs use memory fences, atomics, and other things you might find in CPU land. Why are we in trouble if the demo was true? For every type of hardware, I can come up with a rendering algorithm and data set that makes it look better than similar hardware in a different configuration. If you have more memory bandwidth, you should be using it accordingly.
Then they'll integrate ethernet into video card for high speed traders (also can be marketed as lower latency for gaming). But why not to integrate whole PC into it? With gaming console-style OS on it.
If I was rewriting the same-ish REST helper on every micro-service of a project, I would immediately think of refactoring and yet GPUs now have CPU/RAM/Disk.
Is the future GPU-only or are we in dire need of a better way to use existing PC components from GPUs?
Wait for the Vega cards with 8GB of HBM2 — that will kick ass.