The case for memory-mapped GPU assets
facebook.com
facebook.com
To make it work fully end-to-end, it would look something like this:
1. Shader samples from a sparse texture, and detects that the requested page is non-resident.
1b. Fall back to lower mip map level.
2. Shader uses atomic-or to write to a "page fault" bitmap (one bit per page)
3. The bitmap is transferred to the CPU
4. For each set bit, start async copy from disk to DMA buffer (ie. pixel buffer object in GL)
5. When disk i/o is complete, start texture upload from buffer to a "page pool" texture
6. When texture upload is complete, re-map the texture page from "page pool" to the actual texture
Now this approach works alright, but there are a number of issues that make it impractical for the time being. Off the top of my head: 1. Sparse textures are only supported on Nvidia and AMD hardware. Not Intel, ARM or IMG.
2. Requires Vulkan or D3D12 for step #6 (the demo doesn't do this, so there may be pipeline stalls)
3. One or two frames of latency that can only be avoided if this was done in the kernel mode driver.
4. Poor fit for existing KMD architecture (which has its own concept of residency)
5. Detecting page faults is easy. Detecting which pages can be dropped is hard.
Here's the source code of my demo. It's not pretty because it was a one off demo project for very specific hardware (Android + OpenGL 4.5, which means Nvidia Shield hardware with Maxwell GPUs). The technique is portable, though.https://github.com/rikusalminen/sparsedemo/blob/master/jni/g...
(In the code above, all the interesting bits are the functions named xfer_*)
Based on the experience from writing this demo, I have to agree with Carmack here. File-backed textures would make a lot of sense for a lot of use cases.
Splash screens and loading bars vanish. Everything is just THERE.
I'm not sure I agree with this. It might be more convenient to have a filesystem-like interface, but at the end of the day everything still has to be loaded into the (rather limited) GPU memory at some point.Most CPU applications can handle RAM swapping from disk, but I really doubt that big games could maintain 60fps if even a few assets needed re-loading.
If a frame is 16ms, and the best consumer SSDs are around 2GB/s, You can only load 33MB of assets in a one frame in a best-case scenario.
[1] As an individual. I suspect why we haven't see as much from him since Quake 3 is that his skills don't scale to a team.
He did other stuff like rocket design and all that VR/Oculus Gear stuff
Not quite accurate. Since Quake 3 he made Unified lighting and shadowing (for Doom 3) and Megatextures (for Quake 5 and Rage).
I can't really think of any other major paradigm shifts over the last 10-15 in game graphics engines.
There are lots of little improvements everywhere, but they seem more evolutionary than the revolutionary steps we saw in the 90s (textureMapping, raycasting, BSPs, PVS, lightmaps, Sorted edge rasterization, bilinear filtering, bezier-surferaces, shaders, unified lighting/shadowing and beam trees, etc.).
From the consumer's point of view, all they ever did was have bad AMD compatibility and streaming issues on every platform.
Without globally unique textures Megatexture goes from a paradigm shift to just an implementation detail, unfortunately.
Again, the game was unfortunately a few fun sections interspersed with really un fun things, but that has nothing to do with the tech.
(FWIW I ran it on AMD and it had major problems, and then a year or so later it was perfect. AFAICT they patched most of the issues away on PC)
Phrased this way, this question makes no sense to me. Perhaps you mean specifically the art in Rage. Virtual texture memory strictly increases the capacity for detail, realism, and/or expression.
Deferred shading is a pretty major paradigm shift compared to forward rendering and it's within ~10-15 year range (it wasn't practical on pre MRT hardware which was shader model 3 ?) along with plenty of screen space effects like SSAO.
Linux ppl realized this and implemented MAP_POPULATE but once you're doing that you might as well just eagerly populate the normal way.
Mmap() can be a convenient interface when you aren't latency sensitive but otherwise it's not appropriate.
It's just a lot better than current status quo. Loading more or less full scene assets to GPU RAM regardless whether they're actually present in current frame or not.
Even if some texture is visible, it might be only a tiny fraction of some particular mipmap level is actually needed.
TLB flush cost is pretty insignificant here. A bit like accidentally crushing your finger with a sledgehammer and complaining the hammer was a bit cold. Besides, you'd be talking about TLB flush cost on the GPU. Maybe GPUs can just hide all of that latency with a high number of hardware threads, just like they've hidden RAM latency for over 10 years.
Sure you can figure out if a texture is used in the current visible set (like in a certain part of a game level). But then it starts to get tricky!
Any concrete ideas how to determine which parts (= memory pages) of a texture are actually needed to draw the scene?
If you don't know, you're going to waste a lot of I/O and memory capacity for something you don't need in the first place.
Remember a texture... :
1) ... has multiple mipmap [1] levels. Say you have 1024x1024 texture. You'll also need 512x512, 256x256, 128x128... etc versions. Depending on the triangle orientation [2], GPU might need some spatial areas (not all!) from any of those levels. 1024x1024 for the corner that's near camera and 128x128 for the faraway portion.
2) ... has spatial data order. It's not in row-major order (think scanlines), but in some space filling curve order (like Z-order curve [3]). This helps caching AND paging schemes -- other accesses are spatial and very likely to be nearby memory addresses in the texture.
3) ... is rarely drawn alone. There are often multiple objects using same texture. Sometimes that's true even if actual visible portion is completely different, like texture atlases [4]. This makes it non-trivial to consider points 1 and 2.
4) ... is sometimes huge. Uncompressed 4096x4096 float32 RGBA texture takes 256 MB (4096 * 4096 * 4 * 4) memory just for the first mipmap level. All (traditional "pyramid" type) mipmap levels together would take about 341 MB.
So how are you going to determine which memory ranges of the texture really need to be in the memory?
[1]: https://en.wikipedia.org/wiki/Mipmap
[2]: https://en.wikipedia.org/wiki/Anisotropic_filtering
A very simple solution is to add another logic layer to GPU processing, like a shader. It could be the "asset shader," it would run on each frame and tells the GPU what assets to start preloading, etc. That would give the programmer tight control over asset loading latencies without having to load everything at once.
How does that shader gain information about which portions of the texture are needed at each mipmap level?
Or do you just load whole texture and consume memory you don't actually need to render the image. It'd perform badly due to unnecessary I/O causing a long loading time. You'd also waste large portions of GPU RAM.
Or does your shader try to guess? Do you attempt to reverse engineer exactly how GPU trilinear texture sampler operates, because otherwise you won't know which parts of asset data is needed -- guess wrong and you get weird artefacts, when GPU samples your texture at a memory location you didn't load. Oops. Rounding differences compared to hardware texture sampler would almost certainly get you. Not sure if it's still true, but at least in past different brand/model GPUs implemented texture sampling slightly differently [1], enough to force you to have a version for many GPU vendors and models.
Or do you disable trilinear sampling and use just one mipmap level you somehow pick. You'd get bad image quality, blur and/or aliasing (like moire).
Even after considering all that, how are you going to deal with texture atlases?
Your way sounds really complicated. Unless you're willing to do rather serious compromises. Robustness, quality or loading time performance.
[1]: http://hexus.net/tech/reviews/graphics/747-nvidias-geforce-6...
This description in theory would be much much much smaller than the corresponding assets required to render the scene. It would be trivial to eagerly bind to the shader.
The asset shader API could provide information on the current GPU's texture rendering parameters/quirks.
Slight caveat, Windows 10 rarely pages out to the disk.[1] I'm not sure if it's possible to ask it to treat your mmaps in this way. Regardless, implementing the synchronization required to pull this off would be a nightmare - especially in Vulkan/DX12. The OS would also need some form of API where you are notified that a mmaped page is faulting. Except for the very top studios (AAA) it would likely be a completely unapproachable API. Still, it would be fascinating to see what the very best could do with it.
> You can only load 33MB of assets in a one frame in a best-case scenario.
Carmack indicates that "you could still manually pre-touch media to guarantee residence". Meaning that we're back to loading screens (hopefully shorter ones, though).
[1]: https://channel9.msdn.com/Blogs/Seth-Juarez/Memory-Compressi...
Back the GPU mapped resources with RAM. I have 16GB of it. That's plenty for one level's worth of assets.
GPU <=> RAM <=> SSD.
Obviously, there's the x86/x64 divide there - hopefully nobody is buying new 32-bit systems anymore, and that limitation can go away - although I'd rather get a x64 version of Visual Studio, but apparently that's not a good idea, because reasons[1]
[1] https://blogs.msdn.microsoft.com/ricom/2015/12/29/revisiting...
One game I recently played, a racing game, when you select "restart race" it reloads the entire level (3d models, textures etc) when really all it had to do was reset a handful of variable (car positions, time, a few other things like that).
If its not possible to reset these variables, then take a snapshot!
Its especially annoying when a game, eg, autosaves before a boss battle and then when you die, instead of just resetting some stats and inventory, you have to reload everything... They already know there's a good chance you will die (that's why they autosave!). If its too much effort to reload just the bits that need it, then take a snapshot of non-graphics memory at the autosave point and just reload that.
There's too many games out now where you can die very quickly (within seconds) if you're not that good yet, and then have to sit through a multi-minute loading phase over and over... very frustrating.
(obviously this applies only to "reloading", not terminating the game and loading)
I mean, yeah, when a game does allow bypassing load screens, it is pretty amazing bonus to game play: its the key to success in games like Super Meat Boy, to constantly maintain flow through failure and difficulty. But its hard.
If profiling show you're limited by pagefaults, then preloading by simply touching the memory will have it ready. It simplifies things, which is almost always a good thing.
I think page faulting can work a lot better for graphics assets (= textures) than for general purpose computing, because the system has special knowledge about data and can avoid processing pauses by (temporarily) using lower fidelity mipmap levels.
1) All of those assets aren't present in the current frame. 33 MB might cover the whole scene anyways.
2) Texture aware page fault mechanism can transfer low lower detail mipmap levels first. Two detail levels lower mipmap is just 1/16th of data, but is still going to be good enough for those few tens of milliseconds until better LoD can be loaded.
3) So in 3 frames (5ms) you have 100 MB. This probably covers current scene pretty well.
So everything will just appear to be there. Eyes won't have time to focus in time it takes to load all full quality assets for current scene, even if your SSD can load just 500 MB/s.
Remember, with page fault mechanism you don't need to load all assets, just those that are actually visible in the current scene. So initially there's less data to transfer.
But it doesn't change anything. Lower LoD mipmap level is going to be less data with or without compression. Each lower level data size is just 25% of previous.
33 MB of assets per frame is more than enough for a lot of practical tasks. Modern texture compression achieves 2 bits per pixel (e.g. ASTC 8x8) with very little percievable loss of quality (when you apply texture mapping, lighting, effects, etc. ASTC 12x12 gets 0.89 bits per pixel, but noticeable loss of quality). At this rate, the 33 MB/frame is 16k x 8k worth of texture data, every frame. Vastly more than there are pixels on screen.
And this figure does not take into account the block cache at all. If the data has been used recently (several minutes depending on RAM size), a copy of it might be still around in RAM which would make it nearly instantaneous and almost a gigabyte per frame of data (assuming all memory bw is put to this use).
As for the framerate issues, it would be imperative that this technique is implemented in a stall-free way (which is possible with Vulkan/D3D12) so that a steady 60 fps is maintained. All textures should have at least a few mip map levels resident at all times to fall back on. A single "standard" GPU sparse page (64k) is 512x512 texels at ASTC 8x8, which is quite a large texture already.
In other words: 33 MB per frame is a lot more than what we get today.
Even if you read from RAM during normal operations you'll have very low framerate unless you are using pretty damn good latency masking a good example will be texture streaming where you hold low res textures in your GPU memory and load the higher res ones from RAM or disk and even on very high end systems it causes allot of texture popins which people find rather annoying.
Perhaps disk I/O DMA can even be directly connected to GPU DMA, bypassing system RAM entirely. Even currently, ethernet and disk DMA can bypass memory and go directly to CPU L3 cache. With existing flexibility like that, it's not far-fetched to be able to connect DMA between two devices.
Pretty much same happens when you touch a page that's not present from user mode. CPU will get interrupted and queue disk I/O to get the faulting page.
Even disks themselves have their internal queues (like SATA NCQ). Disks also have CPUs to serve queued I/O requests.
Handling GPU IRQ's is possible but has unnecessarily high latency and pipeline stalling problems. However, the kernel mode drivers do a bit of this behind the scenes (mainly when switching from app to app).
But using sparse textures (ie. GL_EXT_sparse_texture2), the GPU can detect if a fault would occur and react to it. Instead of IRQ'ing for every missed page, a list of all missed pages can be extracted at the end of a frame.
> Perhaps disk I/O DMA can even be connected to GPU DMA, bypassing system RAM entirely.
Afaik Nvidia's NVLink does this but it's aimed at super computers, not graphics.
Time it takes for GPU to access data present in its local internal GDDR RAM: about 1 microsecond. Time it takes from GPU asserting (pulls up) [1] IRQ until CPU handles IRQ: 5-50 microseconds. Time it takes for CPU to insert I/O request into OS internal queue: unknown, but it's almost certainly a small fraction of a microsecond. Hundreds of faults can be handled with one IRQ request.
So there's not really that much more latency than in a normal CPU page fault. Actually probably less on average, because these faults can be grouped. GPU memory access latency is very slow and they are already built to handle high latency with a massive number of hardware threads.
IRQ can for example be edge triggered by the first GPU page fault. Each GPU page fault doesn't need to trigger CPU side interrupt. That'd be pointless, it'd just cause high CPU load for no benefit.
I think some sort of page fault FIFO is simpler hardware wise than to wait for something arbitrary event like "end of a frame".
GPU can just keep pushing new faults in a CPU-visible FIFO. That way it's possible to amortize latency and to avoid IRQ storm (=excessive number of IRQ requests).
CPU side can just pull faults from the FIFO at whatever rate I/O system can support.
Fault data from GPU could also include priority, so that more visually important data can be fetched first even if it wasn't encountered first. For example geometry or textures covering major parts of screen could have high priority.
(I've co-designed a HW mechanism for avoiding an excessive number of IRQs without sacrificing latency and designed and implemented a kernel mode driver to support it. Of course GPUs are quite a bit more complicated than that simple case.)
[1]: Well, technically it's sending an MSI(-X) message. Just using traditional terminology for clarity. https://en.wikipedia.org/wiki/Message_Signaled_Interrupts#MS...
See my other comment above. I implemented a technique like this (on a <10 watt mobile device!) and I had no issues whatsoever with framerate, even my suboptimal technique (using OpenGL sparse textures, which are quite restrictive) that might stall. These stalls can be avoided with Vulkan/D3D12 techniques (using aliased sparse pages, ie. one physical page mapped on several textures) so maintaining a stable framerate simply isn't an issue.
This was done with the CPU orchestrating the whole deal. The GPU isn't required to have access to disk or initiate the DMA transfer.
Latency, on the other hand, is an issue but it doesn't seem to be that bad in practice (warning: anecdotal evidence). In my simple demo, practically every page uploaded was resident on the GPU by the next frame after it was required. A small number of pages (less than 1%) had 2 frames of latency. None had 3 or more.
Hiding the visual artifacts from texture popping is a real issue too, but can be mitigated by speculatively uploading pages and applying filtering between mip map levels.
All of this can be done today using Vulkan or D3D12. What Carmack is suggesting (if I understood him correctly) is to make the kernel driver on the CPU initiate asynchronous upload without userspace intervention, which would improve latency and bandwidth.
Is it so regarding VR? How can you predict a direction and angle where a user turn his head to?
Asset loading latency doesn't affect rendering latency, put simply, it just means you get to play sooner, with less wait.
But does it have to be this way?
I don't think it's necessary for applications that require so much memory you can probably develop some proprietary tech to access storage (NVIDIA does it).
JC isn't talking about compute tho, he's talking about gaming he really likes megatextures but he's in a minority and I honestly can't claim to have enough expertise to judge if he's right or not.
GPU memory management is a mixed bag there are many cards with different speed of RAM, different bandwidths all doing the magic voodoo in the background regarding low level memory access and compression way beyond what you are exposed to on the driver and API layers (e.g. NVIDIA's incard memory compression which is enabled regardless of how you compress or load your textures in the first place).
WDDM allows lower end system to take advantage of more video memory to improve performance in less demanding applications which are most apps today while still enabling all the nice graphics we come to expect from out window manager and applications.
I'm not sure if allowing the GPU to read directly from the SSD or RAM (outside of the current scope of virtualized GPU memory and asset loading) has actually enough benefit to justify both the engineering costs and the potential security pitfall that can happen when you run multiple applications that all of a sudden get DMA access to your RAM and local storage. This is more so the case for gaming considering the amount of RAM that graphic cards come with today and this is only increasing I don't know if anyone but JC actually wants/needs this.
I am wondering how do consoles handle it, IIRC from playing a bit with the UDK (UE3) at the time I recall that it supported streaming textures directly from the optical media on the Xbox 360, so if anyone knows how it works under the hood and is willing to share I think it would add to this topic.
The bandwidth is also not a problem, the entire memory of a 16GB GPU can be filled in under a second.
To get 60fps your video card needs to push out a frame every 16ms. If say 6ms of this is actual in card processing and 10ms is CPU/API overhead and you add to it another CPU call + system memory access that adds 30-35MS latency you are now operating at ~50ms per frame which means that you can only output 20 frames each second.
How come? If it's on the PCI bus it should be able to bus master its way into the system RAM, no?
Or was this capability disabled to prevent bad graphics code from being able to exploit the system or something?
Besides there's a small detail called IOMMU that prevents PCI(-e) devices from writing or reading wherever they please.
You can pull assets from RAM somewhat transparently already (if you want to do it well you need to optimize it further, but to some extent NVIDIA and AMD do quite a bit of optimization in the driver too), shared page table between the CPU and GPU is also already in place under WDDM.
So adding a disk based page file to this won't be so hard, it could more or less work the same way as the page file currently works.
I'm still not sure if this is actually needed, GPU's already come with stupid amount of RAM today, i have 24GB of GDDR5 in my system, and even mid range cards today will not come with less than 6GB of RAM.
Can GPU page table entry point to non-present page(s)? Or does it only work for "pinned pages" [1], that cannot paged out of RAM?
If it doesn't require pinning, then what prevents from mmapping assets on the disk even today?
If, however, it does require pinning, then John Carmack has a damn good point.
[1]: https://en.wikipedia.org/wiki/Virtual_memory#Pinned_pages
WDDM also limited the volume and commit sizes based on some "arbitrary" limits that MSFT set out (IIRC it's something like system memory / 2 or something silly like that), there's a bit more silliness that for example if the limit of the max memory available for graphics is 2GB you can't commit 3GB but you can do 3x1GB just fine.
I'm pretty sure atm WDDM/Vendor Display Driver pin all pages to RAM only so it won't end up in the page file, TBH I haven't had a page file on my system for a long long time windows doesn't use SSD's for paging unless it really has too and considering I haven't been using a system with less than 32GB of RAM for the past 5 years I never had issues with it.
Here is a somewhat decent MSFT doc about memory management under WDDM (this is not for WDDM2.0 so it's not that upto date I'm guessing) http://download.microsoft.com/download/9/c/5/9c5b2167-8017-4...
P.S. Apologies if I used any terminology incorrectly this is stretching both my knowledge and recollection regarding this subject, it's also more or less limited to how GPU/Graphics are handled within Windows.
I think it can only work through pinning, because it seems to rely on GPU bus mastering for accessing pages that are dedicated for graphics.
So you have 32 GB RAM and say 8 GB GPU RAM. Imagine you have a program that has 100 GB of graphics assets without built in mechanism to guess which subset of assets might be required for current scene. The program needs about 50 MB in any given frame -- in other words it has 50 MB working set.
As an operating system, which assets are you going to keep in memory?
Now imagine you have multiple programs running concurrently, each having 100 GB of assets. Each app has that same 50 MB working set.
How can the system handle this situation efficiently (or at all!) if the pages need to be pinned to RAM?
If you can have true on demand GPU paging, all of these apps need to only swap their current working set of data. User would not perceive any delay when switching from app to app. She could even display all of them in the same time without any issues.
If not, the system would grind to halt.
https://blogs.msdn.microsoft.com/greg_schechter/2006/04/02/t...
"In the event that video memory allocation is required, and both video memory and system memory are full, the WDDM and the overall virtual memory system will then turn to disk for video memory surfaces. This is an extremely unusual case, and the performance would suffer dearly in that case, but the point is that the system is sufficiently robust to allow this to occur and for the application to reliably continue."
The only issue here is that as far as i can understand you can't really choose how this is done very specifically.
When you allocate memory the driver pretty much takes over, if you want granular over memory allocation you pretty much have to go through the GPUMMU path which means each process has a separate GPU and CPU address space so while you can control what you store in GPU memory and what you store in System memory I still don't see a way to control mapping an asset to disk specifically other than it being an edge case of you running out of system memory which results in the page file being used.
When the OS runs out of RAM, it can swap pages to disk. If this page also has a mapping in an IOMMU, it can invalidate the mapping there as well.
Then, when the device attempts to touch the page, the IOMMU faults and the CPU would swap the page back in.
I'm not sure if this is possible, but this is the route I'd expect something like this to follow.
Interesting idea. I'd also like to know if IOMMU faults can be acted upon. It might require protocol support between the bus and hardware device (GPU) [1], to tell it the page is not currently present. And a way to tell GPU once the page is available.
As far as I can see this kind of mechanism would require one CPU interrupt per GPU fault. That might be too inefficient.
[1]: Edit: It's indeed possible to handle, if the device supports "PCI-SIG PCIe Address Translation Services (ATS) Page Request Interface (PRI) extension".
Yes, new GPUs allow you to do this. This feature is called sparse textures / buffers in OpenGL (GL_ARB_sparse_texture and GL_EXT_sparse_texture2) and Vulkan (optional feature in Vulkan 1.0 core) or tiled resources in D3D.
This allows you to leave textures (or buffers) partially non-resident (accesses to which are safe but results undefined) and allow you to detect when accessing a non-resident region (EXT_sparse_texture2) so that you can write a fallback path in the shader (lookup lower mip level and/or somehow tell the CPU that the page will be required for the next frame).
The OpenGL extensions are a bit restrictive, but Vulkan/D3D12 allow much greater control (such as sharing pages between textures or repeating the same page inside a texture).
Hardware support for this is not ubiquitous at the moment, but should improve as time goes on.
This feature is somewhat orthogonal to pinned pages and WDDM residency magic (which is afaik more oriented to switching between processes), hopefully it will get more unified in the future.
What I was getting at was the fact that there should be nothing physically preventing them from implementing DMA support, so I was wondering why they didn't already support it. I’m not familiar with GPUs, so I assumed that this is how all large transfers between system RAM and the GPU worked.
I can only speak for the Linux context, but the IOMMU isn’t an issue. You can just allocate memory with the DMA API (dma_alloc_coherent) which will automatically populate the IOMMU tables (if required), pin the pages, and return you a PCI bus address as well as a kernel virtual address which both correspond to the same chunk of physical memory. Or, you can map an existing buffer in page by page using the dma_map routines (I forget the names).
Now, you have a shared pool of memory which can be accessed by both devices at the same time. The coherency fabric (if one exists) will handle all synchronization automatically, though this can be a bottleneck sometimes. If the CPU isn’t cache coherent, then the pages get marked as no-cache in the kernel PTEs so that any read from the CPU side pulls straight from memory.
Passing “messages” can be accomplished by an external notification like an interrupt or something.
You can then even map this buffer into a user space program.
I’m sure there are some security concerns with this approach though.
More complicated things like device to device transfers (GPU to and from disk) would have to be arbitrated by the CPU, but I see no reason that the CPU would actually have to do the copy itself. Why couldn’t the CPU just provide the GPU with the PCI bus address of the disk controller which should be written to?
If the GPU wanted to write to the disk, the CPU would initiate a transfer to disk, but before writing the actual data, you pass the destination PCI address to the GPU and let it write the data. Then, the CPU can resume doing whatever it has to do while this happens in the background.
Just brainstorming here.
The GPU does have DMA to RAM
References:
https://stackoverflow.com/questions/19242711/how-can-i-use-g...
Because it requires both very specific software and hardware configurations only some features work on some systems.
http://docs.nvidia.com/cuda/gpudirect-rdma/index.html#suppor...
Also calling this DMA is kind of misleading it does require the CUDA software layer to work, and for the most part it's a lot of hacks coupled together to form a feature set.
So I would still currently stand by what I said that there isn't a standard and generic way for GPU's to have DMA access.
2560 * 1600 * 3 = 12,288,000 bytes.
1) Hardware doesn't usually support 3 byte pixels. Pretty safe to assume at least 32bpp pixels, 4 bytes per pixel.
2) Hardware might only support power-of-two line pitches (width, number of bytes between screen pixel rows). Horizontal resolution would still be 2560, remaining pixels would just be hidden. That way hardware can always rely on bit shifting when computing display addresses and possibly other tricks.
3) Frame buffer might not even be in row-major format in the first place.
So 2560 * 1600 32bpp might just as well be 4096 * 1600 * 4 = 25 MB. Or something else.
I have been advocating this for many years, but the case gets stronger all the time. Once more unto the breach.
GPUs should be able to have buffer and texture resources directly backed by memory mapped files. Everyone has functional faulting in the GPUs now, right? We just need extensions and OS work.
On startup, applications would read-only mmap their entire asset file and issue a bunch of glBufferMappedDataEXT() / glTexMappedImage2DEXT() or Vulkan equivalent extension calls. Ten seconds of resource loading and creation becomes ten milliseconds.
Splash screens and loading bars vanish. Everything is just THERE.
You could switch through a dozen rich media application with a gig of resources each, and come back to the first one without finding that it had been terminated to clear space for the others – read only memory mapped files are easy for the OS to purge and reload without input from the applications. This is Metaverse plumbing.
Not that many people give a damn, but asset loading code is a scary attack surface from a security standpoint, and resource management has always been a rich source of bugs.
It will save power. Hopefully these are the magic words. Lots of data gets loaded and never used, and many applications get terminated unnecessarily to clear up GPU memory, forcing them to be reloaded from scratch.
There are many schemes for avoiding the hard stop of a page fault by using a lower detail version of a texture and so on, but it always gets complicated and requires shader changes. I’m suggesting a complete hard stop and wait. GPU designers usually throw up their hands at this point and stop considering it, but this is a big system level win, even if it winds up making some frames run slower on the GPU.
You can actually handle quite a few page faults to an SSD while still holding 60 fps, and you could still manually pre-touch media to guarantee residence, but I suspect it largely won’t be necessary. There might also be little tweaks to be done, like boosting the GPU clock frequency for the remainder of the frame after a page fault, or maybe even the following frame for non-VR applications that triple buffer.
I imagine an initial implementation of GPU faulting to SSD would be an ugly multi-process communication mess with lots of inefficiency, but the lower limits set by the hardware are pretty exciting, and some storage technologies are evolving in directions that can have extremely low block read latencies.
Unity and Unreal could take advantage of this almost completely under the hood, making it a broadly usable feature. Asset metadata would be out of line, so the mapped data could be loaded conventionally if necessary on unsupported hardware.
A common objection is that there are lots of different tiling / swizzling layouts for uncompressed texture formats, but this could be restricted to just ASTC textures if necessary. I’m a little hesitant to suggest it, but drivers could also reformat texture data after a page fault to optimize a layout, as long as it can be done at something close to the read speed. Specifying a generously large texture tile size / page fault size would give a lot of freedom. Mip map layout is certainly an issue, but we can work it out.
There may be scheduling challenges for high priority tasks like Async Time Warp if a single unit of work can create dozens of page faults. It might be necessary to abort and later re-run a tile / bin that has suffered many page faults if a high priority job needs to run Right Now.
Come on, lets make this happen! Who is going to be the leader? I would love it to happen in the Samsung/Qualcomm Android space so Gear VR could immediately benefit, but it would probably be easiest for Apple to do it, and I would be just fine with that if everyone else chased them in a panic.
See also the comments from https://news.ycombinator.com/item?id=8704911
Considering that prefetching schemes allow the programmer to spread asset loading evenly over many frames, and cheap rendering approximations can be used in troublesome frames, there should also be enough low-level control.
My disks are usually encrypted though and sometimes I can choose faster or slower encryption methods (thus affecting throughput when loading). I don't see how this can work reliably without forcing the user to reserve specific disk areas just for GPU assets.
Seriously though, the company he works for is owned by Facebook. This may be a factor.
At the time it happened, many had already looked into the architecture and realized that to them there were no real benefits: 32-bit ARM could already address 1TB of memory, you could get the accelerated crypto instructions with an architecture extension and 64-bit ARM implementations were not very power efficient (a problem just recently solved by the latest A7x devices).
But when Apple switched to 64-bit ARM everyone else just had to follow along. This resulted in weaker Android phones for about a year. The funny thing is, the reason Apple switched so early is because they needed some head start since they use native apps. They didn't really needed the 64-bit ARM at that time yet.
I'm definitely no expert though, any inside knowledge on this?
Now, I am not saying that 64-bit was just fluff. The new design has a much nicer pipeline (specially thanks to the instructions they removed) which is MUCH better suited for things like speculative execution. But the implementation to make use of this wasn't really there until very recently. Here is a fun fact for you: the 64-bit Cortex-A53 and the 32-bit Cortex-A7 are 80% the same CPU. What does that tell you about the first generation of 64-bit devices?
So if I understand the problem right, it's because copying data to the GPU is made through the PCI express bus, and done "piece by piece", instead of larger batches ? A little like grouping draw calls ? That's funny how that problem can be seen everywhere in hardware, where multiplying queries will make latencies snowball.
Generally DMA to/from the GPU is cache coherent (either via DMA sniffing for cache invalidation or software managing regions for DMA, e.g. marking relevant PTEs as nocache).
So accesses are _coherent_, but the cache is simply irrelevant (or even more costly, if it's using snooping).
Been there, done that, doesn't work. You start a level and every corner you take invokes the hard drive. Chug city. New enemy, new textures, chug.
Horrible.
Source: worked at a GPU company. Saw the experiments.
GPUs churn through data once it is across the bus.
A heirarchy of GPUs, outputs wired to inputs, mirroring the heirarchy of deep nets would be useful for real time robots & cars, NVIDIA's other big market.
Nvidia introduced a new, faster SLI bridge for the new 1000 generation, aren't they used in GPGPU setups?
Isn't that exactly what they're doing with NVLink?
https://devblogs.nvidia.com/parallelforall/inside-pascal (NVLink High Speed Interconnect section)
If you want to publicly make your content available in a free and seamless way stop using facebook. Please.
I regularly use it e.g. for the cookie acceptance boxes. They are a major annoyance -- as I don't save cookies, the setting for accepting cookies isn't preserved either and I have to approve them again and again. If I visit any such page for a second time it's easier to just hide the whole question with uBlock.
Leave the asides aside please.
This allows both for the community to drive the discussion to what most interests the community, and to allow those with different tastes than the majority to still participate without alienating or frustrating them or those involved in said discussion.
I'm of the opinion that HN doesn't implement features like this on purpose, because they expect you to find a way to work around the limitations that works best for you, and that will likely result in a more customized and comfortable experience than them forcing their opinion of how the content should be presented on you. HN is somewhat notorious for having these soft limitations (such as not showing reply on very active threads for a few minutes, which also has an easy workaround for those that know).
(The other very useful comment at the top is "Warning, it autoplays a video")
Perhaps there ought to be some feature on the part of HN itself, to automatically recommend the better links, when possible.
Having everything the app needs available to the GPU at all times without having to explicitly load them from disk to graphics memory beforehand in userland code (and cycled out as needed) - having it happen automatically so less code needs to be written & debugged in the app/framework, and having it happen in the kernel so it is potentially more efficient (allowing for more complex scenes and/or more detail and/or faster frames, on the same hardware).