Use OpenGL to get more RAM
github.com
github.com
Motherboards only support directly accessing a (up to) 256 MB segment of VRAM directly from the CPU. This is the BAR1 space/aperture space.
So attempts to create allocations larger than that that are resident in VRAM but also accessible from the CPU will land in system memory, in order to ensure they are accessible from the CPU. Graphics drivers will either have the GPU read the data from system memory, or will do hidden copies to the GPU when they detect the resource is bound, etc.
The author sort of suspects this could be happening:
"There is no guarantee that the persistently mapped buffer technique actually references video memory. The worst case it's shadow memory and this actually wastes memory."
Is this a 32-bit BAR / low 4GB space issue?
https://www.pugetsystems.com/labs/articles/Will-your-motherb...
Bottom line, most desktop machines only use a legacy 32-bit PCIe window below 4GB. That is why its limited to a fairly small amount of address space (256-1G generally). The ugly problem is that without purchasing a product its generally quite difficult to determine if it supports a 64-bit aperture correctly. Take the Asus x99A board I have, no mention of support one way or the other, but with an i7 (not just a xeon) it actually works with the xeon phi.
(BTW: Nvidia did a better job with this, their tesla cards have a GPU compatibility mode where they restrict the BAR to 256MB for machines that don't support 8GB BARs).
So I created a ramdisk on the old machine, exposed it as a nbd (network block device), mounted it on over the network (one hop over a 100 Mbit switch, didn't have Gbit kit) on the new machine and created a swapfile on it. I couldn't believe how well it worked. I never actually had any flakiness issues and I used it like that for probably a year.
https://wiki.archlinux.org/index.php/swap_on_video_ram
http://ft23.pmenier.net/docext/mtd/TIP_Use_memory_on_video_c...
This may be the very first occurrence of the idea on Linux:
http://web.archive.org/web/20081227185259/http://hedera.linu...
It was surprising slow, though still faster (and quieter!) than HDD.
"Extended memory" drivers were a common thing back then but that was to circumvent the 640KB/1MB limit of DOS and they had nothing to do with video card memory. In that era few people even knew how much memory was on their video card because it was rarely a concern, most of the time it was just barely enough to hold a single frame buffer, typical of most VGA and "Super-VGA" cards.
Extended memory is that which is beyond the boundaries of the original IBM PC design: https://en.wikipedia.org/wiki/Extended_memory
It's not video card memory in any way. If I recall correctly the video card memory was often mapped into the 640KB-1024KB memory zone, but using that for system memory was futile since you'd be unable to see anything on screen.
Not confused, just not talking about PCs in general (although I have heard cases of this happening there but that might be just indeed demo groups); I still do this on those old computers today for hobby purposes (the last work I did was somewhere end 90s on a legacy system which counted on video memory and which had to be ported so I had to remove & test that 'feature') ; there just isn't another way to get memory if you don't want to cheat (and solder in 1024kb).
Edit: I don't know if it was used in practice because there were so many different VGA cards, but [1] shows that it was common to have 256kb on a card which, depending on the era, was significant. And like said above, for business applications in text based environments (screen 0) you can use a large portion of video memory for storage while you still see something. But I can see the issue there; with other computers we didn't have that issue as you could trust a certain amount to be there for sure.
Unfortunately, they couldn't turn it on universally; there were a lot of programs that "knew" that malloc() would never return a pointer with the top bit set, so they used it as a flag for their own purposes. Such programs would break horribly in 3GB mode when pointers could have the top bit set. Also, restricting kernel-space meant the kernel had less room to do stuff; if you were running an app that was heavy on kernel-resources but light on RAM usage, 3GB mode would be a downgrade[1].
So applications have to opt-in to 3GB mode by setting a special flag in the executable, and 32-bit versions of Windows had to be put into 3GB mode at startup to support it at all.
Of course, the real solution to "3GB mode at startup" is just to run a 64-bit OS, and the real solution to "2GB is not enough address-space" is to run 64-bit applications.
[1]: https://blogs.technet.microsoft.com/askperf/2007/03/23/memor...
Fascinating story. Never knew this. This dedication to backwards-compatibility is something Microsoft was known for. Reminds me of how they took great care when writing Windows to ensure perfect compatibility with old, horrible DOS programs, going so far as re-creating DOS bugs that these broken programs relied on. I suppose they were just accepting reality: if a user upgraded to Windows and some third party program broke, the user would blame Microsoft, not the incompetent application developer.
Windows has broken legacy programs though. In Windows 7 a lot of old programs stopped working in the consumer editions of the OS. I don't know if they've continued making their "professional" editions that provide comparability. I suppose that's a reasonable strategy to encourage developers to update their software without killing user productivity, which could workout for the best in the long run.
As for Windows, best i can tell the problem was Win16 on 64-bit Windows. My father ran into that regarding some old solitaire programs he had been keeping around.
"Just for comparison, Linux defaults to 3 GB for user mode and 1 GB for kernel mode. There are experimental patches to implement a separate virtual address space for the kernel, thus granting 4 GB for user mode AND 4 GB for kernel mode. Unfortunately, there is a 10-20% performance hit because every system call requires an address space switch (and TLB cache flush)."
Take it with a grain of salt, that quote is a decade old.
https://pdfs.semanticscholar.org/4489/18f01d17628cb94ac69f27...
They don't call it "Slowlaris" for nothing...
The flag in the PE header says that the application can handle addresses over 0x80000000. It don't change reservations.
It may appear that they use less than 4GB if you look at the working set or private bytes figures instead of virtual size because they reserve more address space than they allocate. Or it might crash before reaching that limit due to large allocations failing.
But 64bit builds have been available for years on their ftp server, they just weren't supported. But since last december they are[2], so there really shouldn't be any need to run 32bit builds on a 64bit system.
[1] https://bugzilla.mozilla.org/show_bug.cgi?id=556382#c34 [2] https://blog.mozilla.org/futurereleases/2015/12/15/firefox-6...
Maybe the 64-bit version doesn't support the Adobe Flash plugin?
The another problem with 64-bit versions is that all the pointers are twice as big. So if you have only 8 GB of physical ram, you won't necessarily be able to handle much more than with a 32-bit application. But with 16 GB or more the 64-bit version must be better.
64bit flash works for me.
> The another problem with 64-bit versions is that all the pointers are twice as big.
You're assuming pointers dominate the footprint. if you're seeing crashes due to address space exhaustion that's more likely due to large strings/blobs of data and not a huge amount of pointers.
Or to put it differently, do you really think OOM crashes are preferable to some background applications or memory-mapped files being paged out?
No, I just wrote:
>> if you have only 8 GB of physical ram, you won't necessarily be able to handle much more than with a 32-bit application
That means, I absolutely know applications that won't be better on a 8 GB RAM computer as 64-bit applications, and they were written using standard C++ libraries. Typically, such applications weren't optimized for cases when they use a lot of memory. If somebody competent spends enough time and energy, even such application can be made to use memory better but it's not something that "just happens."
And sometimes 32-bits application is a good trade-off. See for example
https://blogs.msdn.microsoft.com/ricom/2009/06/10/visual-stu...
"A 64 bit address space for the process isn’t going to help you with page faults except in maybe indirect ways, and it will definitely hurt you in direct ways because your data is bigger."
What I'm trying to suggest is that the topic is more nuanced as "64 doublegood 32."
Under 4G worth of size pointers are simply zero extended, so that the higher order bits are just zero and the lower order bits stores in 32 bits.
Above 4G but below 32G it switches to "compressed oops" mode. This right-shifts the address by 3 bits and ensures/mandates that all objects are aligned to 8 bits (2^3). When converting back again to a memory address it multiplies by 8. This effectively multiplies the top address space by a factor of 8 and so you can store an object in the same 32 bits but can be wider.
It's possible to shift the alignment (-XX:ObjectAlignmentInBytes=16) which will then right shift the address by 4 instead giving you 64Gb memory with 32 bit addressing.
Although it's possible to go higher the advantages tend to scale out; effectively you end up padding the object space more and more so simple objects take up larger and larger slots.
The bottom line is: by controlling where pointers are used, Java can use a different addressing scheme by unpacking and repacking object pointers on the fly. And for some operations (equality testing) it doesn't even need to unpack the pointers in the first place.
VRAM is also where many "online" development suites—the BASICs and Pascals of the time—expected you to write exception traces. Rather than trying to "break into" a debugger (i.e. cram a debugger into the same address space as your software), you'd simply have your trap-handler persist the stack-trace to VRAM, switch your tape drive over to the devtools disk, and reboot. The "monitor" would load, notice that there's a stack-trace in VRAM, parse it, read the pages it mentions from your program (now on the slave tape) and display them.
The bottom line is, that the OpenGL driver will create a backing store in system memory for each and every buffer object. So if you allocate 4GiB of OpenGL buffer objects, it will allocate 4GiB of system memory; if it's a shared memory GPU that's it, if it's a dedicated GPU the GPU RAM is actually more of a cache to the backingstore.
glMapBuffer will usually just give you access to this backingstore so that writes can be coalesced into a single transfer when unmapping. Also you don't want a full round trip read-modify-write. In case of coherent mappings the coalesced transfer is triggered by a GPU side data read operation on that buffer.
TL;DR: RAM on graphics cards is a cache on top of system memory (as far as OpenGL is concerned).
That's a so negligible minority of devices (underutilized gaming rigs and nothing else) that OSes just don't bother.
No smartphones, no tablets, only a small minority of notebooks (not a single macbook, e.g.), none of the business desktop PCs (which are by far the majority, volume-wise), no servers nor workstations (except those that have them installed for GPGPU purposes).
The exception are gaming laptops/PCs (and those will need the VRAM for actual gaming) and servers/workstations with dedicated GPGPU cards (which, again, would need the VRAM for GPGPU purposes).
That only leaves a few cases where you have beefy hardware and don't use it, which isn't significant enough for OS vendors to bother with. Thirdparty VRAM-as-RAM (or RAM disk) kernel modules / FUSE layers pop up from time to time (this isn't nearly the first), but they never gain any traction.
> I guess those systems correlate with high-end systems with lots of system RAM anyway though...
Usually VRAM is less than a quarter of the system RAM, or even less. And usually those systems have so much system RAM that they have enough spare to make RAM disks an a valid option.
GPU malware is going to become really interesting once memory-coherent GPUs (like PS4's APU) become more available. That combined with something like packet mmap [1] may allow for data exfiltration from a GPU compute kernel without ongoing userspace assistance (MMIO packet send/recv => no need to have the CPU userspace side dispatch syscalls on behalf of the GPU).
[1]: https://www.kernel.org/doc/Documentation/networking/packet_m...