Virtual Memory Tricks
ourmachinery.com
ourmachinery.com
IIRC, it attempts to use end-of-page allocation by default for blocks of a certain size. It is also possible to enable guard pages on a per process basis or globally.
Thanks to this, memory corruption is much easier to detect and debug on OpenBSD.
https://msdn.microsoft.com/en-us/library/ms220938(v=vs.90).a...
IIRC, OpenBSD is slightly more sophisticated than issuing one syscall per call to malloc(3) and friends. Just slightly, but still. ;-)
Still, I do not like expressing disagreement by downvoting. Sorry about that.
Any reads / writes to that virtual memory address work in both CPU and GPU code. So the CPU can do highly sequential tasks like balancing a binary tree (with all of the pointers involved), while the GPU can do highly parallel tasks like traversing the tree exhaustively. Pointers inside of this shared region remain legitimate, because each address is correctly identified between the CPU and GPU.
See this for example: https://software.intel.com/en-us/articles/opencl-20-shared-v...
I'm describing "Fine Grained Shared Virtual Memory" btw, in case anyone wants to know what to search for.
In effect, it works because the CPU does an mmap, and then the GPU does an mmap, and then the PCIe drivers work to keep the mmap'd region in sync. With modern acquire / release memory fence semantics + atomics properly defined, its even possible to write "Heterogeneous" code where the CPU and GPU works simultaneously on the same memory space.
In theory. I haven't tested these features out yet, and its quite possible that current implementations are buggy. But the features exist and its totally possible (to attempt) to do this sort of stuff on hardware today.
I found this recent CppCon video pretty interesting. It's about running vanilla C++ on (nvidia) GPUs and what the implications are for the C++ memory model:
The address space is not actually 64 bit and it's not just the OS limiting it. The CPU/MMU/IDK (I'm not low level enough to get all this stuff) does too in some ways. I assume these values mean that I can have 2^39 bytes of actual RAM and 2^48 bytes of virtual address space:
>[frex@localhost ~]$ cat /proc/cpuinfo | grep address
>address sizes : 39 bits physical, 48 bits virtual
>address sizes : 39 bits physical, 48 bits virtual
>address sizes : 39 bits physical, 48 bits virtual
>address sizes : 39 bits physical, 48 bits virtual
This also came up back when I was looking through PUC Lua 5.1 code for a paper and (out of curiosity) looking into code of PUC 5.2 and 5.3 to see the changes and into code of LuaJIT for comparison.
LuaJIT uses top bits of a pointer to store type info in case of light userdata pointers: https://www.freelists.org/post/luajit/Few-questions-about-Lu...
This reminds me I should probably do my write up of how PUC Lua 5.1 VM works and implements it's features in English finally. If someone's interested in any way I'd appreciate dropping me a word to let me know that there's market for that so I don't fear that I will be writing down stuff I already grok for no one to read and benefit from.
Yes. There’s no point having 64 physical address lines if 26 of them aren’t going to be used by anything
What is inconceivable today, will be reality at some point.
And I vaguely remember that ESA/390 is actually 32bit address architecture in hardware and it's 31bitness comes from using highest bit of address as a flag that such tricks are taking place.
On Linux (and POSIX in general), you can do the same thing by allocating the guard page as described, then setting its protection to PROT_NONE using mprotect(2). Any access to the page will cause an error.
>Since Windows will incorrectly assume that our int 3 exception was generated from the single-byte variant, it is possible to confuse the debugger into reading "extra" memory. We leverage this inconsistency to trip a "guard page" of sorts.
* http://wiki.osdev.org/Page_Tables#Recursive_mapping
* https://medium.com/@connorstack/recursive-page-tables-ad1e03...
When you attempt to access the memory at one of those locations for the first time though, you will experience a performance hit because your program will page fault due to the mapping for that address not actually existing yet. The kernel will then allocate and map a new page and then return to your program so you can execute the instruction again. After that first access there is no other performance penalties.
The unique ID thing may get expensive if you need tons of them, but also since each address within a page is its own unique ID, you get a few thousand unique IDs per syscall, which should cut down on the number of syscalls you have to do. You could also allocate larger blocks of addresses if you wanted too, rather then just a single page.
Not every object needs a unique identifier. Objects that need them tend to be larger and have indeterminate lifecycles. Otherwise the overhead of storing a unique id would by itself be egregious, regardless of the process of generating the id. If you can afford to store the id then you can afford the overhead of an occasional syscall.
You may wind up slowing leaking vmas (the kernel needs to track what’s been mapped, even if it doesn’t allocate physical pages). This makes future mmap calls and page faults slower.
Generating a uniqueID is trivial and super fast (one atomic operation). What’s the benefit here?
also I think that double mapped ring buffer implementation is bogus without a volatile keyword in there somewhere
I agree the ring buffer is problematic. The `read` function for the `ring_buffer_t` is a bad design - it returns a pointer into the actual ring_buffer_t buffer, but also increments the read pointer. There's no guarantee for how long the returned pointer is actually valid - eventually writers will overwrite it and you will have no idea. I suppose if you're only looking at a single-threaded environment though then it's not as a big of a deal, but for the normal implementation "read" should just do a `memcpy` like `write` does.
As for `volatile`, making the `data` pointer `volatile` should make this valid - though you can't use `memcpy` after that, so you have to do the loop yourself unfortunately.
That said, I think you could "cheat" (IE. don't do this) and ensure the correct behavior by changing `ring_buffer_t` to have two `data` pointers to the buffer instead of one - one for reading, and one for writing. The compiler would have to assume the two pointers can alias each-other at any point within the buffer, and thus it should force the compiler produce correct code even with multiple reads and writes in the same function (Because when you do a write, it has to assume data it previously read may have changed). I don't think there's any guarantee though - a sufficiently smart compiler could possibly make inferences that would break this. In particular, the C standard says that the C memory layout is completely linear, so the compiler could make an assumption like "If address 1 aliases with the first byte of the other buffer, then address 4097 doesn't" and that would obviously not be true. Unless you hard-coded the addresses though, I don't think any current compilers would ever make such an optimization - but again, nobody should do this, I don't know for sure. It's more of a curiosity then a good idea.
I'd be really curious why it was removed myself.