See this page (about half way down) https://ruby0x1.github.io/machinery_blog_archive/post/virtua...
See this page (about half way down) https://ruby0x1.github.io/machinery_blog_archive/post/virtua...
This clever ring buffer has hidden costs that don't show up in microbenchmarks testing only the push() and pop() operations on an initialized ring buffer. These can be very high and impact other software as well.
The initialisation requires multiple system calls to setup the aliased memory mapping and it has to be a multiple of the page size. This can be very expensive on big systems e.g. shooting down the TLBs on all other CPUs with interprocessor interrupts to invalidate any potentially conflicting mappings. On other end small systems often using variations of virtual indexed data caches and may be forced to set the aliased ring buffer pages as uncacheable unless you can get the cache coloring correct. And even for very small ring buffers you also use at least two TLB entries. A "dumber" ring buffer implementation can share part of a large page with other data massively reducing the TLB pressure.
Abusing the MMU to implement an otherwise missing ring addressing mode is tempting and can be worth it for some use-cases on several common platforms, but be careful and measure the impact on the whole system.
* https://web.archive.org/web/20140625200016/http://libh.slash...
VRB:- These functions implement a basic virtual ring buffer system which allows the caller to put data in, and take data out, of a ring buffer, and always access the data directly in the buffer contiguously, without the caller or the functions doing any data copying to accomplish it.
* The "Hero" code
http://libh.slashusr.org/source/vrb/src/lib/h/vrb_init.c
* Front & back end source directories:
https://web.archive.org/web/20100706121709/http://libh.slash...
//-- Limit request to available data.
if ( arg_size > vrb_data_len( arg_vrb ) ) {
arg_size = vrb_data_len( arg_vrb );
}
//-- If nothing to get, then just return now.
if ( arg_size == 0 ) return 0;
//-- Copy data to caller space.
memcpy( arg_data, arg_vrb->first_ptr, arg_size );
If at the time of the check for the amount of available data there is less data available than requested but additional data gets added to the buffer before arg_size gets updated, then this might get more data than requested and overflow the target buffer. At least vrb_read() and vrb_write() have the same bug.Concurrency issues existed in the days of yore .. but they arose in different ways with different timings.
It's an odd bit of code archeology recalling a concept ( mmap ring buffers ) from decades past and then hunting to find the best remaining example - much of LibH was 'clean' rewrites of code from the authors past work.
Otherwise, yes, that's what I mean by the OS bending over backward.
Linus has choice words for such architectures.
All modern caches PIPT, VIPT, or even VIVT must work with this scheme as it's semantically transparent to virtual memory. The performance of line-crossing accesses is a totally different issue and I would never used this "trick".