> One aspect was that get_user_pages doesn't give you a linear block of memory. I wrote the grotty C code to dequeue messages from the non-linearly-mapped circular buffer. The message headers could straddle page boundaries, so I copied those out into temporaries before working with them. I think there was some attempt to avoid making unnecessary copies of the bulk data.
Open question: would it have helped to have a CMA[1] region managed by a kernel module and mapping that memory/regions of that memory to user-space using remap_pfn_range[2] instead?
[1] CMA documentation file: https://lwn.net/Articles/396707/
[2] https://www.kernel.org/doc/htmldocs/kernel-api/API-remap-pfn...