(The kernel thread was actually provide by a system call from user space. So there was a daemon as a user space program, but that just called into the kernel, where it then sat in a loop. This is nice: you have something in user space that you can kill with SIGTERM or whatever and restart by running some /usr/bin/executable.)
If the daemon terminated, that didn't destroy any of the buffers; they would just sit there getting backlogged. New ones could be created, too (via ioctl calls on a character device in /dev).
One aspect was that get_user_pages doesn't give you a linear block of memory. I wrote the grotty C code to dequeue messages from the non-linearly-mapped circular buffer. The message headers could straddle page boundaries, so I copied those out into temporaries before working with them. I think there was some attempt to avoid making unnecessary copies of the bulk data.
Open question: would it have helped to have a CMA[1] region managed by a kernel module and mapping that memory/regions of that memory to user-space using remap_pfn_range[2] instead?
[1] CMA documentation file: https://lwn.net/Articles/396707/
[2] https://www.kernel.org/doc/htmldocs/kernel-api/API-remap-pfn...