You basically only interact with the kernel on init/shutdown or outside of the fast path, and do something like isolcpus to delegate the kernel and interrupt handling to some garbage cores and give you the rest to do what you want with.
Perhaps [1] is a good resource to start with (page nr. 7). Example code is here [2]. And [3] makes some experiments with it.
[1] https://www.kernel.org/doc/ols/2006/ols2006v2-pages-83-90.pd...
[2] https://github.com/intel/iodlr/blob/master/large_page-c/exam...
[3] https://easyperf.net/blog/2022/09/01/Utilizing-Huge-Pages-Fo...
Finding out the beginning of the stack should be possible to deduce from memory addresses upon entering the main().
Finding out the end of the stack depends on how big the stack is and whether it grows upwards or downwards. Former is of dynamic nature but usually configured at system level on Linux and for the latter I am not sure but I think it grows downwards.
EDIT: Putting the code segment into hugepages will relief some of the pressure of VADDR translation which is otherwise larger with 4K segments. Whether this will impact the runtime execution time positively or stay neutral I think it greatly depends on the workload and cannot be said upfront.
Stack access is another story as it's usually local and sequential so it might not be that useful.
You can outright install the openonload drivers, preload to intercept epoll, and it literally just works.
Going to efvi can cut out ~1us but that requires specifically targeting efvi, and has more operational / code setup pain. Works on stock Linux all the same though
I am pinning the cores, disabled SMT and turbo boost, but haven't tried isolcpu because this has to reboot the computer.
The gold standard (in my experience) for latency measurement is setting up a packet splitting, marking your outbound packets with some hash/id of the inbound packet, taking hardware timestamps of all these on some dedicated host, and putting it all together after the fact. Ultimately packet-in to packet out is all the matters anyways
You can actually get pretty solid internal timestamps (how long did I take to fully process event X) with TSC counters, but you then have a coordinated omissions problem: https://www.programmingtalks.org/talk/how-not-to-measure-lat...
Also, PREEMPT_RT is the worst option for low latency because it's about execution time guarantees and not speed specifically. If you're on PREEMPT_RT and give your critical thread highest prio, be prepared for some serious OS-level lock-ups.
PREEMPT_RT includes priority inheritance, specifically to avoid this scenario. So your app should indeed be favored if you tune accordingly. What you also seem to be saying is that using PREEMPT_RT may lead to lower throughput, but that's not the same thing as latency.