To properly handle all cases means forcing a memory fence whenever you are done writing code. I haven't thought too much about which fence is proper, but at minimum a write fence needs to ensure all data is written out as expected. You might need a full fence.
EDIT: Maybe a sfence at the location of the "writing" thread, while the JIT Code needs an lfence (to ensure all caches are properly updated). IIRC, fences on both sides are needed in these cases.
The "dynamic linker" idea seems better to my instinct, because it won't require these fences. Of course, testing would be required (Instincts are typically proven wrong in performance questions). IIRC, Intel Skylake and AMD Zen perform branch prediction on dynamic addresses, so you should get okay performance out of the dynamic lookup.