On a multithreaded implementation, you can do it with 2 stub functions and an array/hash of function pointers (hereto refered to as the "table") that itself acts as a spinlock array. Initialize the table with stub 1 and atomic write stub 1 to the table on eviction.
stub 1) upon entry, atomic write stub 2 into the table. Then compile or enqueue onto a compilation thread/process. At the end of compilation, atomic write the newly compiled address to the table.
stub 2) spinlock on function pointer array entry != Stub 2 address. At end of spinlock, call new function pointer. If you care about power consumption or timing attacks, the spinlock can have a backoff corresponding with the speed of your JIT compiler. This function should be able to fit into a single cacheline, so it should usually be hot.
inlined thunk: Given a function, read address from table and call it.
The client uses the thunk to call the function address in the table without knowing if it's compiled or not. On Intel, you shouldn't need a read fence due to MESI rules.
The branch/fetch predictor works by hashing the address of the branch instruction into its own lookup table. So by forcing stub 3 to be inlined, you're relying on the hash to prefetch the correct function address and call it directly.
You should also be able to put the JIT into a separate process that mmaps with the NX bit set, and have the client load without the NX bit set. So it would be difficult to exploit the JIT compiler itself (but the produced code could still be unsafe).
You might want to cachealign the function table or pack similar chunks in the same function table entry to prevent false sharing whenever there's an atomic write. Also, this table is going to thrash your cache by incurring an extra cacheline load for the trampoline that it does every time.
Anyways, that's my back of the napkin design for how to do it.