But re-entrancy might be a problem ... e.g. if some crazy application writer has called fork() from a signal handler, and that happened right between those lines. For a case like this, I think moving the guard to post-generation (but pre-use) and using an atomic CAS option to have the guard effect the return value of the generator, can do the trick at the cost of having to very occasionally throw generated data away. But I haven't tried it on all platforms.
With a thread-local RNG (and a thread-local page which gets zeroed on fork) it sounds like you're probably best off extracting data from the RNG and then checking the flag to see if you need to go back and try again. The only cost of doing it this way is a very marginal extra cost whenever you need to regenerate entropy -- but you can think of that as just being a slight increase in the cost of the fork.
The only other question which comes to mind is whether you need to worry about breakage on systems which have weaker-than-x86 cache coherency. I'm 99% certain that you're fine, but I'd prefer to be 100% certain.
I assume you'd like to see this functionality added to other platforms as well?
A rng-per-thread combined with initialization check before and after might cover all the cases I think...
I'm not sure what you mean by "arc4random isn't marked signal safe" -- you mean it can't be used in any process which receives signals?