> IR-level volatile loads and stores cannot safely be optimized into llvm.memcpy or llvm.memmove intrinsics even when those intrinsics are flagged volatile. Likewise, the backend should never split or merge target-legal volatile load/store instructions. Similarly, IR-level volatile loads and stores cannot change from integer to floating-point or vice versa.
Even in a baremetal/OS context, the compiler is generally allowed to split, combine, reorder, and eliminate memory access however it likes as long as the end result is consistent. Being able to do that is very important for performance and code size, which especially matters in embedded or OS dev -- for example, inserting memcpy calls can save on code size for things like passing a medium-sized struct to a function. The memory model for any reasonably low-level programming language nowadays requires you to specify whenever you actually need memory access that look exactly like you wrote.
The OP suggests:
> I know that this is a long shot, but it would be nice if there was some way to specify how memory in certain address ranges need to be addressed.
The downside to this approach is that the compiler doesn't usually know what address range an arbitrary pointer will acccess; it's not resolved until link-time or often runtime. So this would require dynamic checks around every single memory access, which would kill performance. Thus, the solution is to use volatile loads and stores for any memory ranges that need special treatment.