It does make a difference of course if you're running fetch_max from multiple threads, adding a load fast-path introduces a race condition.
Afaik on all of x86, arm, and riscv an atomic load of a word sized datum is just a regular load.
We're talking about collating a maximum, by definition every write to that is an increase.