Modern volume normalization isn't done based on the highest peak of the track. Instead they normalize based on the average perceived loudness of the whole track (to a level below 0 so there is headroom for peaks). It's intentionally designed to avoid the exact issue you describe.
If this was not done users of streaming services would have to be constantly adjusting the volume to deal with perceived volume differences between tracks due to different levels of dynamic range compression.