For music quality webRTC you need 3 things: disable audio processing, stereo=1 in the SDP and a way to limit bandwidth usage so it doesn't saturate the available bandwidth and create errors.
Disabling video is also really the best thing to do when recording for this reason (bandwidth saturation), and also Chromium will give you much superior experience. Safari and Firefox isn't quite there yet: Safari can't let you choose your output device and lacks some other useful features, and Firefox doesn't yet seem to allow stereo Opus, maybe that's changed since I tested. Microsoft Edge is now Chromium so you're good to go.
Of course all the chain has to be stereo, that goes without saying: input signal is stereo, negotiation has been done in stereo, having enough bandwidth is important (otherwise opus goes mono), and then playback has to be on stereo hardware (but that's the easy part).
Will this also fix this issue? So everyone will be able to hear everyone?
You will think you are in time with someone, but you will react when you hear/see them on your screen, which is maybe .15 seconds after they actually made the sound/movement. And then they will hear/see your reaction .15 seconds later again.
With rtt < 20ms that should make musical performances possible. After all, sound only travels less than four meters in 10ms. So this is just like singing in a choir (with more visual delay - but that can be solved by having a conductor).
Unfortunately I'm not aware of any software making that a practical reality, even with ftth.
There are all sorts of other latency that need to be taken into consideration too, and unfortunately in practice those do add up to live music being unplayable on pretty much any network.
There's a really interesting project called NINJAM https://www.cockos.com/ninjam/ which is designed for live music jam sessions. It flips this fundamental constraint on its head - instead of being real-time, it streams everyone else's output delayed by one bar (theoretically any interval >RTT I guess?). I haven't tried it, but it's a really cool idea.
Opus minimum frame size is actually 2.5ms: https://tools.ietf.org/html/rfc6716#section-2.1.4
Of course there's a ton of other potential sources of delay that make my fantasy hard to achieve, probably already starting at the typical USB microphones (in headsets/cameras).
USB should not really be an issue here, however.
Minimizing latency is certainly technically feasible, it's just hard for stupid reasons.
I just would like to hear everybody at the same time, but what I hear is always one person's sound getting preference over others. Or sounds just alternate randomly based on the volume, I'd guess.
You could try an app like Mumble, where you can turn that off (and also have other detail controls, at the cost of a bit more initial setup).