I work on the Web Audio API at Mozilla, here is an article that I wrote that explains what goes wrong here, distilled to the core: implementing a metronome in literate programming in a way that will _never_ glitch:
https://blog.paul.cx/post/metronome/. The result is nowhere as interesting as this app, but this is to get the point across.
https://www.html5rocks.com/en/tutorials/audio/scheduling/ is an article written by my former fellow co-editor of the Web Audio API specification, that talks in length about this.
In 2021, with the AudioWorklet [0] being available in all major browsers (including Safari as of a few weeks ago!), arbitrary computations can happen on the real-time audio thread. It is now possible to send a copy of the sequencer data (which note when, what sample, what velocity, etc.) to the real-time thread, and let it play, unaffected by the load of the machine, be it from the browser or other programs (until everything breaks down because you're also overloading the audio thread).
I think this app is very good and does the best it can under the programming model it's using, considering the environment it's being run in, but if more robustness is needed, there is now a way !
On your last paragraph, it is still _absolutely_ critical. An audio software that doesn't use a real-time thread for rendering cannot render audio without glitches, even under no load, even on a fast machine. You don't need to tune the OS [1] and pick special hardware these days: today's computers (and even computers from 10 years ago or older) are perfectly capable or running reasonably complex audio processing workloads _if the software is written correctly_. This is regardless of the use case. When we disabled the real-time thread on Windows by mistake in Firefox (adjusting priority of the process level, and then that was inherited to all threads, very quickly reverted), we had reports in a matter of _days_ (if not the next day), and the Firefox Nightly population is not that big. This was I think playing regular videos on social media, not even anything complicated.
If the author of the app, or anybody, wants to profile the app with the right tools, I've written a blog post [2] about this as well. I hear this tooling is used by audio professionals, now, and that they're happy with it. Happy to hear about any suggestion though :-).
For your Firefox-specific question, you're right that increasing the content process count will make everything better, and we're constantly improving memory usage so that it's possible to do so even on lower-specs machines, maybe it's worth a try. Please look at the "resident memory" values, and not, say, virtual or shared, because we make heavy use of shared memory and copy on write facilities of the OS to limit the memory usage when using multiple process (which we need to do for origin isolation and avoiding Spectre issues and sandboxing various bits of the browser as hard as we can, for example).
[0]: https://developer.mozilla.org/en-US/docs/Web/API/AudioWorkle...
[1]: (but you can to improve robustness further, and it helps to have good drivers, such as Jack on Linux, WASAPI with MMCSS on Windows or ASIO, or just the normal stuff on macOS)
[2]: https://blog.paul.cx/post/profiling-firefox-real-time-media-...