A tale of two clocks – Scheduling Web Audio with precision
web.dev
web.dev
The "up-and-coming" (back in 2013) high-resolution timer performance.now() exists on all browsers for a long time by now, but is clamped to milliseconds on some browsers because of the Spectre/Meltdown mitigations. The web audio timer is also affected by this precision reduction, but I don't know what's the current situation across browsers (I was actually hoping that the article would talk about this).
Anti-fingerprinting settings can also heavily reduce timer precision.
In short: timing on the web platform is (currently) an incredible mess.
...
...
...
...
timers
see: https://developer.chrome.com/blog/timer-throttling-in-chrome...
This secure context makes it harder to support various random ad networks, but if you're deploying your app via the web and don't need to monetize it by ads, it's a way forward.
(there is a workaround however by injecting the required response headers on the client side, but who knows how long that's going to work: http://stefnotch.github.io/web/COOP%20and%20COEP%20Service%2...)
The way I understand it, Web Audio API only lets one schedule audio source nodes with start and stop methods. As I needed to schedule something that was not audio related, I ended up creating silent oscillators of almost zero length and relying on the 'ended' event.
But running a scheduled callback is better! It's not immediately clear to me how you do it. Can you maybe explain your code a little?
It runs your callback every 16th note, slightly before it is actually due to play, this can vary by a few milliseconds each time but that doesn't matter as the actual audio context start time for the note is passed in to the callback, and you use that to actually schedule the events for that 16th note, eg: osc.start(time).
You can schedule 32nd notes etc too by using the stepLength property that is also passed in, time + (stepLength/2) would be a 32nd note.
Hope that makes sense? I do need to write a better description on the github page of what it actually is.
The inner workings of the library itself are mostly just as described in the article with a few tweaks.
It is mindboggling to me that a floating-point representation was designed into this. The required precision is well known and won't change. It is also well understood that this precision is required "locally" at possibly large future values. Even having the value wrap is mostly not a problem. Floating point is the opposite of designing for these things… (and no other audio API I'm aware of does this.)
(Okay, after doing some quick math, a double precision float should be OK for more than 100 years. The mismatch still irks me…)
I can't speak for Windows and Mac but in ALSA this isn't possible.
The ALSA documentation seems to disagree with you?
https://www.alsa-project.org/alsa-doc/alsa-lib/group___timer...
(disclaimer: I have not used this API and don't know if it is missing some functionality — is it?)
https://developer.mozilla.org/en-US/docs/Web/API/AudioWorkle...
...doesn't change the fact that WebAudio is an incredibly overengineered contraption of course.
PTP precision time protocol, albeit implemented on the wrong layer to be directly useful here, solves it by timestamping frames.
Beyond what the article discusses, I can imagine long-running audio applications where it might be useful to build an estimate of the rate at which the system time drifts relative to the sample rate clock. Even so, we're still in the realm of the application software having to take these slightly divergent timing sources as given.
The issue really is that all your code is running on the UI thread with JavaScript, so you can’t accurately time stuff, whereas in C++ you’d write this stuff to execute on the audio thread callback. That said, this is now possible with AudioWorklets - but I don’t think you could interact in a sample accurate way with browser-supplied WebAudio nodes from there, so ultimately you still end up needing to schedule them for a future time if accuracy matters.
The other problem of (traditional) WebAudio is that all work needs to happen on the browser thread which runs in time slices at a very low frequency (usually 60Hz or the display refresh rate).
There's now a solution called audio worklets, but those aren't exactly trivial to work with unless all work can happen in the worklet thread (for instance you need to write your own ring buffer to push commands or sample data from the browser to the audio thread).
The problem isn't synchronicity, it's syntonicity, albeit only to a limited degree. What you need to know is: how much is your local audio sample clock off from the source audio clock, and from your (distinct) video output clock.
But microseconds precision is indeed perfectly fine, ideally with some way to get simultaneous timestamps from the different clocks at least for the local clocks.
I don't see PTP playing into this either, if anything the associated hardware timestamping features might be a "nice to have to make it perfect", but noone I know wires the audio clock into the PTP timestamping unit.