Elk OS – Audio Operating System
elk.audio
elk.audio
Here’s a download link if anybody is interested:
For more than 10 years, Linux with the PREEMPT_RT patch has been capable of more-or-less hard real time, with latencies in the 10's of usecs range. That can be thwarted, however, by a variety of hardware features (SMI being the worst offender, plus devices that hog the PCI bus for too long), so you are not guaranteed this performance just by running that kernel.
Also known as soft realtime.
For hard realtime, you need guarantees. Linux is too complex and cannot possibly offer them; this is the territory of formal proof, with complexity growing exponentially with code size.
Look at seL4 for hard realtime.
Fortunately, for realtime audio, you don't need hard realtime unless you're very close to overloading the CPUs. Soft realtime (i.e. 99% of all interrupts handled within N usecs, 100% handled with N*2) is good enough.
AFAIK Linux RT kernel is designed to hit every deadline, but they just cannot offer a formal proof.
So "more-or-less" hard real time.
Usually when you're entering that realm of such requirements, you're rarely talking about off-the-shelf devices.
And that's why it is soft realtime. No guarantee, no hard realtime. It really is that simple.
Linux is still Linux, and it is still fundamentally not capable of hard realtime.
As a teenager I had heard about BeOS and I tried it out. I must have gotten it from a magazine's CD.
And by mistake, clicking around, I opened all of my mp3 collection's files at the same time. Each of the mp3s opened in their own little player window - and they all started playing at the same time, without the UI starting to lag at all.
It gave me an appreciation of what a computer can do
Oh well.
The Arch Linux wiki has some description of their real time patches:
These kinds of tradeoffs don't make sense for a lot of common Linux workloads, which is why it isn't the default.
The ability to compile a (close to) real-time kernel seems to be one difference between Linux and BSD not often discussed.
Ooh, a rare latency claim that actually specifies which type of latency they mean! One millisecond round-trip latency, if true in practice, is quite impressive.
I'm going to continue reading now... :)
Do you have any research you can suggest for that 5ms number?
[0] http://www.klabs.org/history/history_docs/reports/dfbw_tomay... (Keeping the delay at or under 10ms was the recommendation I got from the study P.I.)
One issue is when you get some acoustic sound too (open headphones, bone conductance, singing) and pretty much any delay in the headphone signal causes phase interference.
Did you ever see people who try to sing with headphones with a delay of around 200ms? That really messes up your performance.
https://www.jstor.org/stable/10.1525/mp.2006.24.1.49?seq=1
And another (follow the links to the article, but I thought that particular graph was most informative) :
https://www.researchgate.net/figure/Audio-latency-tolerance-...
Jitter is killer though - if the latency is randomly changing it's really annoying.
Or is this just meant as a visualization of latency in these scenarios versus what it would be like with musicians playing x ft apart from each other in a room?
Since I only used acoustic drums until then I never gave much thought to the idea of having your drums sound delayed. It was quite eye opening.
It is relevant, and monitoring via headphones or a floor speaker is a mitigation.
It's also a problem for large orchestras. If everyone else just plays along with what they hear from the people near them, the timing is going to be inconsistent across the orchestra pit even if you have some obvious audible reference like loud percussion. The mitigation is to orient yourself as much as possible around the visual cues from a conductor.
"If the system processing your signal introduces a 20ms delay, it would be the same as standing 20feet away the loudspeaker you're using as reference."
At 10ms (and possibly less), you will start to notice the lag and it will affect your performance.
https://www.soundonsound.com/sound-advice/q-how-much-latency...
1. If you strum the strings of a guitar and want your laptop to make you sound like Jimmy Hendrix, it had better output sound quickly enough that you perceive it happening "immediately." (I.e., it should sound and feel just like an electric guitar that's hooked up to an amp.)
2. Such a laptop like the one above had better make you sound like Jimmy Hendrix for every moment you are playing the guitar. The website doesn't state it, but they are making a claim to 1ms round-trip latency in the worst case. If your first strum of the guitar convinces you that the resulting sound is happening immediately, then every subsequent strum must also sound like it's happening immediately, too. Users don't merely rant about a sporadically janky digital instrument as if it were your run-of-the-mill Windows 10 UI. They throw it out.
3. There's something I call El Dorado latency, which is always half the latency being claimed and is eternally sought after by users. (Works for system latency, round-trip latency, etc.) Since the unit for measuring audio latency is milliseconds, and since "1" is the last whole number greater than "0", a claim of 1 ms essentially short circuits the search. E.g., if they had claimed "2" then El Dorado latency requires the user to imagine how much better the system would sound at "1". :)
4. Okay, I'm being a butt in #3 above. Round-trip latency in audio seeks to go down to "ultra-low" latency measurements for the same reason gamers want frame-rates to go up to gazillions. In each case it means developers can choose one of the following for their end-users: a) add more hungry, hungry hippos per unit of time, or b) make the same number of hippos hungrier.
Edit: clarifications
Half of 1 millisecond is 0.5 milliseconds.
Just because their published latency is the “smallest whole number” as you seem to be pointing to - doesn’t mean that it can’t be halved. What a strange thing to say.
Ableton measures latency down to the tenth of a millisecond. It’s really not a big deal.
No one cares if your pad sounds are a little late, but variable latency absolutely destroys drum grooves, sequenced electronica, and the kind of super-precise playing needed for good classical piano.
You might think most people can't hear the difference. And they can't. But - as with other audio production values - they can feel the difference. Tiny variations in timing are the bedrock of many kinds of musical expression. If they're not there, people notice.
For keyboard input, one paper (whose citation I can't seem to find) used a clicky mechanical keyboard, pointed a mic at the computer, and measured the offset between each corresponding pair of keyboard click and sound output.
Bela is indeed fantastic, but the CPU is a bit weak.
It never went anywhere, and I'm no longer in contact with the blind friend who was helping me, but I often wonder if today's solutions such as Apple's Voiceover aren't the ideal solution.
They seem to be trying to adapt the GUI for the blind, rather than trying to build a text/voice based system from scratch, like I was doing.
Siri / HomePods and similar appliances are getting there, I suppose, but they still feel like computing accessories, rather than fully fledged computers themselves.
Do you remember which speech synthesizer you were working with in the 80s?
</nostalgia>
SPEECH SAY Hello *SAY How are you?
My comment was intended as an incredibly subtle reference to the trivia that the BBC Micro game "Citadel" 'famously' had speech generated by "Speech!" in its intro screen (but appreciate you taking the time to share the link :) ) as shown here : https://youtu.be/Hu5vu9SgGZI?t=44
For anyone else wanting to try out the emulator link in the comment above, HN messed up the formatting, so the asterisk prefixes were missed on the first two commands:
*SPEECH
*SAY Hello
*SAY How are you?For users, good look finding more than a couple of plugin developers that make their plugins available for (a) a Linux based system AND (b) the ARM architecture.
For developers, VST is just an SDK, and doesn't require OS support in any meaningful way. So sure, if you want to create VST(3) plugins for a Linux+ARM platform, you can do that, with or without Elk.
and an operating system that runs on pi
Less diversity decreases competition and overall technological development. In the 1980s, some OS developers actually blamed Unix for stalling the progress of OS research.
Wouldn't a world of Arm based synthesizer be kind of stagnant? Sure you can dream that compatibility allows more open platforms, but we all know hardware companies always end up making the end product proprietary anyway.
To be fair, I don't know a lot about hardware. I'd also like to guess that the world of music electronics has been longing for some sort hardware medium to sort out the mess of obscure ISA's and poor documentation that hardware engineers have to put up with.
Once performance is comparable for the workload, ARM-based synthesizers aren't inherently any more stagnant than DSP-based synths, in the same way that ARM-based tablet apps aren't necessarily stagnant; if anything, the freedom from crappy compiler lock-in, incomprehensible instruction pipelines, assembly code, and (as you say) poor documentation. Moving to ARM, GCC/LLVM support, and (eventually) Linux makes all the things other than making neat sounds (UI, interconnect, networking, power management, etc) easier, so that developers can spend more time on the creative bits.
I mean stuff like Zephyr, NuttX, mbed, RTOS, Azure RTOS.
The computer in a digital synthesizer is an implementation detail to the consumer. What they're buying is the synthesis engine, the way it sounds, performs and interoperates, the case, the user interface etc. In that sense the CPU used seems about as relevant to stagnation as the screw heads used for the case.
- The website says: "With Elk hardware companies can move away from dedicated chips and use general purpose ARM and x86 CPUs without any compromise in terms of latency, performance and scalability". Is this intended as a ready-made OS for hardware synthesisers, or as a way to host VSTs on consumer SoCs?
- Assuming you're not targetting an obscure platform, or have a cumbersome toolchain, is writing a bare-metal executive the most difficult or expensive aspect of engineering a synthesiser? I've done some experiments breadboarding simple synthesisers with Arduino and STM development boards plus a 16-bit DAC. In my case the most challenging aspects weren't getting a real-time executive loop or handling I/O. Mind you, I didn't go as far as implementing a GUI.
- Is real-time Linux not overkill for most synthesisers? From the outside looking in, it looks like it would add more overhead and complexity than solve problems. If your objective is just to host a VST, this might be a decent trade-off. Or maybe if you intend to use high-level languages for the GUI.
Seems like both and more.
> is writing a bare-metal executive the most difficult or expensive aspect of engineering a synthesiser?
Likely being a big part of the expense, yes. Obviously the hardware parts being a big chunk of it too, programming the software side of things are also expensive and sometimes makes the product receive less updates in the future because programmers who did understand the platform left the company. See Octatrack as an example, where the code is so complicated it's unlikely to see any bigger features released to it now as the original programmer left the company.
Also, consider that many people have experience with general computing platforms while not having special expertise about particular chips, this solution on a general platform can help more people get started, learn more and maintain their own setups with skills that easily translate as well.
> Is real-time Linux not overkill for most synthesisers? From the outside looking in, it looks like it would add more overhead and complexity than solve problems
Probably, but probably not. Music making is more about having options available and use those in creative ways. Being able to use a general computing platform opens up more options than specially programmed OSes for particular instruments. Being able to revive old synths with new software would be wonderful.
I did some audio implementations on MCUs of all kind, and if I managed to pull out multitracked real-time mixed audio players under 4ms latency, I learn that the real problem here is RAM (other than computing power).
Many audio effects requires lot of RAM. Think of how would you do a 5 seconds delay effect, or good quality pitch shifting (FT and inverse FT), reverb, echo, etc. Before you know it, 128KB of RAM is not enough any more. So now you are dreaming about 128MB of RAM (as minimum).
If you want to process or output audio from/to network, USB or save the output into a file then you have to add that overhead to the equation.
Before you notice it, you'll start looking at small processor like the i.MX28 or the like. Bare-metal BSPs for those processors are quite expensive in terms of code/time. You get what the manufacturer gives to you, and that is, a BSP for Linux.
My RPi4 is quite happy doing what any audio engineer would consider "high resolution audio", no kernel recompile required.
[1] https://bela.io/
You'd get more or less the same thing by writing a kernel module for Linux, which would also have to use somewhat specialized techniques for interacting with user space application code.
This seems difficult. What is the 5G latency chain like in practice?
My real question, I guess, is how in sync do the speakers need to be? Just on the order of human perception, or do we need them to be in phase?
Whereas the holy grail of the audio world would be analog effects controlled digitally :)
Recall time is what forced a lot of professionals to totally move in the box.
Yes, but also more. For me, the greatest feature is the ability of running VST plugins without having to use a full computer/laptop, since my setup doesn't include any traditional computers.
> Whereas the holy grail of the audio world would be analog effects controlled digitally :)
That's been done for a long time already, bunch of Elektron Analog machines does this already, just one example.
It's been done but it's not really "mainstream" and standardized.
Imagine a standard where all your pedals etc can be recalled / automated from another (central) unit, with a protocol that the hardware can adapt to, motorized faders, rotary faders with leds, or basic components etc.
Maybe that's what Wes Audio is doing with the gcon protocol that has been linked below.
I thought of Ubuntu Studio as well, but it looks like Elk is intended to be more lower-level / for appliances. I'm not totally sure though, but I'm guessing it doesn't come with the full suite of Ardour etc.
I do wish there was something like Ubuntu Studio, except focused on audio-only (and more polished / supported). Mostly just because JACK + low-latency kernels are an ordeal to set-up and maintain on a normal desktop, so it's much easier to just install a pre-configured distro on a studio computer and be ready to jam whenever inspiration strikes. Unfortunately KXStudio as a separate distro isn't around any more. I still use a KXStudio 14.04 install on an old (airgapped) DAW computer.
I wouldn't call linux-rt from the AUR too hard to install, but I can see how it would be a pain for someone who just wanted to pick up their software and do their thing without a long setup. (Since there would probably be more things to setup aside from just the realtime kernel)
A turn-key audio-only OS like what KXStudio 14.04 was truly perfect for pro audio, since it had Ardour + 100s of plug-ins and tools all pre-installed and configured.