SOUL: A New Efficient, Portable, Low-Latency Audio Programming Language
soul-lang.org
soul-lang.org
https://www.youtube.com/watch?v=-GhleKNaPdk
Sorry. No release yet. But in development by Roli. Makers of cool beat making machines. Seeks to be "OpenGL API for audio"
Found via the wonderful "This Is LINES" HN-style forum for sound technology ;)
Am I understanding right that the language _per se_ is not that special, but the integrations with hardware and other languages are what's special?
And audio programming languages themselves have lingered in a research/prototyping/hobbyist zone for decades. CSound, one of the oldest examples, dates back to the 70's, but production hardware and software tends to ends up with funky embedded-code style toolchains because the higher level provisions aren't good enough for professional apps. That said I have seen a few games(indie scale projects) do something like run an instance of Pure Data rather than code up their own custom audio system. It's all a bit related to the difficulty of doing low-level, time-sensitive I/O in a computing world that has towers of abstraction. Graphics have been ushered up that abstraction path while audio is still living in the 90's, so you can see cool modern shader effects in your phone browser, but the same browser can't play Amiga tracker music without stuttering and popping.
The attractive part of the SOUL pitch is in making the base tooling experience better, and then down the line using that to boost performance.
Spore used Pure Data to do algorithmic music. IIRC somebody there went through Fux's Gradus Ad Parnassum and reduced it to a set of branches to automatically generate common practice harmonic treatment of pseudo-randomly generated melodies. :)
> but the same browser can't play Amiga tracker music without stuttering and popping.
You're on a much tighter schedule for the Amiga audio emulator, and the failures are more glaring.
A video stream that moshes for awhile is kind of beautiful as long as the audio remains pristine IMO. On the other hand, pristine video with stuttering audio starts a countdown for canceling the streaming service.
But even given that, yes-- the webaudio interface has been quite lacking.
Also, if it's GPU based (shaders?), that's going to increase latency not decrease it, as anyone who understands bus transfers and that GPUs are optimised for throughput not latency will know.
With that in mind, did you read my comment? I explicitly mentioned how GPU "acceleration" is going to increase latency, which contradicts one of the very few definite statements they make.
Anyway, I can understand the supposition that GPUs are good for audio processing because they have oodles of processing power and memory bandwidth, however this will hurt latency because there's at least one extra bus hop (assuming the GPU can DMA to the sound card, which I doubt, and if not it's two extra hops!) and it's not like modern multicore CPUs lack throughput for audio purposes.
tl;dr They claim some latency advantage and then use GPU acceleration as an example; as someone doing lots of GPU coding I find that contradictory because of the extra bus transfers.
The end result of this would function like existing hardware modules and software plugins: it's assumed to be always running, and the host(sequencer, keyboard, etc.) sends it messages to trigger audio events. The latency you are discussing is messaging and control rate latency, a problem the audio world has considered mostly-solved and production grade(minus a few scaling and edge-case warts) since the invention of MIDI in the 1980's. This would not present more or less control latency than existing options. Audio output would otherwise be desynchronized from the CPU, making the low round-trip latencies described feasible.
For example it can be executed in the kernel driver, in a DSP, or even on an external device (e.g. a wifi speaker).
The concept is hard to explain in a couple of sentences - the keynote is really the best place to go right now if you want to properly understand what we're trying to do - but essentially, it's about using less CPU than the way we currently run audio code (i.e. running on a CPU in a non-realtime OS, in user-space)
And no, GPUs were never part of the plan for this. I've never considered GPUs to be a remotely sensible place to run audio code. We're mainly talking about DSPs or CPUs with realtime OSes.
...however, surprisingly, after we announced it, we had several very serious companies ask us whether SOUL could run on a GPU because they do have use-cases where that kind of architecture fits quite nicely. Wouldn't have expected that, but it may be something we look into if people want it.
Keep in mind that this is more analagous to a GL shader. It is not a general purpose language.
With embedded DSLs you don't get the portability of a library and reusability of a library, but instead get a horribly crooked, compromised, limited grammar based on the meta-programming limitations of the hosting language (looking at you gradle!).
Apart from that, this seems like a fantastic idea. My only concern would be about memory - what if you're writing a sample-based synthesizer for example? DSPs don't have enough memory for that sort of thing.
And yep, some existing DSPs can be tight on memory, but it's a chicken-and-egg situation - unless people want to run complex synths on them, there's no incentive for the manufacturers to put more memory on there! If we're successful with this, we'd hope that it'll start to create a market for DSPs which DO have the oompf you need for that kind of thing.
>SOUL unlocks native-level speed, even when used within slower, safer languages. The SOUL language makes audio coding more accessible and less error-prone, improving productivity for beginners and expert professionals.
This part confuses me. Is it a language or a library or both?
I'm sorry if they explain this in the keynote i'm at work and can't watch it and their site seems to be lacking in information.
From the comments here I understand it's trying to be the glsl of audio...but that doesn't really make sense to me. Audio doesn't really work like video shaders.
Audio doesn't really work like video shaders.
I get the impression it will include a grab bag of common DSP functions, transcoding, etc. So if you have hardware that supports a function, it runs accelerated. The project looks fantastic to me. I don't know why there is so much negativity in the comments here.Probably because as of right now it's vaporware with a marketing budget. Drumming up hype for something you haven't released yet is a pretty good way to get mixed reactions.
(I wish we had a marketing budget!)
Or even whether it's a language or a library. It says it's a language but that it can speed up higher level languages. I don't understand that statement. Will they provide bindings for other languages to link to or use? Will you have to write your own for your language of choice? Do you compile your SOUL binary first and link to it like a .so or .dll? Do you write standalone programs with it? It's fairly unclear.
It sounds more like a library than a language. And yet it's called "SOUL Language"....Docs would be nice.
It's a language AND an API for deploying it to appropriate devices.
We'll release more textual docs soon - sorry there's not much on the website yet (still writing it) but the keynote is probably the best source of info if you want to know more right now.
"Latency" may actually be the only term more subject to confusion than the term "blockchain."
Is the page claiming that this language can achieve a lower latency under a realtime kernel than can be achieved by iterating over an array of function/arg pointers and delivering at regular intervals the output to ALSA at the lowest delay supported by the machine/chipset/audio hardware combination?
Or is it claiming that the architecture of the language allows the programmer to more easily do complex DSP computation reliably and safely at that same round-trip latency achievable using the design I just described above?
I'd imagine the answer is obviously the latter.
But I'd bet most developers who read the advert would would think this mystical language allows the users to access some kind of new "el dorado latency" with commodity hardware that is not currently achievable using Supercollider or Pd.
I bet that because I know users who have implied that running their single audio-generating app in conjunction with Jack-using-ALSA-backend delivers audio with lower round-trip latency than ALSA alone. They think this because the Jack page says that Jack adds zero latency to the system, and they then confuse the concept of Jack's "system latency" with the "round-trip latency" of their use-case.
For example, even on commodity existing hardware, doing your audio processing inside the audio driver itself with kernel-level privileged control over its threading, affinity, and ability to write directly to the hardware buffers is going to be faster than than passing buffers and task-switching between multiple user-space threads, and dealing with all the scheduling jitter issues.
[0] - https://puredata.info/
Obviously there's a ton of technical detail that we've not released yet, but yes, allowing streams running at arbitrary samples rates (not just "a" or "k") to be invisibly handled by the API and runtime is a concept that's deeply baked into the design.
Btw SAOL (structured audio orchestration language) was proposed as part of the mpeg7 initiative, which is itself quite long dead.
SAOL was based on csound concepts. Back then, we had "csound cards" that ran csound units directly on them in real time with a near-zero (read sub ms) control-to-sound latency. Btw iirc even a control signal sent to a Bluetooth device can suffer a 3ms delay, which is low, but not "essentially zero". I get about that much from trigger to sound on the ios. BT Streams have 100x more delay than that (also noted in the keynote).
PS: finally watched the video.