What Is the Future of the DAW?
djmag.com
djmag.com
> Open a DAW from the year 2000 and it’s highly likely you’ll recognise the vast majority of the features — both functionally and visually — from any DAW you might use today.
This assertion that "old is bad" really, really needs to die. Change for the sake of change is bad. I don't see why all these new innovations can't be baked into existing workflows in a natural and intuitive fashion.
But then, in some cases, do I really want them? Like generative AI for music. I'm not sure but leaning towards no.
A local LLM running on x86 or M2 would be great. But API callouts while waiting for the silicon/Cuda migrations are fine too.
Of course things are far, far more convenient and faster nowadays, but the fundamental paradigms are mostly the same.
- LilyPond for writing the final notation
- MuseScore 4 for playing around, MuseScore 3 for playing MIDI
- REAPER as DAW, SFZ plugin for midi soundfonts
- Audacity for processing WAV files
- music21 python lib if I need to programmatically process MIDI (which I commonly do). My scripts using this lib.
- Kdenlive for video processing
- Most importantly: my piano Roland FP-30X
Now among these the only one I don't hate is FP-30X. That piano is fucking amazing, it sounds and feels great.
Kdenlive is also not terrible but I don't do a lot of video processing. Just basic notation videos for YouTube.
Everything else is incredibly, incredibly buggy. MuseScore 4 is thrash, almost every interaction I have with it exposes some bug. REAPER is usually fine but there are pretty annoying bugs when it comes to MIDI import/export that involves time signature change.
My workflow is so dogshit that last month I decided either I'll write my minimalist tiny notation/DAW tool or reconsider this hobby. So far haven't been able to do anything major but I'm confident I'll invest some time into developing my own tools in late 2023 or early 2024.
I'm glad people like their workflow. Unfortunately my own experience with Linux audio processing has been nothing other than encountering one bug after the other one.
EDIT: Whoops, don't know why I said Ardour. I never used it, I actually meant REAPER. Fixed now.
You kind of buried the lede there...
In fact FL just added a python integration that I haven't even begun to explore yet, but you can use it to generate piano roll scores (among other things?)
foobar2000 handles tagging and DaVinci Resolve/Blender/ZGameEditor handles video.
Every time I've tried to do anything in Linux for music its been a headache. Granted its been a while now. But I'm a little too invested in the Windows audio ecosystem to switch.
Just use Windows and all of these problems go away. When I make music my stack is basically Ableton Live, NI Komplete and Spitfire Audio and I have exactly 0 of these issues.
Also I truly hate using windows or osx.
I kind of just use reaper as a dumb recording reel - I don't use much of it's timing features.
At the moment I'm mostly just writing up exercises to fit my current goal of learning everything in any key.
Oh nonsense -- in fact, bullshit. No, it is not. This is a personal problem; seek help.
>Also I truly hate using windows or osx.
And we truly hate hearing about it, and how things don't work well on your platform of non-choice. So suffer it in silence please.
you can pay whatever you want for ready-to-run versions of Ardour (US$1 up). Need it to be better? Pay more! :)
Heh, been using it for years before that feature (which is admittedly now years ago).
Synth graphs are sweet. But I kept going back to Logic to compose compositions ;)
Habit has stuck, unless I have something bitwig-y in mind.
I've worked in a lot of studios with 2" tape or, like Radar or an HD24. In all of those, the engineer had really great chops. Functionally, for most "traditional" (I do a lot of folk, jazz, and country) I just do what they did, and Logic is just a glorified tape machine + mixer + processing.
I am still mostly doing stuff in single takes until I like what I have and then maybe punching in a note or two here and there. It's a really fast way of working. And the old stuff I have is quite good: API/Neve front end, AKG/Schoeps mics, genelecs.... these are all things that have been around a long time, and I am not seeing any benefits there for anything- they are fast and easy to use well without a lot of tweaking.
Like, you could walk into any studio post 1980 and more or less find things will have a 1-1 correspondence to a modern DAW.
There is nothing wrong with those tools, and so I assume folks will keep using them.
However, I can see a couple of areas where I think that smarter tools might help. A big chunk of what I am doing already involves programmed drums, and a smarter "Drummer" in logic would be an improvement.
Also I (and a lot of other folks) aren't super stoked about harmonies sung by one person overdubbing a bunch of lines. I'd be interested in some workflow where I could, like, sing a harmony and it would change my voice in a credible sounding way.
But as far as workflow goes, I really don't need compositional tools- I could do most of my writing tasks with a pencil and notebook.
This would be a godsend
I’ve dropped more than $1000 with Ableton and always recommend Reaper first.
Is the old in this case the wheel, which is fundamentally the same, except there have been thousands of years of technological improvements surrounding it, or is it the computer, where modern day computers have only a passing resemblance to the original thing.
The difference between music and driving a car is that music is art. Many of those who labor in the space enjoy creating from scratch. Technological progress does not automatically make things better
Commuting to work/school/etc in a car might not be art, but driving for sport, at the top of the profession is very much an art form simply because of the technical skill, creativity, and emotional depth and investment. Just as a musician understands the nuances of their instrument to produce a captivating melody, a driver must intimately know their car's attributes to master each track. The racetrack is their canvas, with split-second decisions, adaptability, and emotional connection to the track, car, and other drivers. Like a well-composed symphony, high-level driving's graceful turns and passes resonate with both the drivers and the viewer. It's not just about who crosses the finish line first (despite what the Fast and the Furious movies told you), but the elegance and finesse required to drive at that level is befitting a music virtuoso to educated viewers.
Very much agree on the last point, technological progress often isn't. Electric cars are the future, but the diagonal torque curve of the motor lacks the charm of ICEs.
There are (at least) four categories of DAW users:
1. Professionals who are being paid to make music, and for whom time is essentially money. Tools that speed up the production of that music are both financially valuable to them, and also make their overall lives easier (if done right).
2. Musicians for whom making music is a creative act of self-expression. They are not being paid by the hour (of music, or of effort), they are not under deadlines, but they do want tools that fit their own workflow and understanding of what the process should look like.
3. People who just want to have fun making music. Their level of performance virtuosity is likely low, and the originality of what they produce is likely to be judged by most music fans to be low. They want results that can be quickly obtained and are recognizable musical in whatever style they are aiming at, and they don't want to feel bogged down by the technology and process.
4. Audio engineers in a variety of fields who have little to no interest/need for music composition, but are faced with the task of taking a variety of audio data and transforming it radically or subtly to create the finished version, whether that's a collection of musical compositions or a podcast or a move soundtrack.
The same individual may, at different times, be a member of more than one of these groups (or other groups that I've omitted).
The needs of each of these groups overlap to a degree, but specifically the extent to which the current conception of AI in music&audio can help them, and how it may do so, are really quite different.
We can already see this in the current DAW world, where the set of users of DAWs like Live, Bitwig and FL Studio tends to be somewhat disjoint from the users of ProTools, Logic and Studio One.
TFA acknowledges this to some degree, but I don't think it does enough to recognize the different needs of these groups/workflows. Nevertheless, not a bad overview of the challenges/possibilities that we're facing.
Before Garage Band, I used Tracktion (I think its now called Tracktion Waveform Free) in the same manner. It's been ages since I used it but if you're a Windows or Ubuntu user I think it'd be worth checking out.
There is a big community around REAPER also, and tons of YouTube videos around it. Plus, you can download it and work with it for free, but after a while, it will want you to pay for it, which is only $60...but you can keep using it if you don't (though I would encourage you to pay for it if you like it).
In the same vein, the Novation Grooveboxes[1] offer some expanded capabilities that don't require a computer. Second-hand pricing is quite reasonable for both.
[0]: https://www.roland.com/au/categories/aira/aira_compact/
[1]: https://novationmusic.com/categories/samplers-grooveboxes
I use Reaper as well, but it takes a while to get that "useable" for modern(ish) music production. The benefit is there's plenty of free virtual instruments/VSTs to download. All of them have downsides though as does reaper itself. In Ableton I can make an EDM track relatively fast given the out of the box presets - especially synth drums - but in Reaper using a free VST like HELM makes it kind of a pain to use. YMMV.
No matter what you choose, I do HIGHLY recommend downloading Spitfire LABS though - the free instrument packages are massive and highly customizable. It's truly amazing.
Here's some good VSTs for Reaper:
https://plugins4free.com/instruments/ (when the site works)
https://web.archive.org/web/20181203014924/http://sonic.supe...
https://guitarclan.com/best-free-vst-plugins/
EDIT: oh also trying to master a track in Reaper with free plugins is frankly pretty bad for a beginner vs Ableton's preset limiters and other utilities. The Cuckoos plugins are messy to deal with in my opinion.
But having said that, sometimes the thing that grabs the interest is recording your voice in the windows built in recorder then playing it back backwards. Try audicity?
https://en.wikipedia.org/wiki/Music_tracker#Selected_list_of...
Oh, and samples? Kids used hardware samplers to rip or record their own:
https://news.ycombinator.com/item?id=37376675
And how'd that work out? The end result was a piece of music called a module ("mod"). Strangely enough, I can't find exact (or even approximate) numbers. A snapshot of the MOD archive from 2007 had 120k mods:
https://en.wikipedia.org/wiki/Mod_Archive
So, yeah, a very low barrier of entry... ;-)
If you've got an iPad that's probably the best start (it'll run on any iPad 2 or above). It will run on iPhone's but it's a bit harder to play. There's also a Nintendo Switch version (it's more limited, eg no audio recording or export), and a mac version (but it's pricey). Annoyingly, the Mac and iOS versions are separate, but at least the iPhone and iPad versions come together as one purchase.
I was planning to try Reaper or FL Studio.
My biggest complaints with LMMS are: doesn’t support VST3, can’t see notes for multiple tracks at the same time (although I saw a “ghost notes” patch someone was working on for this scenario).
Unfortunately the open source DAWs don't hold a candle to any of the paid ones. But once you're paying they're all pretty solid. It's like asking if you should move to vim or emacs or jetbrains after starting with Notepad++. They're all good and everyone will have their own favorite. Many people also use multiple DAWs the same way people use multiple text editors. Personally I use Ableton and Reaper
Really, it might make the most sense to just find music similar to what I've made (or want to make), and then ask/research what they're using. Edit: I think that's how I originally found FL Studio and Reaper, come to think of it!
Logic Pro X is the DAW I'm most familiar with and while not "AI", it's "Drummer" plug-in is uncannily good. So good it's indistinguishable from AI. I want more of that. Give me "Bass Player" and "Keyboardist" and "Guitarist", etc, with all the options that "Drummer" currently has, to select style/genre, kit sound, etc.
Another wish list item: Let me point the DAW to a 4/8/16 bar section of multitrack original music I've created, and suggest n number of directions to take it, spitting out each of the individual instruments on their own tracks, so I can mix/match/edit. My imagination is limited; that's where I'd like AI to help.
Is that the case though?
Ableton and such don't charge per revenue as far as I'm aware. They charge same whether you are scoring a $500mil movie or fooling around after hard day of coding.
And I feel in sheer numbers, latter outweighs the formers by several orders of magnitude. All forums I've been to are filled by, at best, "enthusiasts". Thousands upond tens of thousands of us with some disposable income we give to synths and software to tinker with :-).
It's a Digital AUDIO workstation, not a Digital MUSIC workstation
If you want to take this out of HN and talk more specifically, my website is in my profile and I'd be happy to help however I can. If you want to dig deeper into anything specific, I'd be happy to share any experience I have.
It's ok to post your own stuff part of the time, but the primary use of the site should be for curiosity.
The other thing I saw the other day that I thought was cool was a reverb plugin that uses your GPU. Seems like the next step for modeling could easily be in that direction. Especially since the bar there is low, pretty much just the positively ancient UAD hardware acceleration cards, although UAD themselves seem to be going the opposite way and pushing native stuff now.
We (Ardour) abandoned it, because the universal experience of non-programmers was that they had no idea how to even begin to use this sort of feature. The majority of DAW users don't come ready to deal with the complexities of a branching workflow, or even a desire to learn it.
There is at least one band out of Madison, WI that uses/used git with Ardour during the height of the pandemic to facilitate remote collaboration on new pieces. They gave a talk (and played) at the Ubuntu Summit in Prague last year.
As a very entrenched Reaper user, I haven’t tried Ardour, but I’m glad it exists and continues to exist. Thank you for your work :)
Turns out their pivot to being Netflix for samples/plugins is more in demand :)
`bup` notionally does this a lot better, or git-lfs.
https://raw.githubusercontent.com/bup/bup/main/DESIGN
git really needs textual representation for any kind of meaningful commit, and binaries totally break that.
This is precisely what git does NOT do.
My priorities are more in line with turning music production into an integrated and interactive development environment using modern design principles. Sub-modular capability, A/B testing, git integration, non-local collaboration, scientific visualization, notebook style experimentation, integrated synth build / play, web embedded optional interface, social sharing & tutorials, polyglot open source interface (primarily Rust), programmable behavior / macros, higher order signal dependency optimization, algorithmic mastering, targeted oversampling, creative process reusability, etc. You can solve the plugin issue by just synchronizing the output audio by the user that has it installed.
Quality music production is an opague art and everything is way more daunting than it needs to be. Most producers just mess around until it sounds good and that gets people stuck in a local maximum of clarity. If it takes too long to experiment then you won’t get through the effort of trial for understanding. I have spent 15 years building tools as a research quant dev and also a dj (Extrn). There is a huge unaddressed gap in the audio space and huge barriers to entry in accessibility and cognitive burden.
My post-graduate research concerned signal-rather-than-event-based generation/transformation of compositional data, integrated with textural/timbral synthesis.
My current focus is building a DSP framework for this purpose in C++20 [1].
In any case I'm interested in following your progress, and happy to contribute code/ideas if you feel like collaborating (links in profile).
DAW : Digital Audio Workstation https://www.masterclass.com/articles/what-is-a-daw
Caring about their readership is exactly what they're doing, just that you happen to not be what they think of when they imagine the typical reader. The typical reader is already into music production and with a 99% certainty know what a DAW is.
I wouldn't expect every tutorial on "Google's Official Android Developer Blog" to explain that "JVM" means Java Virtual Machine, some resources really are for people who already know a bit about the subject area.
Maybe their philosophy is that the Top 100 should be an open thing and they shouldn't restrict who can enter based on music style... to me, it makes DJ Mag way less credible, but I guess they probably make money out of the Top 100 being so big.
Users here probably feel the same way about HTML, FIFO, DAG, etc
But in general I'd imagine written language to be a pretty infuriating tool for describing what you want musically, when the most interesting parts of music are just about always the ones that you can't really capture with language. You can kind of outline things with written language and traditional music theory, but it's usually just a blurry version of why a specific piece of music resonates.
I think that AI tools for music will most likely just stay as plugins within the more traditional DAW structure. There's only so many ways to represent an audio file, and a fader that controls the volume of a track or some other parameter.
Like mentioned in the article, most of these additions take quite a bit away from the amount of control the artist has over the music, and lowering the amount of 'input resolution' in this sense is a block that's almost impossible to overcome.
Writing into a prompt feels opposite to the creative process but as a first pass its a cool tool, kudos
I don't think it's combined with AI yet, but I have to imagine it's on the horizon. At the moment it basically just MIDI-fies your voice. You can raise or lower your voice pitch to turn a parameter knob, beat box to lay down percussion notes, etc.
This is why these applications have lifespans measured in decades, and it's extremely rare for a new player to be able offer anything new, different, and valuable because the design space has already been solved for the problems these applications are solving.
I wrote a piece on this subject, e.g., why, how, when software transitions do happen for these kinds of apps: https://blog.robenkleene.com/2023/06/19/software-transitions...
First, "the right approach" to building a software program is wildly unspecified: it could refer to the UI/UX aspects and/or the internal design, and these both have dramatic impacts on long term evolution.
Second: the "right approach" for "making music" in the early days covered things as distinct as MIDI sequencers, trackers and early ProTools. It was far from obvious whether all 3 would continue to exist or some hybrid would become dominant (that's actually what happened - early ProTools did not do MIDI; the eventually archetype for DAWs turned out to be a blend of ProTools and MIDI sequencers, and trackers were discarded).
Third: As I alluded to in my comment here about user groups, the right approach is going to differ for different workflows and use cases. FL Studio is not used by many audio mastering engineers; ProTools is not the choice of beat producers.
Fourth: the goalposts keep moving with increasing compute power. The current idea of infinitely elastic audio that has become common among the most popular DAWs would have been unachievable in the early 2000s. Network bandwidth may have a similar impact.
Fifth: the right approach (especially visible today) for some people who are generally "in DAW space" isn't a DAW at all, but hardware designs that bypass most of the functionality associated with traditional DAW design. The Elektron and similar h/w sequencers of the last 5 years are in some senses closer to plugins than they are to DAWs.
Sixth: plugins - the ones associated with compositional elements (you could say sequencers but it goes beyond that) - have long been where the innovation has been taking place. These have evolved quite differently and more diversely than the DAWs that host them. For many users, plugins are the real workhorses and the DAWs are just the scaffolding around that. It would be hard to take a look at compositional plugins and conclude that the "right approach" emerged early.
> I started thinking about this question, of whether software transitions ever really happen, when I noticed just how common it was for the most popular application in a category to still be the very first application that was ever released in that category, or, they became the market leader so long ago that they might as well have been. The Adobe Creative Cloud is a hotbed of the former: After Effects (1993, Mac), Illustrator (1987, Mac), Photoshop (1990, Mac), Premiere (1991, Mac), and Lightroom (2007, Mac/Windows) are all market leaders that were also first in their category. Microsoft Excel (1987, Mac) and Word (1983, Windows) are examples of the latter, applications that weren’t first but became market leaders so long ago they might as well be (PowerPoint [1987, Mac] is another example of the former).
So in this world at least, the longevity of the first mover has more to do with actual and imagined barriers to entry rather than anything especially good about the software itself (and indeed, many of its users used to complain endlessly about the software).
I think the factor you're missing is path dependence. The approach that becomes dominant isn't necessarily "best", just a stable equilibrium where switching costs are greater than potential benefits for most users. I'm typing this comment on a QWERTY keyboard, but I don't believe for one second that it's the optimal layout - I just can't be bothered to learn Dvorak or Colemak or whatever.
The piece I linked to makes all these points, as well as addressing your others.
Everything is represented as blocks of samples on rows of timelines which are then effected by transforms placed on the rows above the rows containing samples. Edits and transform adjustments all happen in real time, and everything is continuously rendered to a scratch buffer that can also be dragged in to the project as a new block of samples.
It is truly a very creative approach and when you see it you will be wondering why nobody tried this approach before. The developer, Colugo, has a new video on YouTube showing how it’s main features work.
A video about Blockhead (experimental digital audio workstation)
Why would expect the next steps to come from "real players in the space", when the previous next steps did not?
Look at how quickly Presonus managed to build Studio One to a crazy-level of credibility (it helps that they hired someone who had already done it twice).
No solid quotes from presonus either.
Didn't say that. They absolutely did!
It took a small startup in Berlin to introduce the idea that maybe everything should be tied to the groove, always. That was not 100% revolutionary (Acid could do some of that), but it upended computer-based music production entirely.
The incumbents are very, very unlikely to have access to people with a deep/serious interest in "AI & music" technology (maybe Apple?). Startups in this space are started by those people.
The article mentions those "extremely capable incumbents" are now stuck with legacy codebases.
> music production capable workstations didn’t really exist
Logic and Cubase were highly capable in 2002. They haven't progressed much in 20 years. The midi drum track editor in Cubase today looks the same as the 90s version.
Look, FWIW, my favorite music making "stack" is
- Fruityloops (as in, sure I'll use FL Studio but it was pretty much solid for me when it was called that)
- Sony ACID. Still unbeatable to me for quickly layering/previewing multiple tracks of loops
- Cool Edit Pro/Adobe Audition. Still much nicer than Audacity. I don't know why Audacity remains so popular, it's way clunkier than the above
Ableton/Bitwig are also fun to play with and I could see them being indispensable for some, but what more do people need that you couldn't theoretically get "incrementally" or "with plugins?"
Presumably because it’s free, but I’m with you on the clunkiness. The UI looks very dated.
I can recommend OcenAudio, which is free though not open source, and for me offers just the right level of functionality in a clean user interface. It’s very similar to how I remember the old versions of Cool Edit pre-Pro, which I thought were amazing bits of software.
If the UI is clunky to use, that's something that needs to be fixed. If the UI works okay, but only looks dated, that's ... fine and there's every good reason to leave it as it is.
Why? Because altering the look of the UI is strongly coupled with altering the usability of the UI. This leads to regressions in usability, making the product "look better" but "work worse".
Leave things alone. I don't care if the UI looks like Motif or TCL or a TUI, but if it works in a smooth, intuitive and pleasing way, there's no reason to change it.
The VIM interface is the very definition of "extremely dated", and that's fine.
It is a bit (just a poor analogy) like saying you could program anything in assembly or C. Sure but sometimes it is just much more efficient time-wise to just use Python, Kotlin, C# or any higher level language and not have to reinvent the wheel or deal with repetitive complexity.
In the end, what I always recommend is to try things and see what sticks, test new tools and retest the ones you didnt select regularly and when your workflow evolves just switch to the more adapted tool. I also find that really nice to improve myself in my current favorite tool as I can often bring back new tricks that were more obvious in other tools.
There is this quote : "To know your future, you must know your past". I agree the cloud/AI combo is revolutionary, but we also need to dig a little deeper into the past to understand what is coming.
ProTools just announced their version too.
And yes, this model is roughly 20 years at this point (Ableton Live started at about the same as Ardour).
I'd also like to see more music creation nuts & bolts built into DAWs, like integrated lyrics, a quick way to shuffle pieces of a song like "bridge" and "verse", visualizations of key center shifts, chord builders and so on. A lot of that stuff is out there but it's definitely an afterthought.
Present in the about-to-be-released Ardour 8 (using a design very similar to Studio One).
They never stopped
To innovate in the space, we need an ecosystem of interoperable tools which can be used seamlessly to perform the same tasks reliably.
I think of the paradigm shift between apache http server and nodejs. Does the server run the code or does the code run the server?
If DAWs switched to the nodejs model, what would be the entry point?
I mean we're really talking about wave analysis then cutting of frequencies, boosting others. This can be automated based on a desired "genre/sound".
For me the most natural way to incorporate AI would be as "ai people" that have particular skills.
"Master producer Eric" "Loop maker Ryan" "Guitarist Wendy"
You can interact with them in much the same way as human session musician, producer, synth programming etc.
I recently bought a Arturia keyboard in order to make some Trap and pop songs, and I immediately realised how important a seamless experience is between the DAW software and hardware. That kind of experience can't be replaced by AI.
I am sure if Arturia makes a DAW, it would be a great one.
As opposed to "non-legacy" code which is obviously malleable and easily adjustable to whatever new user expectation may arise?
None of these have dissolved completely; the rotary dial, the keyboard layout, the swappable storage medium still exist, especially in music production hardware.
It just takes a company like Bitwig to come along and show how else these pre-existing interfaces can be modularised and combined, and some iterations of the Push controller[0] to uppend the traditional fingering required to play them.
All of that credential waving is to say, I think about this a lot. The software workflow is so bad that I have resorted to buying an insane amount of hardware, so, what makes it so bad? * Latency correction hell. - Only an issue for musicians not working entirely digitally * Limited I/O (albeit this wouldn't be an issue if I could get away with using less hardware) - You can't use more than 1 audio interface at a time in Windows * Interface navigation hell - VSTs all load as individual windows - There's often still issues with DPI scaling * DRM hell on those virtual instruments * Software stability hell - Not all DAWs sandbox plugins * Complex routings are often complicated - In audio-contexts, we often want to pipe signal around in branching - diverging & merging - paths, often with signals from unrelated chains changing things (ex: sidechain compression from a bass drum)
But I can almost forgive all of those issues. The real issue is and will continue to be MIDI. * 127 values for knobs? Really? That's quantization you can hear. * No good support for microtonal - Current hacks often break down realllllly bad with big scale. If you're in 31EDO, for example, you only get 4 octaves of range. * Clocking is a disaster, making swing, polyrhythms, etc. more of a mess than they need be.
Those are issues that legitimately limit the expressiveness of all music creation. You can't fine-tune a value in MIDI, you can't play between notes (barring MPE, which itself is a hack and not supported in all DAWs), and - combined with the latency problems above - it means *both* your traditional instrument (ie Guitar) and digital interfaced (MIDI) will have weird latency issues.
AI is not the solution and there's more than one problem. There are a few things which could help: * MIDI 2.0 / OSC / Literally anything with more resolution * Designing with more than just western music theory in mind (microtonal, swing/complex grooves, etc.) * A standardized framework for UI
I think VCV rack [7], albeit it absolutely isn't a general purpose DAW, does a good job of addressing many of these issues, but it can't handle more basic use cases either. Sequencing a long track in it is like pulling teeth and performance is quite bad due to processing everything as float's sample-by-sample, not as a buffer.
I do have some hope that change is coming. The Blockhead DAW [8] rightly has many musicians following this space excited for it's novel approach of treating audio in the timeline as a much more fluid source for quick modification rather than baked content you have to modify elsewhere.
There is, ultimately, an incredible amount of complexity that music software has to account for. I could muse(score) about why this is the case: it could audio having a temporal element, unlike a text editor or maybe just the inherit complexities that arise from trying to create with a medium where you have to consider such an extreme amount of information that caries with it cultural context and an ease of offense to the senses that no other medium suffers (Bad audio is much less tolerable than bad video).
No matter the case, it means that a DAW and the interacting with it needs to be very tightly integrated and intentionally designed with good standards while also being highly flexible. We nailed the flexibility, but totally dropped the ball on intentional design and standards that work for us. I don't think that can be un-done now, not without throwing away support for legitimately amazing tools.
While I'm risking going pure-tangent, I also want to mention that there being no good method for hardware acceleration is a limiting factor in the audio world, along with better standards for UI and data transport, we're in desperate need of an open hardware acceleration standard that's easy to use and not crazy expensive. I should be able to throw in a DSP card just like I can throw in a big GPU and have enough umpf in my system not be constantly worrying about freezing tracks or having audio underruns. This is an extra reason for having as much hardware as I do: A lot of DSP distortions eat CPU and sound really bad.
[0] https://opguides.info/music/midi/ [1] https://opguides.info/other/hci2/intro/ [2] https://opguides.info/music/instruments/#expressiveness-and-... [3] https://opguides.info/posts/aiartpanic/ [4] https://opguides.info/posts/ai2/ [5] https://opguides.info/music/software/daw/ [6] https://www.youtube.com/watch?v=Zpq7g9CcGv4 [7] https://vcvrack.com [8] https://twitter.com/ColugoMusic
Your opguide site looks like an incredible resource, thank you for creating it!
[0] https://ardour.org/ <= a cross-platform open source DAW that has been around for more than 23 years
I'd love to chat! I'll make my way onto the forums. Hopefully there's a relevant spot to post about this :)
PS. Not sure if you're aware your email is not currently in your profile.
paul@linuxaudiosystems.com for email
You reviewed three DAWs. Glad you put the link to Admiral Bumblebee's page on yours, because that's actually "a list of DAWs".
Re: MIDI - you can already "play between the notes" using either MTS or even just simple pitch bend - the issue with the latter is building a controller that makes this better than the wheel on most keyboards, the issue with the former is finding synths that understand MTS and use it.
> no good method for hardware acceleration is a limiting factor in the audio world
DSP processors have been available for this since before ProTools, which was fundamentally based on hardware acceleration. In the present era, UAD and others continue (for now) to carry that torch, but the principle problem is that during the period where Moore's Law applied to processor speed, actual DSP could not keep up with generic CPUs (every DSP system was the speed of 1 or 2 generations ahead of generic CPUs). Current generic processors are now so fast that for time domain processing, there's really no need for DSP hardware - you just need a bigger multicore processor if you're running into limits (mostly - there are some exceptions, but DSP hardware wouldn't fix most of them either).
There's definitely more than 3 reviewed there? Sure, a lot of them still say [TODO] on getting a review, but then I don't want to review something I don't have a significant amount of 1-on-1 time with. Maybe you didn't realize there are click-able tabs?
Also, the way you phrased this came off as quite rude. There is someone on the other side of the screen, and you'd do well to consider that in the future.
> Re: MIDI - you can already "play between the notes" using either MTS or even just simple pitch bend [...]
Pitchbend is global to the track, if you play a chord and bend, all of the notes bend equally. With MPE it is possible to bend a single note, but then MPE isn't supported in every DAW (FL Studio doesn't have it) and my gripes with MTS are explained already: You're still working with only 127 possible notes so you limit your octave range. Worse, with all of the microtonal solutions, the UI will still typically look like a 12-tone piano. This is sort of okay for 24TET, but It's immensely confusing for anything else.
> DSP processors [...]
They really haven't been true for audio for a while now. Moore's law isn't dead, but definitionally "Moore's law is the observation that the number of transistors in an integrated circuit (IC) doubles about every two years" doesn't equate to all work loads seeing an improvement. Audio is a pretty latency sensitive, mostly strictly sequential workload. Making audio code that uses multiple cores often isn't even possible. Audio needs clock speed and IPC gains, which we have gotten, but that's not always enough. I have a 5900x and still hit limits. What would you recommend I do, get a Threadripper or Xeon so that I can have even more cores sit idle when making music? If anything, the extra cores have been a hindrance lately - on my 3900x I had before at high loads I had to pin the processes to one chiplet or I'd me more likely to get buffer underruns. It's not as if anyone is arguing that CPUs getting faster so quickly means that we don't need graphic cards.
UAD exists but then you're limited to their plugins and their accelerators are quite expensive for not being really all that powerful. I'm also not convinced that kind of accelerator is even the right approach. Field Programmable Analog Arrays, FPAAs, for example, could be setup with a DAC and ADC on either end. Or we could make DAWs/OSs capable of handling digitally connected analog "plugins" better - think effects like the Big Muff Pi Hardware Plugin [1] or Digitakt with Overbridge [2] (These are the only two examples I know of!). Using the word "Acceleration" was wrong, what I really meant is offloading. We need a way to offload the grunt work to something we can easily add more of or better fits the task. I think this is particularly true of distortions, as slamming the CPU with 4 or 8x oversampling to get a distortion to not sound awful hurts.
[1] https://www.ehx.com/products/big-muff-pi-hardware-plugin/ [2] https://www.elektron.se/us/overbridge
I agree with you about pitchbend, but you're narrowing what "play between the notes" means: you seeem to mean "polyphonic note expression", which is a feature that quite a few physical instruments (not just piano) lack.
MPE doesn't need to be supported by the DAW, only by the synthesizer. It's just regular MIDI 1.0, with different semantics. It's more awkward to edit MPE in a DAW that doesn't support, but not impossible. Recording and playback of MPE requires nothing of the DAW at all.
> the UI will still typically look like a 12-tone piano
We just revised the track header piano roll in Ardour 8 as step one of a likely 3-4 step process of supporting non-12TET. Specifically, at the next step, it will not (necessarily) look like a 12-tone piano.
> Audio is a pretty latency sensitive, mostly strictly sequential workload
It's sequential per voice/track, not typically sequential across an entire composition.
IPC gains are not required unless you insist on process-level separation, which has its own costs (and gains, though mostly as a band-aid over crappy code).
If you're already doing so much processing in a single track that one of your 5900X cores can't keep up, then I sympathize, but you're in a small minority at this point.
Faster CPUs don't help graphics when the graphics layers have been written for years to use non-CPU hardware. Also, as you sort of implicitly note, there's a more inherent parallelism and also decomposability of graphics operations to GPU-style primitives than there is for audio (at least, we haven't found it yet).
Offloading to external DSP hardware keeps popping up in various forms every year (or two). In cases where the device is connected directly to your audio interface (e.g. via ADAT or S/PDIF), using such things in a DAW designed for it is really pretty easy (in Ardour you just add an Insert processor and connect the I/O of the Insert to the appropriate channels of your interface. However, things like the BigMuff make the terrible mistake of being just another USB audio device, and since these things can't share a sample clock, that means you need software to do clock correction (essentially, resampling). You can do that already with technology I've been involved with, but there's not much to recommend about it. The Overbridge doesn't have precisely the same problem in all cases, but it can.
There needs to be significant number theory and computer science algorithmic work related to sound and how we represent sound with data.
GPU's currently cannot work with sound data in the processing chain, and multicore is basically just used to scale horizontally (ie to have more plugins or instruments)
New algorithms are needed to scale out audio processing, as well as make use of new hardware types (for example, using the gpu)
However, audio has a very different set of constraints from other types of workloads - the hallmarks being one worker doing LOTS of number crunching on a SINGLE stream of floating-point numbers (well, two streams, for stereo), that processing necessarily happening in SERIAL, and getting the results back INSTANTLY. Why serial? Because for most nontrivial audio processing algorithms, the results depend on not just the previous sample, or even chunk of samples, but are often a rolling algorithm that depends on a very long history of prior samples. Why instantly? Because plugins need to be able to run in realtime for production and auditioning, so every processing block has a very tight budget of tens of milliseconds to do all its work, and some of them make use of a lot of that budget. Also, all of these constraints apply across an entire track as well - every plugin on a track has to apply in serial, one at a time, and they need to share memory and act on the same block of audio.
One thing you might notice is these constraints are pretty bad conditions for GPU work. You're not the first to think of trying that - it's just not a great fit for most kinds of audio processing. There are some algorithms that can run massively parallel and independent, but they're outliers. Horizontally scaling different tracks across CPUs, however, works splendidly.
Saying that audio has "different set of constraints from other types of workloads" and giving up on fundamental algorithm research is just defeatist, throwing in the towel, and frankly really insulting to human advancement.
Come on, we need some new algorithms and just saying "whelp it can't be done" is kind of ... not the hacker spirit.
It could be that we need quantum algorithms for parallel processing what previously was thought to be serial. Just from reading your well reasoned paragraph, I can see we desperately need fundamental algorithm research in sound processing.
An imaginary scenario might to to invent an algorithm to convert/transform sound/pressure wave information into another domain, one that is not dependent on serial time, and then do the operations, and then re-convert it back to the time domain that we usually associate with sound processing. Where, within this alternative domain, parallel processing is possible.
We do stuff like this all the time in other disciplines. Even stuff like the FFT was an attempt to transform and make certain "unsolvable" problem solvable in another form.
That's the kind of math research that I'm referring to.
But for 90% or more of the things people do in DAWs (and currently want to do in DAWs), current processors can already do it fast enough.
So the sort of innovation your dreaming of/imagining isn't going to come from new algorithms - it's going to come from people wanting to do new things.
This is already happening to some extent with things like timbral transfer, but even there, the most important part of it is well within current processing capabilities.
> Come on, we need some new algorithms
If you don't have a "why", that doesn't make much sense. Start with "Come on, we need to be able to do <this>" and then (maybe) the algorithms will follow.
Necessity is the mother of invention, but so is desire. What do you desire?
A lot of the cutting edge of acoustic research is from people wanting to make materials and spaces do certain things with sound.
For example, this article describes a system that lets sound through one way This was described in https://www.scientificamerican.com/article/a-one-way-street-... ("an acoustic circulator") - the need will come from things like this.
And for musicians, artists, and instrument makers to make use of materials, devices, and spaces like this. That would be my answer to your question of where the need/why will come from.
Imagine needing to write a DAW module to deal with "one way sound" that results from using an acoustic circulator in a musical production or a song.
We already do this. It's called FFT, which transforms the data from the time domain to the frequency domain. You can, if you want/need to, parallelize frequency domain processing. There's oodles of interesting audio software that does this.
But again, parallel processing is only interesting for speed. And we mostly have plenty of speed these days.
https://www.nvidia.com/en-us/on-demand/session/gtcspring22-s...
> New algorithms are needed to scale out audio processing, as well as make use of new hardware types (for example, using the gpu)
What kind of "scale out" are you referring to here if not "to have more plugins or instruments"?
Bitwig is mentioned three times in the article.
That said, I don't think the article author would consider Bitwig to be substantially novel compared to the other mainstream DAWs that inspired the sentiment behind the article.
A Synthstrom Deluge, a humble Eurorack (morphagene!) and a microFreak ..
A Zynthian! Holy molies, live Zynthian compositions are binoculars!
1010music Bluebox+Blackbox.
SonicWARE SmplTek.
Monome. Oh dear, monome is a hell of a platform for next-generation DAW tomfoolery!
Notice a pattern? Its the hardware. You can no longer call it a workstation when, after all, it behaves like an instrument.
* a four part acapella group
* a traditional western european orchestra
* someone who wants a recording of the call to prayer
* anyone who needs to make a recording of anything at all?
The original purpose of DAWs like ProTools was recording, editing and mixing. Composition and instrumental performance crept in later. Even if that aspect were to be completely replaced by hardware (possible, but unlikely), it would leave the original DAW functionality just as in in demand.
Yes, that'd be a microphone setup per singer, probably, or at least two channels of audio.
>* a traditional western european orchestra
Yes, tons more channels ..
>* someone who wants a recording of the call to prayer
Only one channel needed really, but could be multiple if warranted.
>* anyone who needs to make a recording of anything at all?
As a professional designer of recording devices, I can tell where you are going - not that you might have missed the notion that there is literally no reason that a DAW-like instrument cannot be multi-channel capable - but also that portability to any environment is key.
Physicality is the new frontier for digital capture.
Recording/editing/mixing no longer belong in the workstation.
Portable instrument-like DAW's, with extraordinary new and interesting interfaces to accomplish those goals you've set (although they are all basically the same thing), are already on the market.
For example, the 1010Music Bluebox, alone, can be applied to any of those scenarios. Just add microphones (and USB powerbank...)
Incidentally, PaulDavis: acknowledging your position as the originator/BDFL of the ARDOUR workstation software, I mean no disrespect for your stature and point of view -- just that I believe era of the DAW-less approach is upon us, and there is a rather large opportunity for device-makers, such as me, and software-makers, such as you, to align ourselves...
The workstation is dead. The kids want portability, reliability, and power. This can all be done on non-standard operating systems, in a bespoke case, for fun and profit.
However, I don't think that they are going to eliminate the current concept of "a fairly big piece of software running on a relatively general purpose computer that is used for recording, editing and mixing".
Por que no los dos, eh?
I expect the future of the DAW is a Xeon-powered machine from Dell or HP with Adobe Creative Cloud (CC.)
There are some AI tools that work outside the main workflow, like for mastering after you're done with the DAW. But it's quite difficult to improve and bring new ideas beyond the typical signal processing modules without completely revamping the current workflow.