Generate unique drum samples using artificial intelligence
audialab.com
audialab.com
I am fully aware of apple’s drummers in GarageBand and Logic, but they’re weak in my opinion. I’m looking for a quantum leap forward in the same way GPT chat makes Siri look like a toy.
If you’re curious, our original generative models used GANs, and we’re incorporating diffusion approaches now.
Drums are the start - we’re currently training models for instruments, synths, vox, foley, etc.
There is no shortage of samples and sample packs (millions), and most pros are picky : context in king, and style transfer is more contextual.
For instruments/synths/vox, "playability" is important, so the best approach IMO is cloning a sample to playable Midi instrument like midi-ddsp, midi2params or Mawf
I couldn't find a license or terms of service on the website, so I must ask: who owns the copyright on the generated samples?
Perhaps "royalty-free" are the terms they're offering if the law does eventually allow an AI (or its creators) to own a copyright.
It's too bad the cofounder hasn't answered.
In a similar vein, it would be very cool to use AI with synthesizers. Text to synth, but use an actual synthesizer as an intermediary step. Don’t just create a sound from nothing, tune the knobs in the right ways. Start out with subtractive synthesis and additive synthesis.
For a drum sample, I guess some are so generic and simple that no one can claim ownership. Still, if you manage to reproduce a particular sample, perhaps because of an overfitted model, then you may have some issues. However, the music industry seems to be fine with sampling.
Yes, as long as the sampled are paid and credited.
I produce hiphop and those 909 sounding drums seemed a bit bland for me.
I think the answer comes down to the fact that it would take, at minimum, one highly skilled audio engineer with specific audio programming skills to develop reliable variations that sound good. The sounds need to be realistic and appropriate literally every time, despite also being "random". This process would be deceptively hard, and take a significant amount of time.
How great or necessary is the payoff in most cases? For example, Dota 2 is widely regarded as having superb audio engineering - experienced players can tell what's going on even in frantic ten player battles involving 100+ different heroes and hundreds of abilities by ear alone. Dota doesn't use any procedural generation at all, and I've never heard anyone complain despite the thousands of hours they put in.
There is only so much you can do with pitch changes and creative use of filters, and having multiple samples for each sound can burn memory (and sound designer hours) fairly quickly. As you point out there is a wide variance in the technical skills of sound designers (some love playing with the game engine side, some are better at making sounds and throwing them over the wall).
Interesting that Audialab is focused on drums only.
No need for rolling a dice when you can generate all. All one needs is a taste which will be compatible with niche audiences.