157 karma · joined February 24, 2022
Sortof ML, Evolutionary algorithm w/ fitness determined by a ML model.
Yes, that kind of visualisation could be performed. The authors of the prior work we are building on (DDSP) have made some great visualisations which I think you will find useful! Here: https://storage.googleapis.com/ddsp/index.html
Speaking of wavetables, have you seen this relatively recent work?
Here is the link I think you are looking for:
https://erlj.notion.site/Neural-Instrument-Cloning-from-very...
Thank you!
The current architecture imposes some strong assumptions on the source (such as harmonicity) which does not necessarily apply to all instruments (bells for instance). The harmonic + noise model is not that good with transients either (which is a key ingredient in piano sounds) but of course you could solve that through different means.
I think you would be interested in the work of Jatin Chowdhury[1] as well as the work of Christian J. Steinmetz[2].
[1].https://arxiv.org/pdf/2009.02833.pdf [2] https://arxiv.org/abs/2102.06200
Short answer is yes! Previous work has shown that we can obtain very good results from controlling DDSP models from midi input. The solutions I am familiar with employ a two stage approach where the first stage takes midi and turns it into control signals (pitch & loudness contours etc..) and the second stage turns the controls signals into audio (like the particular model I discuss in the blog post)[1][2][3]. I actually think that the first stage could also benefit from the transfer learning techniques we discuss in the blogpost. In terms of actually releasing a MIDI playable VST plugin I believe that Magenta have something like it in the works[4]. I hope that it will come with some ability for users to quickly create their own instruments, presumably using a transfer learning technique similar to the one we have presented.
Real-time rendering poses multiple challenges. For one, some instrument sounds occur before a note properly onsets (for example the sound of the fingers pressing the keys of a saxophone occurs before the first note of the piece). Secondly, the research models are quite heavy and considerably more compute intensive than a standard VST instrument which poses a problem if you want to use it inside a DAW. I think this latter problem can be solved with some clever engineering and the general trend of hardware being more and more accommodating to machine learning applications.
[1] https://erl-j.github.io/controlsynthesis/#/ (Our previous work) [2] https://rodrigo-castellon.github.io/midi2params/ (Focuses on realtime rendering) [3] https://arxiv.org/abs/2112.09312 (Magenta's recent paper on the subject)
In terms of actually releasing a MIDI playable VST plugin I believe that Magenta have something like it in the works[4]. I hope that it will come with some ability for users to quickly create their own instruments, presumably using a transfer learning technique similar to the one we have presented.
Real-time rendering poses multiple challenges. For one, some instrument sounds occur before a note properly onsets (for example the sound of the fingers pressing the keys of a saxophone occurs before the first note of the piece). Secondly, the research models are quite heavy and considerably more compute intensive than a standard VST instrument which poses a problem if you want to use it inside a DAW. I think this latter problem can be solved with some clever engineering and the general trend of hardware being more and more accommodating to machine learning applications.
[1] https://erl-j.github.io/controlsynthesis/#/ (Our previous work) [2] https://rodrigo-castellon.github.io/midi2params/ (Focuses on realtime rendering) [3] https://arxiv.org/abs/2112.09312 (Magenta's recent paper on the subject)
For the time being there is a colab where you can play with a pretrained model.
Immediate use case would be sampling. Say you like a certain sound in a song and would like to use it as a starting point for your own sound patch.
I also believe that transfer learning has benefits even for making great sounding instruments in cases where you have access to lots of data. That’s my intuition at least.
At the very least, it saves you a lot of memory/bandwith. Instead of having one large model per instrument you only need one large models with a few extra instrument specific weights.