HNHacker News
TopNewBestAskShowJobs

abdljasser2

157 karma · joined February 24, 2022

submissionscomments
abdljasser2··on Show HN: Neurosymbolic music generation with Claude Opus 4.5
code here: https://github.com/erl-j/neurosymbolic-music-generation
abdljasser2··on Ask HN: SWEs how do you future-proof your career in light of LLMs?
My plan is to become a people person / ideas guy.
abdljasser2··on Maximum likelihood estimation and loss functions
Excellent post
abdljasser2··on Show HN: I made an AI soundfont generator
Thank you! I can also recommend this work which had a similar idea:

https://instrumentgen.netlify.app/

abdljasser2··on Show HN: Downloadable AI Musical Instruments
Good question. In my experience combining generic descriptors is what works best. This is probably due to the text captions used during training mostly consist of generic instrument names, genre names and adjectives.
abdljasser2··on Jsfxr: 8-Bit sound maker and sfx generator
https://erl-j.github.io/textsynth/

Sortof ML, Evolutionary algorithm w/ fitness determined by a ML model.

abdljasser2··on Show HN: Cloning a musical instrument from 16 seconds of audio
[4] https://forum.juce.com/t/ddsp-tone-transfer-vst-possibility/...
abdljasser2··on Show HN: Cloning a musical instrument from 16 seconds of audio
[4] https://forum.juce.com/t/ddsp-tone-transfer-vst-possibility/...
abdljasser2··on Show HN: Cloning a musical instrument from 16 seconds of audio
Hello! I think this will be possible!
abdljasser2··on Show HN: Cloning a musical instrument from 16 seconds of audio
I can't say for sure but I think the issue with why the reverb sounds off in that particular example is that the reverb present in the recording has a longer decay than the maximum reverb duration I set for the experiments (1s). I will set it longer in the future.

Yes, that kind of visualisation could be performed. The authors of the prior work we are building on (DDSP) have made some great visualisations which I think you will find useful! Here: https://storage.googleapis.com/ddsp/index.html

abdljasser2··on Show HN: Cloning a musical instrument from 16 seconds of audio
Thank you! Will do.
abdljasser2··on Show HN: Cloning a musical instrument from 16 seconds of audio
Very interesting! Thank you for this.
abdljasser2··on Show HN: Cloning a musical instrument from 16 seconds of audio
Very cool! Thank you for this.

Speaking of wavetables, have you seen this relatively recent work?

https://lamtharnhantrakul.github.io/diffwts.github.io/?

abdljasser2··on Show HN: Cloning a musical instrument from 16 seconds of audio
Hello! I think you accidentally mistook the first citation link for a link to the work we are presenting.

Here is the link I think you are looking for:

https://erlj.notion.site/Neural-Instrument-Cloning-from-very...

Thank you!

abdljasser2··on Show HN: Cloning a musical instrument from 16 seconds of audio
Great question! I think this approach (and DDSP more generally) could work well for most instruments if you make some adjustments to the synthesis architecture.

The current architecture imposes some strong assumptions on the source (such as harmonicity) which does not necessarily apply to all instruments (bells for instance). The harmonic + noise model is not that good with transients either (which is a key ingredient in piano sounds) but of course you could solve that through different means.

abdljasser2··on Show HN: Cloning a musical instrument from 16 seconds of audio
Hello! I think this can be done but I would not advise you to do it with this particular algorithm.

I think you would be interested in the work of Jatin Chowdhury[1] as well as the work of Christian J. Steinmetz[2].

[1].https://arxiv.org/pdf/2009.02833.pdf [2] https://arxiv.org/abs/2102.06200

abdljasser2··on Show HN: Cloning a musical instrument from 16 seconds of audio
Copy pasting my a reply to a similar question

Short answer is yes! Previous work has shown that we can obtain very good results from controlling DDSP models from midi input. The solutions I am familiar with employ a two stage approach where the first stage takes midi and turns it into control signals (pitch & loudness contours etc..) and the second stage turns the controls signals into audio (like the particular model I discuss in the blog post)[1][2][3]. I actually think that the first stage could also benefit from the transfer learning techniques we discuss in the blogpost. In terms of actually releasing a MIDI playable VST plugin I believe that Magenta have something like it in the works[4]. I hope that it will come with some ability for users to quickly create their own instruments, presumably using a transfer learning technique similar to the one we have presented.

Real-time rendering poses multiple challenges. For one, some instrument sounds occur before a note properly onsets (for example the sound of the fingers pressing the keys of a saxophone occurs before the first note of the piece). Secondly, the research models are quite heavy and considerably more compute intensive than a standard VST instrument which poses a problem if you want to use it inside a DAW. I think this latter problem can be solved with some clever engineering and the general trend of hardware being more and more accommodating to machine learning applications.

[1] https://erl-j.github.io/controlsynthesis/#/ (Our previous work) [2] https://rodrigo-castellon.github.io/midi2params/ (Focuses on realtime rendering) [3] https://arxiv.org/abs/2112.09312 (Magenta's recent paper on the subject)

abdljasser2··on Show HN: Cloning a musical instrument from 16 seconds of audio
Short answer is yes! Previous work has shown that we can obtain very good results from controlling DDSP models from midi input. The solutions I am familiar with employ a two stage approach where the first stage takes midi and turns it into control signals (pitch & loudness contours etc..) and the second stage turns the controls signals into audio (like the particular model I discuss in the blog post)[1][2][3]. I actually think that the first stage could also benefit from the transfer learning techniques we discuss in the blogpost.

In terms of actually releasing a MIDI playable VST plugin I believe that Magenta have something like it in the works[4]. I hope that it will come with some ability for users to quickly create their own instruments, presumably using a transfer learning technique similar to the one we have presented.

Real-time rendering poses multiple challenges. For one, some instrument sounds occur before a note properly onsets (for example the sound of the fingers pressing the keys of a saxophone occurs before the first note of the piece). Secondly, the research models are quite heavy and considerably more compute intensive than a standard VST instrument which poses a problem if you want to use it inside a DAW. I think this latter problem can be solved with some clever engineering and the general trend of hardware being more and more accommodating to machine learning applications.

[1] https://erl-j.github.io/controlsynthesis/#/ (Our previous work) [2] https://rodrigo-castellon.github.io/midi2params/ (Focuses on realtime rendering) [3] https://arxiv.org/abs/2112.09312 (Magenta's recent paper on the subject)

abdljasser2··on Show HN: Cloning a musical instrument from 16 seconds of audio
Thank you! It was made with figma.
abdljasser2··on Show HN: Cloning a musical instrument from 16 seconds of audio
I think this is a very interesting question which I currently don’t have the answer to. This is something we hope to answer in the upcoming paper.
abdljasser2··on Show HN: Cloning a musical instrument from 16 seconds of audio
Thank you for the feedback! We will synthesize longer excerpts in the future.

For the time being there is a colab where you can play with a pretrained model.

abdljasser2··on Show HN: Cloning a musical instrument from 16 seconds of audio
Hello!

Immediate use case would be sampling. Say you like a certain sound in a song and would like to use it as a starting point for your own sound patch.

I also believe that transfer learning has benefits even for making great sounding instruments in cases where you have access to lots of data. That’s my intuition at least.

At the very least, it saves you a lot of memory/bandwith. Instead of having one large model per instrument you only need one large models with a few extra instrument specific weights.

abdljasser2··on Show HN: Cloning a musical instrument from 16 seconds of audio
Hello! Thank you for the interest. There are links to a colab etc at the bottom of the blog post!