Neural networks emulate any guitar pedal for $120
hackaday.com
hackaday.com
Given the wide (and very non-linear) range of settings of a typical pedal, as well as interaction (impedance, etc.) with a real guitar and amplifier, it seems like it would be a pain to get all of the training data.
For a digital pedal, running the actual (e.g. Eventide) DSP code is just going to be better than some ML approximation.
On the other hand, I've been a bit dissatisfied with amplifier and cabinet models based on traditional DSP and physical modeling approaches, so maybe neural networks could fill in some of the gaps.
I wouldn’t discount ML. The nonlinearities are the bread and butter of modern ML models. In fact, two linear layers without a nonlinearity inbetween is equivalent to one big linear layer. So nonlinearities are required.
To put it another way, I would gladly bet any reasonable sum of money that in a double blind test, the listener wouldn’t be able to tell the difference from a genuine guitar pedal. (Not necessarily this pedal, but I suspect ML will model the effects more than adequately for human hearing precision.)
FWIW, I say this as someone who used to argue that graphics programmers were doing gamedev all wrong because they weren’t modeling light, they were approximating light. ML models were the way out.
I also think much of the problem is that ML devs often don’t have traditional signal processing experience, so they haven’t been modeling signals in quite the right way. (I’m trying to rectify that a bit with my FFT tutorials: https://twitter.com/theshawwn/status/1398796224921321472?s=2...) It remains to be seen, but Fourier space has recently been making strides in ML, and it’s likely much easier for a model to approximate a nonlinear waveform in frequency space than as a raw waveform.
To put it another way, if human speech is getting to the point where ML models can trick people, what are the chances that a future model won’t be able to do it for guitars?
DCT is on my radar. But there are several serious limitations that I think are overlooked. For example, convolution is no longer a simple component-wise multiplication. This seems, to me, a big deal.
Complex numbers are tricky to model, but I think most people have given up too easily, or haven't been creative enough in how they're modeling them. Some of my (outdated) ideas: https://gist.github.com/shawwn/c6865fccafac5066e1c7bab672781...
In other words, you're probably right, but I'm focused solely on FFTs on the (very low) chance that people have overlooked something that will work well.
Maybe we can talk later? Not sure.
I guess I didn't get my point across. What I meant was that pedal settings tend to be non-linear with multiple sweet spots (which often depend on the guitar and amp) so you shouldn't just do a linear range from 1-10^N (where N is the number of knobs) for training data, as someone else had suggested. Moreover, there are also dependencies on the impedance chain, gain structure, feedback, reflections, etc., which seem well-suited to circuit and physical modeling. Digital pedals, as I note, are largely software anyway so it doesn't make sense to me to try to model them with ML any more than it does to model Microsoft Word using ML (though I'm sure someone has tried.)
In general ML seems most useful when you don't have good analytical models - but in the case of circuits and software we have very good analytical models.
Have you tried the Fractal stuff? I've been using it since the first generation (consistently for live use since the Axe-Fx II days) and they've been ahead of the competition since their inception. At this point I'd venture to say that the majority of their amp models sound indistinguishable from the real thing with no advanced parameter tweaking.
That being said when I was between Fractal units about a year ago I spent a brief amount of time with the Pod Go and was immensely impressed with how much they were able to pack into a $500 floor unit. Most of the amps and drives still felt a bit caricature-y but were still very usable - a far cry from the Pod Bean days. It's truly a great time to be a musician.
Source: https://youtu.be/MhnkGMBn1tw
https://acris.aalto.fi/ws/portalfiles/portal/41964332/Real_t...
those are MUSHRA tests, meaning only skilled listeners are allowed to participate https://en.wikipedia.org/wiki/MUSHRA
TLDR: the resulting neural networks for given amplifier models were rated as “excellent” by the listeners, some even outperformed the reference!
Contrast with the Fractalaudio approach of modeling each component of a device. Fractal's AxeFx is the gold standard and any geek would gush over the HW and SW engineering. The best part is that the company owner keeps improving his algos and pushing out updates for free. This device costs the equivalent of a good amp head and is loaded with more amps and effects than any of the competition.
Sorry if this sounded like an ad, but I am always surprised how little airtime this amazing product gets in hacker circles.
Modeling real gear is all fine and dandy, but what really has potential is being able to replicate circuits that aren't viable in the real world, like a tube amplifier on the edge of occilation or one run way over the rated TDP of it's tubes, situations that might provide sonic possibilities that aren't viable long term can now be stable and replicatable long term. Additionally, being able to modify a tone stack in software to provide extra flexibility is very handy.
The Quad Cortex is also slated to receive updates that will allow it to capture modulation effects, and will undoubtedly get better with time as the Kempers did.
Umm. You and I are just random people on the internet. I have spent a lot of time trolling the places pro musicians talk about the leading devices and I have found that Fractal is very widely considered to be the best of the best as far as sound quality goes. Kemper is known to be easier to use and Neural is the new kid on the block that has bluetooth, touch screen on device and footswitches that act as knobs. For the limited # of tones you get, it is supposed to sound great. None of those features are advantageous to me at all.
>being able to replicate circuits that aren't viable in the real world
That is definitely something that Fractal does. Considering the two devices you are talking up are based on capturing actual tones from real world devices, I am unsure of your point here.
>capture modulation effects
Modulation and time based effects have been modeled to perfection in the digital realm for a long time now. See the ubiquity of Strymon, etc. Fractal has equally good algorithms and allows incredibly intricate and signal paths as it is an all-in-one device. I have four expression pedals and 10 switches that can be programmed to control any parameter I would desire with a couple clicks of a mouse. No other multi effects device brings this degree of controllable complexity.
The emulation ased on the model runs on a piece of hardware which has some generic controllers on it: some rotary encoders, a couple of switches and whatnot. These get assigned to the parameters of the model.
You could have a MIDI input on it, and use MIDI controllers, which would be cool. There are MIDI foot controllers that you can tilt with your foot to vary a parameter.
The different models could be assigned to MIDI program numbers. You could change the patch number with the foot controller, and vary the parameters with it also.
The foot controller might have, say, only two pedals, so you have to assign which ones you want: if the patch has five parameters, you have to fix the values of three of them and map the two important ones to the foot pedals you have. For the others, you can bend down and tweak the knobs on the unit itself.
this approach is not scalable (that's why the high cost)
the ML approach doesn't require any knowledge about the system (black-box) to produce the result
It may be opposite, most of the amps follow some classic schematic (e.g. jtm, plexi, princeton) with insignificant changes, so after building digital copies of some limited number of classical amps they can add new one rather fast.
As result, fractal has about 100 high quality models already (average guitarist probably uses 5?).
> the ML approach doesn't require any knowledge
ML approach requires you to capture training data: which is different sound samples with all possible knobs positions (8 knobs per amp in average) + different types of speakers and mics, which is very huge number of variations.
> that's why the high cost
cost is driven by market, ML profiling competitors (kemper, neuraldsp) charge about the same for their devices.
correct, ml requires data, but you don't need to capture every possible position to do a good prediction
only a handful would suffice and let's be honest, how many presets do you really need (average guitarist probably uses 5?)
This is very different and more narrow use-case.
Also they have tone match like forever, you build signal chain close to your amp, then add tone match block, which applies ML to voice your digital signal chain close to recording.
Here is example: https://www.youtube.com/watch?v=hZnZ1nJODLo
Also in this example he didn't profile actual amp, but actual AC/DC recording, and result is very good I think.
> how many presets do you really need (average guitarist probably uses 5?)
But how you find this preset for your signal chain (guitar + speakers)? That's one of big points of frustration with kemper: one needs to go through hundreds profiles (not necessary good quality) to find one which will sound good with his signal chain.
With fractal: you take some basic preset, and change knobs to your tastes and goal as with real amp.
I believe you can already create your own impulse responses (from your personal amps).
This is fine if you like that sort of thing. I would note that latency is very important here: it's not going to be nice to play through if it's incurring any significant latency.
As the creator of this project I can assure you that the audio used here is at least CD quality (44.1kHz 16bit). With the HiFiBerry hat the digital audio comes in at 24bit/192kHz. The NeuralPi DSP processes the audio at 44.1kHz with 32 bit floating point precision. No reason the sample rate can’t be higher though. Elk OS claims latency is less than 1ms, but I’d like to test and see exactly what the latency is running the plugin. As a guitarist, I can’t tell the difference between this and an analog effect.
runs and ducks for cover
Not necessarily... why? We have been very successful with amp / pedal modeling through regular DSP methods. You'd be very hard pressed to find an album that doesn't use one nowadays. What makes NN methods fundamentally different?
>it's not going to be nice to play through if it's incurring any significant latency.
Yes but this is a non issue for many of the current systems, actually since a couple decades ago. 1-2ms latency is pretty achievable, especially with a RT kernel. That is the natural latency of a sound source about 1 meter away from you.
But it's pretty hard to beat elections flowing through an analog circuit, when in the digital side you have to: convert analog to digital, run through the kernel to get to user space, run the bits through the RNN, send back to kernel space, convert to analog and finally send to an output.
In order to be competitive with a 1980s guitar pedal, you have to do all of that in under ~10ms latency, and that's just really hard still.
Even though we carry around these super computers in our pockets these days, there are still some things left where analog still beats the pants.
> In order to be competitive with a 1980s guitar pedal, you have to do all of that in under ~10ms latency, and that's just really hard still.
Well, don't use a whole PC with software stack. A custom embedded solution with DSP can easily manage.
> Even though we carry around these super computers in our pockets these days, there are still some things left where analog still beats the pants.
Again, it's not digital vs analog, it's massive software stack versus embedded hardware/software solution.
Things like Clavinova keyboards Piano or Line 6 Pods simulating Guitar/Bass amps & effects have been out their for decades now.
And while they have been quite popular due to the sheer number of sonorities and the convenience they bring (possibility to play with an headset, extremely useful to play at night or in apartments), traditional analog setups remain strong.
Playing on an analog setup still is more pleasant and more expressive IMHO, in particular "simulators" tend to mask the attack when hitting a note, and hides a lot the tension/crispation in the hands/fingers when playing, leading to potential bad habits, specially for people learning to play an instrument. Analog to digital and digital to analog conversion definitely lead to loses in expressiveness.
traditional methods were successful in emulating sounds, but they fail short at replicating the feel and response of the hardware they're trying to replicate
that's why you need neural networks!
Honestly, the only OS that's truly bad at this is Windows. DirectSound is laggy and highly limited in it's capabilities, and even a nice ASIO won't fully alleviate your issues. Your best bet is to get a DAC and hope for the best. Besides that, I've found Linux and MacOS to be very similar in terms of latency, out of the box. However, I've found that tuning Linux with a custom low-latency kernel absolutely destroys CoreAudio's latency. Given that it's something most people won't be doing, I think it's fair to say that both OSes are tied, but I still give the edge to Linux for having a more modular and adaptable sound backend.
The answer would probably be to reduce the 'learned' output to be a convolution kernel that gets run rather than the RNN itself on the input. Then the kernel only has to change gradually to produce a different sound not continuous processing to produce a particular sound.
This has been solved for over a decade. Linux is a bit tricky but MacOS and Windows have native low latency drivers. There are also a ton of digital effects units running at 0ms on the market.
There's also another element: if you have, say, a vintage Fender Champ and a Klon (or whatever) it's because you mean to project different expressions through your string handling and note-playing. At that level you've made a best effort to produce the most emotionally transparent and responsive signal chain, which you will then not think about once you've got it turned on and tweaked: ALL the settings are liable to sound 'good' and respond for you.
The modeling approach is so often "This is exactly that, but better, because here are twelve other Fender Champs and models of Klon to choose from!" and when the first claim isn't as true as we would like, and the second is a distraction and time-sink, that's not great.
I can tell when I've chosen wrongly in my music-making tools, because I flat-out stop making music. Even in a dilettantish way: it just stops being a thing. That's a concern.
The only strength of analog processors is that they're dedicated. That an iPhone SoC and run dedicated audio-only OS on it, and there we go.
You can just use a Pi with bare metal code or a real time OS.
Can anyone give a rough idea of the actual limitations? I would guess that there is a limit to how non-local the effects it can manage are.
extern char* turing_machine;
int main() {
process_turing_machine(turing_machine);
beep();
}If you aren't familiar with RNNs, think about it like a NN that instead of learning a input -> output function, learns a (input, state) -> (output, newState) function
A call to `y, hidden = layer.forward(x)` (where x has a batch size of 1, and an arbitrary length) produces two hidden states of dimensions `(1, 1, hidden_size)`, where hidden_size is the exact number you passed to the LSTM constructor. Those two states represent the long term and short term memory features.
You would need to have an LSTM with hidden_size large enough to store the samples (or a compressed representation) of your entire loop. Not to mention you'd run into other issues with handling the logic around variable length loops based on a pedal toggle.
But, yeah, at some point your signal has such a complex behavior on long time scales that there isn't a good way to predict it based on a limited state size (or at least gradient descent can't find a function to predict it for you).
If the future input samples have a meaningful impact during loop playback, then it hasn't learned the correct behavior of the original loop pedal.
Note that the linked project appears to use a hidden size of 20. Twenty floats. With that much space we're very much back to "sure, you might theoretically be able to loop if the information fits in the hidden size".
Increasing the hidden size beyond 20 still won't solve learning the complex state machine behavior of an original loop pedal, which can loop variable length audio. You'd need to provide the pedal state to the network in addition to the audio, and probably train need to train it on a bunch of different loop lengths (>thousands?).
This would mostly be an academic pursuit, as it's extremely impractical compared to the other uses of the device.
The hard part is that there are some fundamental limitations to deal with. The biggest is aliasing - distortion effects in particular deal with enormous amounts of distortion (> 100% THD) which creates spectrums far outside the range of hearing. Digital audio systems need to have high orders of oversampling to prevent audible aliasing (8-16x is not unheard of!).
After aliasing is memory. It's too early in the morning for me to do math but I'm almost certain you can't model a looper with a causal NN that has less internal state memory than the length of your loop. Doing so is dumb anyway, since loopers are pretty trivial and their biggest cost is memory. Same goes for digital delay and modulation effects, the algorithms are not expensive.
A 3rd or 4th order polynomial interpolator is pretty darn good and doesn't need a NN to find the coefficients.
Oversampled DSP algorithms work by oversampling, performing the nonlinear processing, then filtering to remove the higher harmonics, and finally downsampling again. We do it this way because it's convenient and easy to understand and based on proven mathematics. But nothing says these steps have to be distinct.
An oversampled DSP algorithm looks like a regular DSP algorithm from the outside, perhaps with some more state and latency required for it to perform the internal oversampling. You can also imolement such an oversampled algorithm entirely at the original sample rate clock; it just means the processing needs to internally process several samples per outer loop sample.
Since neural networks excel at modeling "black boxes" as one amorphous blob that we don't understand, I wonder if a NN could learn to model such an internally oversampled algorithm fairly accurately, and what the computational complexity would be.
Since you can model the oversampling/filtering/etc steps as linear convolutions with wider internal state at the original sample rate, I'm almost certain this will work with the right NN topology. It's obvious an NN can implement oversampling.
And so my question is: could treating the combined oversampled processing as one step, and training a NN on that, potentially result in a more efficient implementation than doing it naively? Especially for heavy distortion that needs high oversampling ratios.
It can’t do time based effects such as delay, reverb, flange, chorus, etc. The LSTM can model distortion, overdrive, compression to an extent, and amp circuits (including vacuum tubes).
The reason I stopped using it is that it is not that the fuzz doesn't react to the guitars volume pot like a Fuzz Face or because the Tubescreamer was a TS9 not an 808. It was because having a single box with all of your effects in is a faff to tinker with. I like the dedicated hardware of my pedals, I like having the right number of knobs. I like being able to turn off the fuzz but keep the delay by stepping on the fuzz switch. I like being able to run the TS before or after the fuzz and to see it happen. I mentally am much more at home with little boxes for each stage, it is like a real world flow chart!
I'm running a hybrid setup, using a Line 6 Helix in front of a Soldano SLO 30 (and on its effects loop as well). Each effect can be adjusted (with all sorts of knobs, more than 12 for some), moved around and triggered individually or together. You can even use the Helix to toggle channels in the amp, manually or as part of presets.
It drastically eased up my workflow to experiment with and mix effects. No more tinkering with cables. No more wondering which pedal is plugged in wrong or doesn't have enough power. No more velcro strips. Plus, selling my hoard of pedals felt nice.
I'm not a musician, but even I could easily tell the difference a lot of the time. It didn't sound bad, but there were clearly aspects the Neural Capture failed to, well, capture.
Still, early days.
[1]: https://towardsdatascience.com/neural-networks-for-real-time...
[2]: https://towardsdatascience.com/neural-networks-for-real-time...
I've been using Jetsons as RasPi replacements wherever I can, they are not only more capable but also much more reliable than RPI4s, in my not so limited experience.
A board with a m2 port costs an extra $300 and only exists since last year: https://antmicro.com/blog/2020/04/updated-jetson-nano-xavier...
It should have been standard and by default from the beginning.
They can also boot from USB3 as of recently, which boosts storage access speed tremendously.
1 - Kemper Profiler
2 - NeuralDSP
Both of them are above 1000€/$. We are talking about 10x the price of this thing. Add some Multi-FX Pedal (like Line6 HX-Stomp) where you put this in the FX-Loop and you end up with something equally good for still half the price.
And in general, its not about what is better, digital or analog. It's about the use-case. In the studio or when noodling around with the knobs when practicing: Real Amps and Pedals. But on stage, you don't play with the settings of your pedals or you want presets. This is where you go digital. No one will notice the slightly different sound there anyway.
That said, if you use a preset based effects setup like rackmount gear, I could see how this would be cool.
I know some who change their settings during playing, but I wouldn't consider this as standard. Also, you can still add a real Pedal to a MultiFX Board if you really want to change its settings live. On the contrary, automatically changing the settings controlled by your DAW and triggered via MIDI into your MultiFX unit is also pretty common. You don't have to hassle with your Effects at all and can concentrate on playing, what you should, as you need to be absolutely on point.
Given that this was the only really active process on the pi, it ended up working really well. I simply converted modules I had written for vcv rack.
I also built one that enabled usb host mode and acted as an audio device that worked with any daw. Ended up being pretty cool for about $20 of parts.
While not as cool as an ml system, given that I was already writing dsp, it ended up being pretty neat.
another is add more params to the models to take the knob states into the account when doing predictions (which can impact the performance)
And yeah the sonic possibilities are endless. On a physical pedalboard, it's pretty involved to rewire everything to change the routing, or add/remove pedals, etc. With the digital modeling ones, this all becomes trivial and you can try all kinds of different setups much more quickly, save them and go back later, share them with friends, etc.
All that said, it's kind of like asking an acoustic guitarist in the 1950s why they would use an electric guitar. Electric guitars obviously have lots of advantages, but it's not like people stopped playing acoustic guitars. I still think real analog pedals are cool, they're fun to collect, in some cases they sound better, etc. And sometimes you don't need the mega-flexibility -- if you have a 4-pedal setup that does "your tone" maybe that's all you ever need.
But as an engineer I just nerd out over all of it. Analog. Digital. If it makes good music it’s all cool to me.
That's my primary motivation. On my AxeFx, I have signal chains that are impossible with real gear. If I want to tweak that chain, it is a couple clicks of a mouse. I have four expression controllers and 10 foot switches that are tied to different parameters. I can tweak this functionality on the fly. All in a single $2k box.
Beyond all, tube amps are stupid for bedroom players as they generally require gig level volume to get the killer tonez...
I have had modeler amps in the past, but every so often I'd just nerd away with the dozens of available amp models and the myriads of settings, and come out with the dissatisfied feeling of having just wasted a lot of time instead of engaging with music.
I find modelers have their place when it comes to replacing a set of analogue pedals, which is the reason I traded my three Boss pedals (compression, reverb, delay) for a Boss GT-1000Core. Overkill for my purpose, as I really never use any amp sims, cab sims, of any of the advanced signal chain stuff that the device is capable of. I just have one patch with my three pedals for practice, going into the AC4, occasionally turning the same knobs as on the analogue versions, but enjoying that I now have built-in a tuner as well :-)
you can run the produced models in browsers, on mobile, on embedded hardware such as in this case and even on calculators (joking)
Because if those two things were squared away, I could see this being an extremely viable project.
Another benefit to a software defined pedal is that it can express sounds that cannot be replicated in analog. Emulation is boring. Train it to do stuff that I can't buy in a pedal!
this thing runs on consumer-grade hardware
Are they perfect emulations? No. Did anyone notice that in my music? No. Would I take that rig on stage? Also no.
It did help me find sounds I liked though, and over the years I've bought hardware equivalents of some of my favorite emulations, and I've bought hardware that goes beyond anything Line 6 can do.
As to whether Line 6 is cool or not, NIN/Trent Reznor toured with Line 6 rigs to sold out arenas 20 years ago, and today the live rigs are managed by the FOH using software emulations that can be automated or managed from on-stage controllers. Maybe you don't think that's cool, but the important takeaway is that you should use the tool that's right for you.
I think it's good to see people tinkering with new ways to reduce the cost around modulating audio signals in interesting ways.
While you can emulate any guitar pedal with AI and you don't have to pay 120$ I'd like to introduce you to some understanding.
MUSICIANS love to purchase shit they can show off to others. it's in their ego nature. While 120$ for software might compete with quad pro cortex pedal, you can't show off with 120$ but with a pedal you bought fo 2g you can. That is why BOSS made a big travel box for physical pedals this year, this is what they released at the NAMM2021.
How do I know that? I make software for them. GarageBand has 1000s of different pedals already and those are free. most musicians have apple or abelton live or some similar.
Also, not to discourage anyone - live music is dying. More and more you see musicians dancing on the stage with their instruments pretending they are playing, while pro-edited backtrack is going through the speakers.
If you want to make this world a better place think how to improve housing/farming/healthcare with technology. Things like roof above the head, food on the table will be always in demand. There is no money to be maid with musicians, cause they don't have any. it took me few years to realize, you can learn on my mistakes.
If you have some cash and time and thinking if you need to build some software product - kill that, buy some land and build a cabin. In the end you'll have another property. you can sell it or rent it. you can live it to your kids, there is value there. There is no value in your software idea unless people use it like crazy, and they really won't cause they already have 100 pedals and that new pedal bag from boss
I've seen entirely analog circuits designed and built with the PCB parts and techniques used in these devices, that are miles away from the quality of good pedals, without even being digital or modeling at all. I'm sure the problem circuits measure completely fine, and then you A/B them with a truly great pedal and it's chalk and cheese.
If you try to argue the point, people committed to the digital modeling model get quite fierce, so tone hounds learn to just roll their eyes and not engage. You can also think of it as learning to rely on secret weapons.
Case in point: the old Neve desks
https://acris.aalto.fi/ws/portalfiles/portal/41964332/Real_t...
those are MUSHRA tests, meaning only skilled listeners are allowed to participate https://en.wikipedia.org/wiki/MUSHRA
Sure if you run a null test these will fail, but in real life it's really up to how honest the tone hound in question is, and if you can trap them into being honest.
If you're talking product viability, something like the Pepsi challenge, double blind testing, would probably be effective marketing.
There have been a few recent advancements lately (Boss SY-1), but even the supposedly "ideal" solutions, that require a new polyphonic pickup, are not good at all. I have a Fishman Triple Play and a plugin whose name I forgot, and tracking is frankly terrible.
Here's what Metheny said himself: "But the guitar‑to‑MIDI part has always been a problem. It's a question of physics. On input, I sort of have to rush. But I know how to rush. I play ahead."
It's fun for lots of things, and you can make lots of cool music, but it's still very limited to certain styles, dynamic ranges, phrasing, tempo, speeds...
But "solved" here means "when not doing the analysis in real time". The realtime solutions are not as good. NN's are not typically great at realtime either, so this may not help very much with this particular goal.
I'm just disputing the GP general assertion about NNs that "NN's are not typically great at realtime either", which is quickly disproven by the TFA which uses NN for realtime audio.
Am I missing something?
"It’s a well-established fact that a guitarist’s acumen can be accurately gauged by the size of their pedal board- the more stompboxes, the better the player."
As both a software engineer and a guitarist, I'd say the opposite is true. Or at least truer. You can't do math-rock without a lot of pedals but the hard part is to acquire the chops. A lot of pedals, and production effects generally, quickly become cliches and it's like dropping down a musical black hole. Someone like Hendrix could take a new effect (superset of pedals) and make it work musically brilliantly but most pedal users buy the pedal to get someone else's sound.
GarageBand has 1000 of pedals, people don't need more pedals, they need to practice more, but they don't understand that. That's why they are buying more pedals while keep dragging about how much they care about green/clean planet.
"Jokes aside..."