Making music using new sounds generated with machine learning
blog.google
blog.google
Case in point: AudioSculpt manual, IRCAM 1996 (p62) (Second edition!)
http://homepages.gold.ac.uk/ems/pdf/AudioSculpt.pdf
Machine learning is gonna give us some awesome new sounds but you've got to do your research first on what's been done.
Of all the computer synths I tried out back in the day, Absynth was a standout. BUT it's hard to 'get physical' with a computer keyboard ... let alone a finger on a tablet. When the technology gets in the way, musicians run away.
Seems more like a hipster-ish take on making music. "I only make music on instruments that I can build myself."
There are also any number of ways to make cool new sounds, even using old tech. (My current favourite is Aparillo by SugarBytes, which uses FM but sounds absolutely amazing.)
This Google project seems to be a technology demonstrator made by people who seem curiously unaware of the domain they're trying to work in.
There is a lot of competition in this space and this project is doing nothing remotely new, including the ML element.
All of which might be forgivable if the sound was unbelievably awesome... but it isn't.
Seriously though, this announcement from Google does reek of "invented in ignorance"-style NIH'erism .. but then again, it may be an effective marketing ploy, given that there are so, so many kids getting into synthesisers, who don't have a foggy clue what decades of technology has produced, in terms of the state of the art.
Its a common thing in synthesisers though - someone comes up with some 'amazing new technique nobody has ever thought of before', and ends up .. producing sounds just like the 303's and DX's before them ..
I wish people would stop using this tired and bland description of people's motives to enjoy fresh, new, creative and diy. It seems pretty intellectually shallow to me.
So does spending a crazy amount of time and tech to implement something that could be done in other existing synthesizers with a fraction of the time/effort, just so that you can say "I made this with ML."
Spoiler alert, it's the buzz around ML.
The device is a nice toy, but even they say it's unpredictable. Usually you need about 4 samples per octave and 4 samples for different velocities (note loudness) for professional use. Would NSynth generate all those samples in uniform way so they can work together as a single instrument?
Me: "So .. how does it work?"
Axel: "It doesn't matter how it works, musicians don't care, they just want to hear something when they do something.."
Me: "Sure, okay .. but how does it work if I am a musician who does want to know.."
Axel: "Real musicians don't want to know."
Me: ...
I'm honestly not convinced he knew how it works, either! :P
This is a bland marketing demo of course they're not going to get to the really interesting parts. Oh, this combo of instruments didn't really sound that good...yeah, that's because they chose an arbitrary combination to demonstrate that they could, not because they're claiming that every new sound will be amazing. An important aspect of any new medium is the ability to make novel mistakes.
Music synthesis has been a thing for a long time now, so of course there are going to be sounds that you can get through some other technique. That doesn't mean there can be no surprises using a similar (but still novel) approach.
Andrew Huang did a video a while back about nsynth (just playing with the algorithm, before the nsynth super) and lo, despite nsynth's absolute dearth of musical merit, he actually manages to make something cool.
Edit: link to that video:
Isn't the point of marketing to emphasize the interesting parts? If the demo isn't compelling, who will take the time to make something that is?
As with anything, there's prior art, so everyone should _chill_ :)
Being passionate about synthesis, I have to mention vector synthesis which has a similar vibe in terms of interaction. More like simple mixing between sounds but nonetheless really neat and powerful.
https://en.wikipedia.org/wiki/Vector_synthesis Yamaha TG-33 https://www.youtube.com/watch?v=8DK7K5sFqWg SCI Prophet VS https://www.youtube.com/watch?v=1lJL3blZKVM Korg Wavestation https://www.youtube.com/watch?v=i1fokDelaxM
Didn't hear one single interesting sound out of that box that can be used to create music. Most producers and musicians already went back to analog because it sounds so much better.
I would be really impressed if some machine learning algorithm can only generate a kick that sounds better than the original analog TR808/909.
Perhaps I'm naive or 'old school' but I'm struggling to believe we've exhausted all possibilities of what can be done with just 12 notes (and their octaves) and rhythm subdivided into 32nds or 64ths.
In other words I'm more interested in how ML can help us discover new combinations of notes (melodically and harmonically) and rhythms rather than new kinds of sounds.
For rhythm there's a wide range for experimentation: from really slow tempos [2] to really fast and (seemingly) chaotic [3]. I like this study from Conlon Nancarrow [4], where the tempo between two voices evolve in time in coordination with each other.
So this 'different' music exists, and has existed for a long time in human history, it's just not very popular. E.G. In [1], apart from Wendy Carlos's invented microtonal scales there are some songs that take from other cultures [5] (See how the 7th track is similar to [3] interestingly enough).
And it makes sense for not being popular actually, given how human ears perceive chords as really fast polyrhythms. The simpler the more harmonic they sound. [6]
0: https://www.youtube.com/watch?v=JxWUgqduqYM
1: https://soundcloud.com/roberto-la-forgia/sets/beauty-in-the-...
2: https://www.youtube.com/watch?v=wEiRKpflgQA
3: https://www.youtube.com/watch?v=Zg1sgrw1PGM (1:50 onwards)
4: https://www.youtube.com/watch?v=f2gVhBxwRqg
You could view it as supervised learning, but with the user having only a passive (listening) role.
As a long-time player and designer of digital music instruments in academic laboratory and underground contexts, i can say it's almost impossible to produce a hit directly from a laboratory. It has to go through a commercial layer first. Then an artistic layer.
That's not a dig at anyone or anything. Music is a projection of life's experiences into sound... A lab can inject some tech into the music world, but we won't know for approximately 1-2 decades whether any instrument or technique is a hit or it.
FM Synthesis, formalized by Chowning at Stanford ,comes to mind. His initial implementations were idiosyncratic. His math was an organization of community knowledge. It took Yamaha's need to make sound cards and keyboards to direct all this effort those sounds into the chips and instruments we ended up buying and loving in the 80s. And it took artists and producers to bring out the best of the techniques. Also, you'll find when the tech passes through the artistic layer, there are many more mundane variables than the technology, such as re-sale value and ability to get them re-paired.
For example, if you think these sounds are non-compelling, I dare you to look up compositions by Chowning himself! [1] They are nothing like Samantha Fox's "Touch Me," [2] a famous example of the Yamaha DX-7 FM synthesizer.
The point of research like this is not really to inspire music fans so much as to inspire the next generation of commercial instrument designers and eventually artists and producers.
But, in every practical sense, the sounds are seemingly random to the listener.
Abstract expressionist art can be enjoyed regardless of the fact it's visually just a bunch of random scribbles.
But the sound this machine produces has to then be sculpted - lest we listen to a droning tone.
Would be cooler if you could describe the timbre of a sound you want to hear, and it produces that...
"I want a woody.. tinny.. percussive sound" (out pops some kind of glockenspiel marimba hybrid) or "I want a breathy sounding noise that sounds synthesized with the tone of tenor vocalist but a punchy distorted entry".
source: https://github.com/googlecreativelab/open-nsynth-super/tree/...
I got the feeling that since like 1920 until now all sounds are generated and we don't experience any 'new' sounds anymore.
Maybe AI could be used to explore sounds we really never heard.
To some degree, it's about how you group things. If you believe a guitar and a banjo are similar enough to qualify as the "same" sound – then no, there's probably not too much left to discover. But if small subtleties in timbre are fair game, then there's basically a limitless supply of new sounds.