Music and Machine Learning
ai.sensilab.monash.edu
ai.sensilab.monash.edu
As someone who's worked a lot on procedural music, I think this is definitely true. I'm always surprised to see ML-based approaches where someone has just trained a system on a bunch of songs, and then hopes the system will produce music with a recognizable structure - even though all the training songs will have had (in general) different chord progressions, different numbers of voices or melodic lines at any given time, etc. Such approaches strike me as akin to training a system on a bunch of short stories and then hoping it will produce a new story with a recognizable plot.
It seems like it would make a lot more sense to remove these hidden dimensionalities, e.g. by annotating the source data with chord or structural information, or by training on lots of different melodies that all share the same chord progression, etc. But it's hard to imagine that with enough layers the network will eventually grok all these hidden details.
Edit: As a matter of fact, you don't even need a whole DAW, really. You just need to be able to read existing DAW files and give users a reason to upload them.
I think the deeper problem is that musical structure that's obvious to the listener (chord progressions, modulations etc.) are realistically going to vary too much across the training data for any ML approach to figure out.
Much to digest here, but I especially like this thought: "There are no words in music." A whole lot of people think that music is nothing more than a carrier they have to modulate with some important message.
> The only way to recognise which pop song (from one popular training set) this score comes from is the lyrics. The melodic line has little similarity to the original recording.
...strikes me as plain nonsense. I can play that line on my guitar or keyboard and know immediately which tune it is even without the lyrics.
Am I misunderstanding something?
That is to say, I think you don't give people good enough credit if you think a bad transcription will preclude many from being able to recognize something.
Similarly, a novice butchering a song on guitar is likely completely wrong, but people will still recognize it if the general shape of the song is preserved. Heck, most novice guitars probably are not tuned properly.
For instance, take
http://i17.tinypic.com/4uneoft.jpg
Remove the title, add a mistake or two in the score, and give it to some jazz fan who can read music, he'll recognize the tune instantly, yet nobody ever played it this way, so blandly.
You'll almost never hear any seasoned jazz players refer to a chart like this as a "score", unless they have a very very classical background, or are sort of making a joke about the quality of a particular lead sheet. Scores are for orchestras and films and things like that. It's strange to see people keep referring to these sh*tty transcriptions as 'scores'.
It's like calling Kraft Mac 'n' cheese "pasta". Like, ok maybe technically it could qualify as a pasta, but you don't really refer to it as that in practice.
I think almost anyone who's spent significant time jamming/improvising with other musicians will disagree.
The statement is an oversimplification. It is an oversimplification that may be necessary for music to become tractable to machine learning algorithms, but it does more to highlight the limitations of current understanding of AI than anything else.
Otherwise a very interesting read!
Won't most people will be able to tell you with good resolution why they don't like the music they don't like?
I briefly experimented with procedural music generation many years ago and will relate my experience in the hope that some may find it interesting or take inspiration from it.
I had read the Byte magazine article called "A Travesty Generator for Micros." which works with text files and realized that Markov chains could be applied either to whole word or individual letters. Sufficiently long chains of letters almost always produce actual words. Sufficiently long chains of words generally produce complete (although nonsensical) sentences. Excessively long chains copy the input to the output. See [1] and [2]
At the time I was playing LOTRO [3][4] which uses ABC files [5] which are a text representation of music. I used the .abc files as input to the travesty program and got very interesting output. I used the rescan method which reads the input file for each note to output. It is slow but uses far less memory than the array method which reads the input once and generates a complete table of all transitions.
Running travesty on a single .abc file produces an output which is very similar to the input and only mildly interesting. Chaining together 2 or more input files is when it gets more interesting. It did not work well unless the input files had the same key signature.
I considered the possibility of transposing all input files to a common key signature but did not implement it. Nearly all music representation is an abstraction of the music. Music is generally quantized into notes of the even tempered 12 note scale. The tune is recognizable regardless of the instrument it is played on. I wondered whether there were further abstractions which could be used similar to the way that either letters or words could be used for text but am not sufficiently musical that I could discover them.
If you try this I think you will quickly get results which encourage you to continue.
[1] https://en.wikipedia.org/wiki/Parody_generator
[2] http://runme.org/project/+travesty/
[3] https://en.wikipedia.org/wiki/The_Lord_of_the_Rings_Online