Making Music: When Simple Probabilities Outperform Deep Learning
towardsdatascience.com
towardsdatascience.com
It strikes me as analogous to generating text from an ML model trained on a bunch of short stories. A short snippet of the results may sound convincing, but obviously a whole story generated that way would be gibberish. I think the same sort of thing is happening in the example songs - any given chord change or three-note phrase may sound musical, but over the course of several measures it's not distinguishable from random notes.
Listening to his examples what struck me was exactly that: You could hear the outline of something, but it was killed instantly by the lack of internal consistency: What at first sounded like the outline of phrases were self similar, sure, but there was way too little outright repetition.
It also fell flat at capturing well established rules of chord progressions - I have the same impression as you: individual changes were ok, but even with my very, very superficial understanding of composition I know that certain progressions add tension and others release it, and, and pairs of chords that work well together sound just plain weird if they're not part of a bigger structure that also follows the rules.
I'm reminded of a program printed in Compute! or similar in the 80's that was meant to let anyone play music. It worked surprisingly well. The way it worked was simply to only allow note progressions that were likely to result in something that at least sounded "like music". But you still needed to understand rules for composition to go from something "like music" to something that sounded pleasant.
I feel like his example falls in that category: He's made it sound "like music", but it's something most people can beat by buying a book on improvisation that explains a few simple chord progressions and some very basic rules. Going from that to something that sounds like an actual proper piece of music is another matter.
On the other hand, Markovian Models can segment and learn grammars. (Specifically probabilistic context free grammars.) The problem with these is overfitting.
You can have markov models at various levels of granularity...
This is neither here nor there. The article already applies models to two separate levels of granularity (chord progressions, and melody notes over a given chord). But for higher order structures the hard part is decomposing source music into data points that typify its structure - if you did that then applying a markov (or other) model to the data would be straightforward.
And you need to understand what you're Markovising - which apparently OP didn't either.
DAWs were created for this point. But for a lot of artists nowadays, their whole window into the music-making world is their DAW, and as such, the roles have somewhat reversed, the DAW serving as inspiration for the work and providing the limits of it.
I remember reading a paper a few years ago where they showed that just the color skin used for your DAW had an effect on the music you produced. Humans are fairly influencable.
I'd love a link for this.
- a whole phd on exactly the subject of this debate: https://yorkspace.library.yorku.ca/xmlui/bitstream/handle/10...
- http://www.arpjournal.com/asarpwp/experiencing-musical-compo...
- https://www.researchgate.net/profile/Josh_Mycroft/publicatio...
Now the DAC is an instrument.
That's the difference.
From a classical producer's perspective, we could say that the 'new music' now is a function of production.
A slightly different view would simply be to say that the Engineer/Producer is the artist.
Someone made a comment about rock/jazz not evolving much ... well, you need people who can actually play instruments! It takes massive work, talent, skill. People who isolate themselves for long periods. I know being a DAC/artists requires work ... but it's altogether of a different kind.
I appreciate soundscape art, but not not as much as live music.
FYI though maybe Brian Eno et. al. were the first to really take advantage of the tech, I have to think of DJ Sasha's album, Airdrawndagger as kind of seminal. It was a fully a producers album.
The instruments are the samplers, synthesizers, loops, and FX.
The proportions vary - Ableton plays up the sampling, ProTools plays up the linear tape and FX - but they're all more or less the same product, which is a digital implementation of a late 1980s recording studio with project recall and much more powerful automation.
To use a DAW as an instrument the design would have to be opened up and made much more configurable and programmable. Products that do this a little exist (e.g. M4L in Ableton, which is descended from Max/MSP, which is used by Autechre) but they're always limited and/or clumsy to program and invariably not very popular.
It turns out most DAW users are very comfortable with the studio metaphor and only a tiny minority have any interest in exploring beyond it. They either find their way to modular synthesisers, which are a different kind of dead end, or Max/PD, which (IMO) are a bit of a nightmare to use (PD less than Max), or maybe one of the code based environments like SuperCollider, which are much more open but so user-hostile you need a good grounding in CS and command-line development to use them at all.
To date no one has made a code-based album that has had any mainstream success.
"What do you use to produce & perform? I use Ableton Live for both. I use almost solely Ableton’s built in synths/effects in my productions. My live setup is a sort of DJ-style set, in that I mainly mix using 2 audio channels and then have several more channels of acapellas, drum loops, risers, 1 shots, drum samples, etc."
Obviously it's not happening in 'real time' - but rather, by tuning each bar, each phrase, applying different effects etc. hence they are 'playing that instrument'.
Songwriting, melody, possibly musicianship and definitely lyrics are kind of secondary. You can say anything in a song these days it doesn't matter.
I like the way he presented his explanation, and his approach to finding a solution.
However, listening to some of the songs I would say that he needs to find a better model, or like mister Burns would say: "Smithers, continue the research!".
Is this project the result of a high school computer music assignment?
https://www.crowdai.org/challenges/ai-generated-music-challe...
Also check out the recent work from Google's Magenta team. Their MusicVAE seeks to model not just the instrument, but the expressiveness of the musician. With "style" emerging from the MIDI representation alone.
Latent Loops Beta
Just think how big a model should be to understand self similarity in 2 different parts of a song.
(500 internal server error when clicking on Generate Pop Music.)