My approach to automatic musical composition
flujoo.github.io
flujoo.github.io
The worked example also doesn't follow the stylistic rules that would make it satisfyingly authentic as a variation in the style of Beethoven. To calibrate: if you were to submit it as an exercise for music A-Level in the UK (age 16 pre-university) I don't think you would get a passing mark unless your teacher was feeling particularly generous.
But on the other hand, Classical style is incredibly refined and specific. They could have had much more success producing something following the example of Messiaen, who wrote a specific set of harmonic and rhythmic rules he was going to use for all his compositions which would be relatively easy to encode in a program. ("The technique of my musical language" is the 2-volume book I'm talking about and it's completely amazing btw. He really was an incredibly extraordinary person).
https://www.scribd.com/document/355450046/Messiaen-Olivier-T...
So while it's true that classical style(s) are very specific, it's also true they're all different. Bach, Mozart, Beethoven, and Chopin are not speaking the same language. They don't even speak the same language across their careers.
So while there are common techniques - like the idea of elaboration - they're really just the most obvious surface features, and implementing them is not nearly enough to produce decent pastiche.
AI/Music research is littered with the the corpses of projects that have attempted to do this. It's an incredibly hard problem - because unlike a board game, the "rules" change all the time, and there is no clear definition of "winning."
You can certainly mass-produce musical content in a fairly simple style - say minimal techno - and then cherry pick the best output. But even that turns out to be harder than it seems, and it's not real AI unless you can formally define why some output is better than other output. Beyond some fairly minimal perceptual and psychoacoustic basics, there's almost no research into that.
Really, we know much less about music than we think we do. Theory, Schenker included, is a very basic introduction. The interesting bits happen elsewhere, and we really don't understand how they work.
Reverse engineering music is hard, music has patterns but it also has patterns of breaking patterns. I think the heuristics here for compostable repetitive elements that repeat, reduce, and elaborate is a neat approach-very fractal-like in a way.
I'm looking forward to reading the other articles submitted by you!
https://psmag.com/social-justice/triumph-of-the-cyborg-compo...
This ch0pi1n python library looks supremely interesting, as when we listen to music, we're really expressing structures and shapes in consonance and dissonance with each other where each elaborates facets of the others. These are if not functions, at least algorithms composed over types. The author's description of these is just the right level for deriving and applying a logical architecture without diving into some bonkers numerological gematria. The post is a beautiful way of thinking about these forms. I look forward to revisiting it and playing with the library.
What I would really like to see is an attempt to write multipart contrapuntal works like Bach with some kind of AI. The rules are fairly well understood, but Bach knew how to adapt and even violate them all but still wind up with amazing pleasant music.
I think it would be cool to combine the two. Instead of generating raw midi, your GAN or reinforcement learning agent or whatever could try to generate sequences of transformations to melodic fragments. Neural program synthesis type stuff.
Or maybe one could build an automatic music analysis tool that can start from the score and try to infer the program that generated them. (Is that a thing already?)
This project is quite similar to something I've been working on, but I realised early on that you can't split up features like this and get credible results - credible meaning "appropriate for the style grammar."
It's a bit - only a bit, but let's go with it - like trying to generate sentences by swapping out nouns and adverbs. You end up with something that is grammatically correct in theory but makes no sense in practice.
Classical music particularly is fundamentally integrated in a way that textbook analysis doesn't fully explore.
As for deep learning models which create good contrapuntal music, see e.g. 'Biaxial RNN' https://github.com/danieldjohnson/biaxial-rnn-music-composit... by Daniel D. Johnson, who is now at Google Brain but wrote this as an independent(!) researcher. (Note that the existing code requires Python 2.x It would be interesting to forward-port it so it can work with Python 3.x and a maintained version of Theano. Replicating the model using Tensorflow would also be quite worthwhile.)
If you're interested in Bach's work specifically, the "BachBot" and "DeepBach" projects are also interesting but less accessible.
Example output for all of these models can be found on the Internet, just look around for it. The proprietary system AIVA is also worth mentioning because even though it's so proprietary and secretive, the compelling and "serendipitous" music it manages to come up with is a tell-tale sign that it's actually doing well-founded deep learning stuff behind the scenes, much like the aforementioned open systems. Note that much of the released output has been orchestrated (AFAICT) manually by humans, but at some point I was able to find some piano-format reductions that are most likely very close to what the AI actually created, somewhere on the official site.
Any prior information about the domain can be encoded in a way you prepare the input samples, and this would be equally useful and relevant regardless of the DL architecture you decide to use (RNNs, autoencoders, GANs, transformers).
Using raw audio samples has two advantages: First, you provide all available training data to a model, and trust that it will learn to extract what it needs from it. Second, unlike any other music representation formats, you have huge amounts of training data. So essentially all you need is a huge transformer and many thousands of GPUs to train it on, using all mp3s you can get your hands on, and you can reasonably expect a similar level of quality as the quality of text generated by GPT-3. What do you think?
Adding features to provide further information about the input data can be helpful, but not nearly as much as ensuring that strong priors are reflected in the actual model architecture, whenever relevant to the domain. Johnson's paper about biaxial RNN provides one example of how to do this in the context of music.
Forked it and spent the day upgrade to Aesara (the maintained Theano fork) and python3.
Haven't fully trained/tested the model, but feel free to take a look: https://github.com/kpister/biaxial-rnn-music-composition
Also, the old version could use a special incantation:
theano.config.experimental.unpickle_gpu_on_cpu = True
to run CPU-based inference on a GPU-trained model, even without a supported hardware GPU. It would be nice to check how this stuff works with the newer packages, and document it properly. Most likely nothing much has changed.Here is another very good one: https://openai.com/blog/musenet/
I'm a trained musician myself and interested in automatic music composition following the progress for the last thirty years, but only recent work (like the ones referenced) produce convincing results (besides Cope's work of course, but which required manual selection and editing).
You might also be interested in this survey paper: https://arxiv.org/abs/1709.01620
If so, the most important reason that it sounds terrible is the music is generated with MuseScore, without adjusting dynamics, tempos, etc. Actually, the Beethoven's original sonata in this blog is also generated with MuseScore, and it sounds not so good even with dynamics added.
However, this is why I agree that deep leaning is more promising than this manual approach, since too many variables you need to adjust to make music sound good rather than syntactically correct.
And I think that is not a very nice thing to say about other people’s work.
What's cool about this library isn't the quality of the (midi, retch) output -- what's cool is the actual Python library behind it, and the writeup, which are both super easy to read and follow.
Art needs intent.
Actually, art IS intent. (Contemporary art is pure intent: art without artifacts.)
With music, what works (for me) is a mix of algorithmic generation and then selection / arrangement.
Here's a piece I made following that approach: automatic generation of ideas, then me doing the selection, ordering and interpretation.
The music of earlier composers, Bach especially, may be more robust when put under this type of algorithmic manipulation since much less sense is lost in Bach even if you only have the melody.
Here are my responses to some comments on Hacker News.
Limitations
To be completely automatic, ch0p1n should be able to
1. analyze 2. generate 3. manipulate 4. select
musical materials to generate music.
Specifically, it should be able to analyze musical structures, generate core musical materials, manipulate these materials to produce more, and select musically good materials to make semantically rather than syntactically correct music.
For now, ch0p1n can at best provide only a framework for manipulating given materials to generate music.
Deep Learning
I have only very general knowledge of deep learning, but I think it is more promising than any other manual approach in automatic composition. I will spend time study it.
Terrible Result?
Some comments say the final generated music sounds terrible, which I can only partially agree.
All music pieces in this blog are generated with MuseScore. To make the music sound less mechanical or less terrible, you need to carefully adjust dynamics, tempos, pedals, etc. And even so, the music may still sound unsatisfactory. The original Beethoven’s sonata generated with MuseScore sounds bad even with dynamics and articulations added.
However, ch0p1n can only deal with pitch and durational aspects of music for now, but to make music sound good, you need to adjust a lot of variables. This is also why I said deep learning is more promising.
Some comments think the terribleness results from that ch0p1n can only generate syntactically correct music which may be musically meaningless. While I agree that ch0p1n has this limitation, I do believe in most cases, syntactically correct music is good enough to sound good.
I will generate more convincing music with ch0p1n in the future.
Would this be an early version of digital upload?
You're not replicating Mozart's mind. You're replicating your interpretation of some painters' interpretations of what he looked like; a selection of word sequences bearing some statistical similarity to what he wrote down; a fantasy of your own devising as to what his voice sounded like, and you don't have rules for producing Mozart's music -- as this article demonstrates, pretending that you can produce satisfying music from a set of rules is not convincing, in the same way that GPT-3 output cannot convincingly pass a Turing Test as soon as you ask questions about meaning.