Generating MIDI Music with GPT-2
gwern.net
gwern.net
Music has many rules, not only theoretical but unwritten rules about what works. You must incorporate them somehow into the program, either by code or offering something to the program to deduce them.
At the moment, it probably would be more practical to train a model to predict ratings and use that to screen generated samples or possible completions and throw out too-low-scoring ones (the 'ranker' approach worked out very well for the Meena chatbot recently: https://arxiv.org/abs/2001.09977 )
Music falls somewhere in between text (as a sequence of chords or PCM samples) and image (as a piano roll or a spectrogram), so maybe some hybrid of image and text generators is needed.
If you wanted to improve my ABC-MIDI GPT-2, the most straightforward ways would be to do data cleaning (I'm sure there's tens of thousands of awful MIDI files which should be removed! data cleaning with RNNs or GPT-2 or GANs always makes a large difference) and increase the model size (the fact that loss bottomed at 0.20, which is still quite bad, suggests that MIDI is hard enough that GPT-2 is struggling). More interesting would be to use Reformer or another long-range Transformer and try to operate directly on a more raw representation, like the the piano roll representation of MIDI. I think GPT-2 makes a lot of syntax errors which cripple outputs when a 'voice' goes silent, and a piano roll representation would be a lot more robust (at the cost of being like 10x larger).
Personally, I'd rather try redoing BPE encoding for ABC-MIDI specifically.
It doesn't, incidentally. It's just a standard StyleGAN dumping random images, trained to to model the average image/distribution, and not optimizing for human ratings or anything. Almost all the 'X Does Not Exist' things operate that way, including my own https://www.thiswaifudoesnotexist.net/
If you can use a piece as inspiration, browsing around until you find the 5% or whatever % that you like, there's really some benefit to be found. I felt a bit like Simon Cowell, auditioning works and being really picky, last time I did this. But in the end you can discard an item, go with it as-is, adapt it somehow, or use it as reference material. Eventually you build a gallery.
Do people ever train at different hierarchical levels? I've done very little ml, but it seems to me that it'd be beneficial to train a net on "plans" and then separately train one to interpret plans.
It's definitely interesting to me. Not sure I see a product in there... but if it could be made more efficient and set up to generate smaller, tighter clips with better instrumentation (after being a trained a bit more on what people like), and had a few key features like time stretching, it could prove useful to creative types. Writing a melody or progression doesn't always come easily, and a little push can do wonders for writer's block.
There is this project https://github.com/CorentinJ/Real-Time-Voice-Cloning
But I found it pretty hard to run, and it doesn't have voice input.
Depends on how much fidelity you want and how much lag you are willing to accept. Our current voice style transfer state of the art is sufficiently capable though the results may still need anywhere from 6 months to two years of development to be considered "production ready" i.e think poor quality, noise, and artefacts in output audio with existing tech. Pasini. has a pretty good blog post and paper on this:
https://towardsdatascience.com/voice-translation-and-audio-s...
Here's a random snippet of a real tune where you can clearly hear the AABB structure: https://www.jefftk.com/contras/tunes/cast64__starabovethegar...
It would be nice if the network learned these patterns (and transitions) on its own, but it's far more important to generate quality phrases. That is the hard part and as you can see from the samples we are not there yet.
I do disagree, though, that phrase repetition is as simple as you say. A good 16-count phrase has patterns within it, and a good AABB tune has relationships between the A and B parts.
2. The "too dangerous to release" refers to the text generation mode of GPT-2 and was mainly targeted at spammers etc.
3. Of course there was a hype component in it. They wanted people to start thinking about the question before it's too late, which they did. I doubt that we'll get any kind of attempted world control takeover by AI (unless someone explicitly tells AI to do it, at which point it's a guy trying to seize world control with the help of an AI tool), but even if the chance is very small, if you multiply it with the damage, you should at least start wondering about countermeasures.
To get philosophical I won’t say that I “know” it’s happening. Most of the world consists of things I don’t know about. The best I can do is build a mental model based on what I do know.
Why aren't bad actors using it in the wild? I think it's a combination of them being technically unsophisticated and conservative, propaganda not actually working nearly as well as people like to think it does, and lack of detection of competent actors (things like StyleGAN being used for fake FB profiles are detected through carelessness like leaving the faces exactly aligned).
One trained model is not enough to judge transformers. Listed to this:
Algorithmic performance data is a thing, but generally not very musical. There've been a couple of exceptions.
Why can't I just as easily say that
> There's no such thing as 'sheet music'. Sheets are a communications protocol to transfer encoded performance data between human sources and sinks.