MuseNet
openai.com
openai.com
Overall: stylistic coherency on the scale of ~15 seconds. Better than anything I've heard so far. Seems to have an attachment to pedal notes.
Mozart: I would say Mozart's distinguishing characteristic as a composer is that every measure "sounds right". Even without knowing the piece, you can usually tell when a performer has made a mistake and deviated from the score. The mozart samples sound... wrong. There are parallel 5ths everywhere.
Bach: (I heard a bach sample in the live concert) - It had roughly the right consistency in the melody, but zero counterpoint, which is Bach's defining feature. Conditioning maybe not strong enough?
Rachmaninoff: Known for lush musical textures and hauntingly beautiful melodies. The samples got the texture approximately right, although I would describe them more as murky more than lush. No melody to be heard.
Overall, IMO it wildly gyrates from the best I have ever heard all the way to the return of Microsoft's thematically lifeless Songsmith without warning...
Most previous attempts at neural net composition restricted the training set to one style of music or even one composer, which is pretty silly if you understand how neural nets work. It was obvious to me if you used a very large network, chose the right input representation, and most importantly used a complete dataset of all available music, you would get great results. That's exactly what OpenAI has done here.
It is still lacking some longer term structure in the music it generates (e.g. ABA form). But I think simply scaling up further (model and dataset size both) could fix that without any breakthroughs. This seems to be OpenAI's bread and butter now: taking existing techniques and scaling them up with a few tweaks. (To be clear, I don't mean to minimize what they've done at all. "Simply" scaling up is not so simple in reality.)
What might still need some breakthroughs is applying the same technique to raw audio instead of MIDI. Perhaps what is needed is a more expressive symbolic representation than MIDI. I'm imagining an architecture with three parts: a transcription network to produce a symbolic representation (perhaps embedding vectors instead of MIDI), something like this MuseNet for the middle, and a synthesis network to translate back to raw audio. This would be analogous to gluing together a speech recognizer, text processing network, and a speech synthesizer. Such a system could generate much more natural sounding music, even perhaps with lyrics.
Long-term structure MuseNet uses the recompute and optimized kernels of Sparse Transformer to train a 72-layer network with 24 attention heads—with full attention over a context of 4096 tokens. This long context may be one reason why it is able to remember long term structure in a piece
* Who uses piano as the lead instrument in bluegrass?
* They're only using one one note velocity for the whole piece, which misses a huge wealth of variation through rhythmic accent.
* Timing generally feels a bit robotic?
* The best parts it sounds kind of ok but boring; in bad parts it sounds like nothing (https://soundcloud.com/openai_audio/gaga-beatles is even worse that way)
* A lot of the artificiality seems to be the synth they're using.
* On the other hand it does understand a bit about phrasing and repetition, which many people new to traditional music take a long time to pick up.
I do agree they could have trained it about the importance of velocity though. (That neural net, and most young music students out there too, heh)
It might have been stripped off to ease complexity, or maybe it was unadequate and created odd sounding results. Maybe the midi library they used doesn't do a good job of portraying it and it is actually there! Who knows!
MIDI certainly can represent realistic compositions, when used with good synths, though I agree it's not what it's known for!
> I suppose the background "machine gun" piano notes would be a bassy synth combined with 90s electronic music drums.
Uh, what sort of bluegrass music do you listen to and where can I find some of it?
Another interesting musical form to see how well MuseNet does would be boogie woogie.
[1]: https://open.spotify.com/track/4QGO8m9OCsiW5iNF4ZNqgk?si=E8C...
And then we'll have Utopia!
Every gameplay a different soundtrack. Sounds fun.
http://www.audiogang.org/chance-thomas-composer-of-dota-2s-2...
Tim Larkin: We took several steps to keep the music interesting enough that the players would be inclined to keep it on as they play. We keep it changing so it won't become tedious; to this end, we created a music director that runs alongside the AI director, tracking the player's experience rather than their emotional state. We keep the music appropriate to each player's situation and highly personalized. The music engine in Left 4 Dead has a complete client-side, multi-track system per player that is completely unique to that player and can even be monitored by the spectators. Since some of the fun of Left 4 Dead is watching your friends when you're dead, we thought it was important to hear their personal soundtrack as well. This feature is unique to Left 4 Dead.
For single-player it works really well for building tension leading up to being attacked by waves of zombies and creating calm spots after high-stress encounters.
In the online versus mode however players got really good at using the musical hints as "tells" for when certain actions happen in the game. Most notably there were musical signatures consisting of a few notes that would play when the special infected characters spawn. Coordinated teams could use this to their advantage know when the infected team is about to stage an attack (like an ambush at a choke point in the map).
Besides, this is already doable, without neural networks. E.g. Karma[1]. The fact that most "AI music" enthusiasts don't know/care about such systems is a clear indicator that they aren't really interested in music or helping creators and just want to shove AI in yet another niche where it doesn't belong, then pat themselves on the back for contributing to "progress".
Further commodification of what's left of out popular culture isn't something anyone should be excited about. I just can't wait for the endless wave of trite "style blends" and the inevitable "oh, but all human music is garbage anyway" justification.
[1] https://www.karma-lab.com/
Edit:
Also, some people don't seem to be aware of it either, but there is already a full-fledged genre of generative music where artists use anything from analog circuits[2] to randomized or algorithmic sequencers/arpeggiators[3] to custom-built digital devices[4] for creating entire tracks. Of course, the point of generative music usually isn't to replace the artist, it's to shift their focus.
[2] https://www.youtube.com/watch?v=u8Hr4wBRbaY
Can you elaborate on why AI doesn't belong in this niche? The results are too good? The results are not good enough?
This isn't chess or go, where you either win or you don't. Chopin composed for a reason, and it wasn't just an excuse to throw a lot of notes at the page.
It might be possible for AI to work at that level someday, but it's not just a technical problem, and you won't be able to solve it by throwing a corpus of compositions at it.
Aside from that, this still sounds like aimless noodling. It's far more polished noodling with some awareness of genre cliches, but it's still essentially aimless - and so meaningless.
I disagree based on your following quote:
> Aside from that, this still sounds like aimless noodling. It's far more polished noodling with some awareness of genre cliches, but it's still essentially aimless - and so meaningless.
This sort of problem can easily be reframed as a win or lose problem, we simply consider whether the music sounds good or not. More concretely: would this music be able to convince you that a gifted human composer created it?
I'm fine with the answer being no, but I don't understand why the original comment I replied to wrote off the entire exercise. No we may not be there yet, but this seems like a good step forward to me.
Would a computer? If it could, then in theory it could use that as an optimization goal...
You're arguing an academic definition of meaningful: can you design an AI that solves a hard problem?
gambler (and I think TheOtherHobbes, who you are replying to) are arguing a practical definition of meaningful: does it solve a problem people actually have?
It's neat that you can make artificial music, but actually generating music, per gambler's original comment, isn't a problem people have. It also doesn't actually add too much to culture. Essentially, the results are "meaningless" in that, even if it was successful at sounding good, what value would it actually have besides novelty?
That is valuable to me.
[1] https://filmstro.com/ [2] https://news.ycombinator.com/item?id=17132462
"The Double-Edged Nature of Video Game Music" https://www.youtube.com/watch?v=HZB2_hlgKoA
Wouldn't solve all the problems, but I think there could be very interesting results if a composer provides skeletons of themes used, and "AI" adapts them to world and player state dynamically.
Like the game of thrones melody, or star wars.
Consider two tracks that are identical (forget copywrite for a minute). Between one that an AI generated and a human composed, I would personally grant the human-generated version more credit and enjoy it more. The story of how art is created and the stories of the artist are as substantial to appreciating art as a stroke of a brush or a note on a page. Computers will never replicate this until singularity.
Even accepting the premise, what happens when the next artist with a great story is simply using MuseNet to write their emotional pieces and passing it off as human? They'll be functionally the same, yet it still feels like something was lost.
I wonder, if they can compose, how long until passable lyrics are added along?
Given two stories, both identical, where one story is real - the real story will always be more meaningful because it has actually happened within the constraints of our reality, granting it validity and us the ability to relate to it.
Now, consider two stories, both identical, where one story is "real" and the other story is from a simulated universe. Now I'd say that both stories are of possibly equivalent value, since both have happened.
Human potential always intrigues us "What? Human can do THAT?" kind of way.
Yes, the machine and AI can do the same thing at the fraction of time, from the practicality standpoint, but it's not and never be the same — it's empty. It's just lifeless product and we never feel related to it.
ML/DL is coming for a lot of the grunt work. It's coming for us as programmers as well. It's probably a few years away, but ML/DL
I've been training on Stack Overflow and the model has already learned the syntaxes and common coding conventions of a bunch of different languages all on its own. Excited to see what else it's able to do as I keep experimenting.
Some sample outputs (you'll probably want to browse to some of the "Random" questions because by default it's showing "answers" right now and I haven't trained that model as long as some of the older question-generation ones): https://stackroboflow.com
To give an idea how big is the gap between MuseNet and CodeNet, we can consider a simple problem of reversing a sequence: [1,2,3,4,5] should become [5,4,3,2,1] and so on. How many samples do you need to look at to understand how to reverse an arbitrary sequence of numbers? Do you need to retrain your brain to reverse a sequence of pictures? No, because instead of memorizing the given samples, you looked at a few and built a mental model of "reversing a sequence of things". Now, the state of the art ML models can reverse sequences as long as they are using the same numbers as in the dataset, i.e. we can train them to reverse any sequence of 1..5 or 1..50 numbers, but once we add 6 to the input, the model instantly fails, no matter how complex and fancy it is. I don't even dare to add a letter to the input. Reason? 6 isn't in the samples it's learnt to interpolate. And CodeNet is supposed to generate a C++ program that would reverse any sequence, btw.
At the moment, ML is kinda stuck at this pictures interpolation stage. For AI, we don't need to interpolate samples, but need to build a "mental model" of what these samples are and as far as I know, we have no clue how to even approach this problem.
We will definitely get a great code autocompleter at the very least..
The program does not have volition.
Why would you think that using statistics to generate a model of a piece of art (which is just data in the case of MIDI and pixels) would be "off-limits"? People have been doing this for decades.
No one knows the answer to your last two questions, but there is no indication that this program is leading there.
The concept of how different human intelligence is from "AI" fascinates me, as it would seem to say a lot about the nature of intelligence and how far we are from GAI (pretty darn far).
Even several months before Alpha Go beat Lee Sedol there were people saying the AI could be good but never great. Now everyone admits it's super human.
I agree that getting really texture and subtlety in the work is really hard. However a really good system has the possibility to be better than any human has ever been, and IMO that's super exciting.
So long as the training set only tells it what human-composed compositions look like (MIDI files of existing music), a system has no signal to find "super human" territory.
On the issue of generalising alpha go learned go by watching human played games and by playing against itself and is now better than any human. So it's not true that a system is limited only to the skill of the examples it's shown, it's possible to design something which can surpass the examples it learned from.
There was a baroque-pop song just now that had a ritardando that almost gave me chills. Probably copied from Chopin, but still.
If I look at the Marvin Gaye Blurred Lines case: https://en.wikipedia.org/wiki/Blurred_Lines#Marvin_Gaye_laws... that took years to resolve involving humans, I wouldn't personally risk to release machine-generated music that was trained on copyrighted music.
The legal consensus, such as it is, seems to be that (if you did not otherwise agree to a contract/license modifying this in arbitrary ways) you create a new copyright & own it if you use their music-editing tool to tweak settings until you got something you liked, because you are exercising creative control, making choices, and engaging in labor. On the other hand, if you merely generated a random sample, neither you nor anyone else own a copyright on it.
What if that person is a monkey[1]? Is it "animal-made art"[2]?
[1] https://en.wikipedia.org/wiki/Pierre_Brassau [2] https://en.wikipedia.org/wiki/Animal-made_art
As your own links indicate, animals have no more copyrights any more than a computer program would because they are not human, and copyright is explicitly granted to human creative efforts.
Since this is OpenAI, is MuseNet open source?
Doable with a single 1080ti and a couple of hundred midi files?
Also, can you do supervised learning with this - say melody input and chords (with good voice leading) output?
A 1080ti would probably require something like several days or a week. It depends on how big the model is... Probably not a big deal. However, a few hundred MIDI files would be pushing it in terms of sample size. If you look at experiments like my GPT-2-small finetuning to make it generate poetry instead of random English text ( https://www.gwern.net/GPT-2 ), it really works best if you are into at least the megabyte range of text. Similarly with StyleGAN, if you want to retrain my anime face StyleGAN on a specific character ( https://www.gwern.net/Faces#transfer-learning ), you want at least a few hundred faces. Below that, you're going to need architectures designed specifically for transfer learning/few-shot learning, which are designed to work in the low _n_ regime. (They exist, but StyleGAN and GPT-2 are not them.)
I would kill for a VST tool that would take a set of midi tracks, and synthesize a new track for a specific instrument that "blends" with them. I would also kill for something that can take a set of "target notes" and break them up/syncopate/add rests to produce good melodies, or take a base melody and suggest changes to spice it up.
I definitely think creativity is on the radar for AI, see: AlphaGo. Everything we think is based on emotions is ultimately learnable.
To me, creativity is really about generation of "aesthetic novelty" which is hard to get from a ML algorithm that is trying to approximate patterns in training data. Eventually, there will be models trained on a wide variety of art, music, stories, etc that can recognize aesthetic and structural isomorphism between mediums (say between a grizzly picture and death metal), then we'll lose our competitive advantage. I don't think we're nearly so close to that as the singularity types would have us believe though.
Actually I think "inventing new forms of music" is a pretty great musical Turing test. How much neural-network training would it take to make an AI that can take the sum total of existing music, extrapolate the rules, and then deliberately break those rules in such a way as to make something that humans would find interesting?
But, despite this potential greatness, if there's a problem, it's that this AI only produces music...
What this AI composer really needs is an AI lyricist to write lyrics for the songs it composes!
Sort of like an AI Lerner to it's AI Loewe...
An AI Hammerstein to it's AI Rodgers...
An AI Gilbert to it's AI Sullivan...
An AI Tim Rice to it's AI Andrew Lloyd Weber...
An AI Robert Plant to it's AI Jimmy Page...
An AI Keith Richards to it's AI Mick Jagger...
An AI Paul McCartney to it's AI John Lennon...
An AI Bernie Taupin to it's AI Bernie Taupin...
An AI James Hetfield to it's AI Lars Ulrich...
An AI Wierd Al Yankovic... to it's... AI Wierd Al Yankovic... <g>
You know, an AI Assistant for this... AI Assistant... <g>
Well, an AI Assistant to write lyrics that is... An AI "Lyrcistant"... <g>
Come on, I know there's someone in the AI world who can do this! But it might be a bit challenging... the AI would not only have to write poetry, but it would have to match that poetry to all of the various characteristics of the music...
Not an easy task, to say the least!
But, for the right AI researcher... an interesting, challenging, worthy one!
(I think I hear 2001's "Daisy Bell" playing in the background...)
By the way... disclaimer: I am an AI. That is, An AI wrote this message on HN.
No, I'm kidding about that! But... how would you know? (insert ominous sounding music here) <g>
Jukedeck has significantly better AI generated music but since I have not found a description of how their model works, it is hard to compare it to this.
Couldn't one generate music and upload that to Spotify and get paid based off the number of listens?
Generative music is definitely on the come-up. If you like this, also check out https://generative.fm/ , which is from another HN member.
I've used it once or twice, but for whatever reason, nothing sounds better to me than the music I used to listen to when I was a teenager.
Tangental, but listening to the music from your teen years is a form of therapy for people who developed dementia. It's possible the music we impress in our teen years hold a special value in our brains.
Now an unpopular opinion. I'm not an ML expert, so take my words with reasonable skepticism. This fancy GPT2 model diagram can impress an uninitiated, but we are initiated, right? There is really no science there and it's still the good old numbers grinder: an input of fixed size is passed thru a big random pile of matrix multiplications and sigmoids and yields a fixed size output. We could technically replace this nice looking GTP2 model with a flat stack of matmuls and tanhs, with a ton of weights and given enough powerful GPUs (that would cost tens of millions), train that model and get the same result. It just won't make an impression of science. How are these GTP2 models designed? By somewhat random experiments with the model structure. The key here is the GPU datacenter that could quickly evaluate the model on a huge dataset. The breakthru would be achieving the same quality with very little weights.
I didn’t quite get it. How would you feed this variable sized input?
S[0..n] = the raw input, 48000 bytes per second of sound F[1][k..k+48000] -> [0..255], maps 1 second of sound to a "sound vector". F[2][k..k+96000] -> ..., same, but takes 2 seconds of sound as input
Now instead of the raw input S, we can use the sequences F[1], F[2], etc. Supposedly, F[10] would detect patterns that change every 10 seconds. It's common in soundtracks to have some background "mood" melody that changes a bit every 10-15 seconds, then a more loud and faster melody that changes every 5 seconds and so on, up to some very frequent patterns like F[0.2] that's used in drum'n'bass or electronic music in general.
This is how music is composed by people, I guess. Most of the electronic music can be decomposed into 5-6 patterns that repeat with almost mathematical precision. The artist only randomly changes params of each layer during the soundtrack, e.g. layer #3 with a period of 7 seconds slightly changes frequency for the next 20 seconds, etc.
Masterpieces have the same multilayered structure, except that those subpatterns are more complex.
You mean like an autoencoder?
Ok, assuming we have those sequences (F1, F2, F10, etc), how would you combine them to train the model?
We can combine multiple sequences in any way we want. Obviously, we can come up with some nice looking "tower of lstms" where each level of that tower processes the corresponding F[i] sequence: sequence F1 goes to level T1 which is a bunch of LSTMs; then F2 and the output of T1 go to T2 and so on. The only thing that I think matters is (1) feed all these sequences to the model and (2) have enough weights in the model. And obviously a big GPU farm to run experiments.
Why would you try to manually duplicate this process by creating F1, F2, etc?
The idea of skip connections would be like feeding T1 output to T3, in addition to T2. Again, I’m not sure what useful info F sequences would supply in this scenario.
Don't we already do this with text translation? Why not to let one model read a printed text pixel by pixel and the other model produce a translation, also pixel by pixel? Instead we choose to split printed text into small chunks (that we call words), give every chunk a "word vector" (those word2vec models) and produce text also one word at a time.
Would it help to decompose sound into subpatterns with Fourier transform?
Afaik, there is a similar technique for recognizing faces: a face picture is mapped to a "face vector". Yet this technique doesn't need the notion of "sequence of faces" to train the model. Can we use it to get "sound vectors"?
I'm not sure what would be useful "subpatterns" of sound. In language modeling, there are word based, and character based models. Given enough text, an RNN can be trained on either, and I'm not sure which approach is better. For music the closest equivalent of a word is (probably) a chord, and the closest equivalent of a character is (probably) a single note, but perhaps it should be something like a harmonic, I don't know.
Unlike faces, music is a sequence (of sounds). It's closer to video than to an image. So we need to chop it up and to encode each chunk.
Ultimately, I believe that we just need a lot of data. Given enough data, we can train a model which is large enough to learn everything it needs in the end to end fashion. Primary achievement of GPT-2 paper is training a big model on lots of data. In this work, it appears they only used a couple of available midi datasets for training, which is probably not enough. Training on all available audio recordings (either raw, or converted to symbolic format) would probably be a game changer.
I do still think that the good old https://github.com/hexahedria/biaxial-rnn-music-composition (Hexahedria's Biaxial RNN/Tied Parallel Networks, also published a paper achieving SOTA at least for that time, on a variety of midi datasets) has more interesting/compelling output musically, despite starting from a comparatively tiny dataset and using rather elegant and easily-undestandable convolution- and RNN-based techniques. Too bad that the implementation relies on Theano which is quite endangered at this point (doesn't seem to support up-to-date python 3?), but I do think it's a compelling starting point if anyone really wants to work on this domain.
Heck, proving a piece was generated would be hard.
We played a little with transformers inside a browser using tensorflow.js a few month ago, for real-time music transcription.
For those interested : Website : https://gistnoesis.github.io/ Project : https://github.com/GistNoesis/Wisteria/
My project is currently on hold, but will definitely receive update in the future. As it's kind of project for fun with no hope of monetization given that the space is already crowded with Google's Magenta, and now OpenAI is joining the dance with museNet.
On the other hand, I've been much more impacted by visual "art" already being generated by AIs. Perhaps the musical medium is ironically harder to crack since the format is simpler, rather like how it took longer for an AI to defeat an expert goban player than a chess player.
Enjoyable to the trained mind to be sure. Especially as a moment in history. But the tunes themselves lack soul. And jarring for myself as the background music I had turned off to catch the tail end of MuseNet's performance, was of a particularly feverish level of human expressiveness ;)
Bad Brains - Live at the CBGB's 1982 (Full Concert)
So the turn-around time from sci-fi to reality is... less than a week now?
I think it’s a subtle distinction, but imagine a model where we can throw in thousands of math proofs then give the model some initial assumptions and just let it run wild. I think getting a neural net to model the creative spark / the ingenuity is what has been missing.
A part of me doesn’t want to believe that it is possible, but a part of me is genuinely curious as to the consequences.
An exciting time to be alive, folks.
I would LOVE to see where that goes... Is it going to turn into 4 chord pop? Or maybe more dominant, more resolve-y? Or maybe my assumptions will be wrong and we would collectively train for more complex music?
Perfect for youtube ads that must have obligatory music complete with the drone in the background "Pond 5.... Pond 5..." and the producer can't pitch in 5 bucks for it.
From literature, music, physics,... everything unrelated to the computer itself.
That's why every IT book needs more knowledge to teach users how to really understand art and science behind everything we experience.
I'll bet talented musicians and real composers feel the same about MuseNet. The music probably makes them cringe to their core.
Then there is the general public...
They don't notice anything wrong with do-it-yourself websites, and this music sounds amazing.
> Neural network generating technical death metal, via livestream 24/7 to infinity. Trained on Archspire with modified SampleRNN. Read more about our research into eliminating humans from metal: https://arxiv.org/abs/1811.06633
> More albums https://dadabots.bandcamp.com/