Making Music and Art Through Machine Learning
blog.ycombinator.com
blog.ycombinator.com
Try expressing feelings and stop using that word that has absolutely nothing to do with machine learning.
The very very very basic Idea of machine learning is at opposite with art.
These things have nothing to do together.
With my personal experiments with generative art, I think of the machine almost as a work assistant. I have a rough idea of what I'm going for, and the ways to attempt it, and I teach that to my computer using code. But then I go away from the computer and out into the world and drink coffee, sketch in my notebook, and in the meantime my computerized assistant is back home working on draft after draft of my latest piece of art.
I come home and I see what he's done, and most of it isn't great, but occasionally there is a gem! I look at what he did for that gem and then encourage him to do more of that, and then go back out into the world. I highly disagree that machine learning is the opposite of art.
If you go to any traditional artist's studio you will often find notebooks filled with discarded sketches and ideas. Paintings with color combinations or techniques that weren't ideal. This is almost exactly the same process. And I would argue that process is exactly what art is.
You've "selected" something you liked. You are no artist and should not call your self like that. If you do, you will never inspire "simple" people.
At least if you do it would become the MacDonald of art.
Do nothing, sit and say to the bot if it fits you or not.
You're even worse than a DJ. The DJ actually works his selection and works his transitions.
Sketches have been written out of artist mind. Not by some selection that fits you or not.
Art is work.
Slowly coaxing a desired result out of a computer by painstakingly tweaking an algorithm over the course of hours or days doesn't sound like work to you? The computer wouldn't have just done it on its own if an artist didn't tell it what to do.
I think your definition of art is narrow-minded. Who are you to tell someone their work isn't art. Art is (and always has been) in the eye of the beholder.
Still, I see the work putted as an art, not only the result.
What I mean is that tools that allow re-generation without someone putting he heart into it would not result in art.
Dropping 200 pics and having a new one is no art.
If you iterate your self on how these mechanics goes on and create something with it with control and skill then that looks more like it.
But then, if you drop out an algorithm that does that, handle it to users,
would their results be art ?
Tools and way are more important than result. How you use them and how you create something original from it is the art process.
I guess it really depends on how you use this.
in the same respect, did someone not create art in photoshop because they used some fractal noise?
I've had this conversation and thought about it a lot for what I'm doing and come to this conclusion.
art can be directed, it can be accidental, it can be both (every artist loves a happy accident), it can be because of a million iterations whether by hand or computer.
it's up to you, the viewer to decide if you think the end result is good art.
What is the difference between me selecting the specific frame made by a generative algorithm and photographer picking a specific moment in time to capture a frame?
What's the difference between me using a generative algorithm and Michelangelo using assistants to paint the Sistine Chapel?
"What is the difference between me selecting the specific frame made by a generative algorithm and photographer picking a specific moment in time to capture a frame?" -> he captures life and tries to create a embodiment of something that already has a meaning for everybody. He can have multiple ways of underlining different aspects of the same thing. He has total control over How and Why. It is not only technical but it also communicates something in very different ways depending on the subject and how he took the picture.
"What's the difference between me using a generative algorithm and Michelangelo using assistants to paint the Sistine Chapel?" -> these are people not machine.
Generation is a series of unfortunate accidents in search of something that might pass for a goal.
Get interested in some artists and what they have achieved.
A humble human being would never call it self an artist.
That's the first mistake, and it's philosophical rather than technical.
I would say the tool it self is rather an piece of art. Not what it creates.
If you really want to build a tool used by producers then it would be required to be a VST in the first place.
But then I really doubt (as an amateur composer) than any artist could get anywhere with this.
How can I transmit my feelings trough a bunch of data i've no control over ?
As an amateur composer as well, I have found uses for these types of models even without plugins / with "data I have no control over" for small flashes of interesting motifs/progressions that I rework, modify, and expand on my own. Inspiration comes in many forms.
[0] https://github.com/tensorflow/magenta-demos/tree/master/ai-j...
At best it can inspire.
edit : thus the title should have been "rearranging sound with machine learning" ; it's not music nor art.
However, I think it has its place if it's part of the artistic process. Many artists put the emphasis not on the end product, but on the process itself. We can perfectly imagine an artist creating an entire collection using AI, find something wicked in it and expose it as a criticism of our times. (sounds hipster yeah… I'm not an artist.)
Right. Now, my mind is full of thoughts and emotions about ML. And I use ML to make some videos that express my thoughts. And you judge my work and say it's not art? Really I profoundly disagree.
Saying something is "not art" is normally a tough argument to make. To say it about a whole category of practice is nearly impossible.
I think the non-research value will be like the value of chess computers to chess communities. It enriches the community, helping them develop their craft, but the community is still a community of humans.
http://theconversation.com/how-computers-changed-chess-20772
I can think of a few ways machines have beaten humans before, but where human experience is still valued in a way the machine achievement isn't. Modern weapons are valued for what they do, and they can beat any human, but the human practice of martial arts is still valued despite that fact. A player piano with a roll performed by Rachmaninoff will deliver an excellent reproduction of a historic performance, but I'd still rather listen to my friend play, even if what I'm hearing comes down to objectively worse key pressing.
Music is indeed a language, so while it evolves, it still requires a network of meaning to rely upon. So doing something wild like music in 5/6 time will for the most part, bounce off people's heads and not resonate with meaning.
Entire creative industries will be automated away with this, in time.
I liked listening to it, but am disappointed by the field's disregard for existing methods of structure. "We don't use rules".... well, rules have helped mankind progress forward for hundreds of thousands of years.
Just because they don't easily transfer into your ML model doesn't mean they are useless. Maybe the model needs to learn some new tricks.
The rules based approach has (as you mention) years of history, so people are currently exploring the green-ish fields of raw learning approaches with good success on tasks where rules based approaches performed much worse or didn't work at all (cf image recognition, speech recognition), and in some areas it seems like the more you let the model learn / get rid of classical rules based approaches (with enough data), the better it does. Whether that is true for field X, not true yet for field X, or will never be true for field X depends on who you ask.
There is definitely a recent tide of models which are focused more on rule learning, function generation, and so on. The general thing I see is that rule based approaches with good approximators/probability models to guide heuristic or exact search can do crazy things - this is the story of AlphaGo at a 10k foot level. People in the ML community are just more focused on the new-ish part (learning good probability/function approximators from data) right now.
Just because rules aren't incorporated widely yet, doesn't mean they won't be in the future. I am personally very interested in this direction, and a bunch of work from Sony CSL (Pachet et. al.) has focused heavily on this idea in the past.
As an aside, whenever you hear an ML researcher say "prior", it is generally functioning as some kind of soft or occasionally hard rule - so maybe there are more rules floating around than it seems. Soft rules aka priors are generally (much!) easier for gradient descent style optimization and incorporating directly into models, so we tend to have priors rather than hard rules as seen in many other parts of computer science. Even the structure of the model itself can be seen as a prior.
Making physical art based on inspiration from machine learning processes:
https://medium.com/@rememberlenny/digital-processes-inspirin...
>Note that this isn’t a performance of an existing piece; the model is also choosing the notes to play, “composing” a performance directly. The performances generated by the model lack the overall coherence that one might expect from a piano composition; in musical jargon, it might sound like the model is “noodling”— playing without a long-term structure. However, to our ears, the local characteristics of the performance (i.e. the phrasing within a one or two second time window) are quite expressive.
It's understandable to be encouraged by progress, but to my ears, it's not really expressive other than the first thought that I had when listening to the piece:
"That piano player is having a stroke, somebody should do something to help."
I'll put it thusly: As a broke-ass musician barely making ends meet, I'll be more than happy to contribute if Alphabet pays off what's left of the mortgage on my house.
Until then, I've only got the perspective that a project like this is trying to take food off my plate, money out of my pocket, and replace me. If you want my help, Son, you got to come better at me than with Altruism.
Otherwise, the best thing I can do to help is to say "This project is a heap of shit and your time is better spent elsewhere" without tinging it with too much condescension. Your child is brain damaged. Take it to Oregon.
The model is unconditional i.e. any concept of form at all was learned directly from data, with no real structural hints to what it should learn.
Generally you get much stronger "global structure" the more prior information you provide either in the model itself or in the input/targets (such as chord constraints, etc.). This is one reason that harmonization tasks often sound notably better than pure generation - the "backbone" in harmonization was human provided, whereas in pure generation the model needs to be self-consistent.
The analogous task in language would be character level language modeling, whereas something like harmonization or conditional generation from chords would be more akin to machine translation. Character level language models may be a bit wonky looking when generating, but any structure at all was purely model discovered from a simple rule (generally, maximum likelihood), whereas in conditional models a lot of extra help is given by the conditioning variable.
This model is a pretty big jump in quality from previous efforts in the same vein of unconditional generation, and should combine well with techniques for conditioning in sequence models from all over the research community (translation, speech recognition, question answering, summarization, speech synthesis). On top of that, the model is IMO quite elegant, and much simpler to understand than many other attempts at polyphonic generation including my own.
It is worth comparing this output to some previous efforts (from Magenta and others), some with more complicated internal structure such as [0][1][2][3][4][5][6], and also some models with stronger conditioning or different input representations such as [7][8][9], the extreme case being [10] where a human (Benoît Carré) used a modeling/ML toolkit along with standard arrangement and production approaches to produce a wacky, awesome pop song.
Magenta is putting out lots of neat stuff around interactive usage of these models, with web demos and plugins to standard music tools [11]. Indeed, Doug's quotes in this article [12] sum it up for me, excerpt: “I don’t think that machines themselves just making art for art’s sake is as interesting as you might think,” he explained. “The question to ask is can machines help us make a new kind of art?”
Link to the model in question: Performance RNN, Magenta https://magenta.tensorflow.org/performance-rnn
[0] Work from Doug in his postdoc doing LSTM Blues, http://people.idsia.ch/~juergen/blues/
[1] RNN-RBM from Boulanger-Lewandowski et. al, see TF impl http://danshiebler.com/2016-08-17-musical-tensorflow-part-tw...
[2] Early Magenta models, https://magenta.tensorflow.org/2016/07/15/lookback-rnn-atten... , ex https://www.youtube.com/watch?v=nHxr9u9_4_s
[3] Daniel Johnson's Biaxial RNN http://www.hexahedria.com/2015/08/03/composing-music-with-re...
[4] Some experiments only published in my thesis to date, https://www.youtube.com/watch?v=tavHJoum--g&list=PLRMa_gJ8vx... , thesis in question https://github.com/kastnerkyle/udem_masters_thesis
[5] Polyphonic Magenta model https://www.youtube.com/watch?v=s0BVFVqEY4A
[6] RL tuning to improve generation https://magenta.tensorflow.org/2016/11/09/tuning-recurrent-n...
[7] Folk-RNN is similar but uses ABC representation (https://highnoongmt.wordpress.com/2015/05/22/lisls-stis-recu...), https://soundcloud.com/seaandsailor/sets/char-rnn-composes-i..., also see related post https://maraoz.com/2016/02/02/abc-rnn/
[8] DeepBach , http://www.flow-machines.com/deepbach-polyphonic-music-gener...
[9] Lawson He's LSTM approach, http://www.lawsonhe.com/music.html
[10] Daddy's Car https://www.youtube.com/watch?v=LSHZ_b05W7o
[11] Interactive Magenta Jam https://magenta.tensorflow.org/2016/12/16/nips-demo
[12] https://www.technologyreview.com/s/604010/google-brain-wants...