The tone on one of these tools is hypocritical. When it comes to digital art general sentiment is that it's inevitable and artists need to up their game. This sentiment is not being repeated for code generators.
The tone on one of these tools is hypocritical. When it comes to digital art general sentiment is that it's inevitable and artists need to up their game. This sentiment is not being repeated for code generators.
Learning from source code is like learning from Photoshop files or Logic projects. Artists usually do not share them. In music, not even the labels get anything than the final mixdown (the music equivalent of the binary); the stems or projects remain with the studio - just like the source code in software projects tends to remain with the studios.
In software, we started showing and sharing source code under the assumption that it makes other humans making software better. People still rarely do this in music, as holding up industry secrets is deemed more important than fostering a community of personal growth.
We don't open source our code so an AI can grow. We do it so that humans can grow. There's no need to publish our code to distribute our final binary. This breaks the common understanding of why it's a good idea to share your source code, and as a result, people might become more protective of their code again.
In many ways, your comment is similar to all the comments saying "but you open-sourced it, surely you must know that people will then put it in their commercial code?"
Also the code is the end product. Similarly, artists don't share their project file because it’s not the point. This is like saying programmers not sharing their vim macros.
Apparently it's directly available in Logic (or was)
https://old.reddit.com/r/Logic_Studio/comments/gic2r2/logic_...
Hold on just a second here. How do you know what "people" "want"? Who put you in charge of speaking for them?
Yes, there have been some loud complaints, but given that millions of people are involved, the overwhelming majority of whom haven't expressed an opinion one way or the other, I think it's a bit premature for this kind of blanket statement.
Secondly, there's a difference between what people want and what people are legally entitled to. If you ask people if they want a million dollars, most of them will say "Sure!". That doesn't mean they're going to get it.
Existing copyright controls the making of copies. That's it. There's a fudge factor in there called "fair use" that controls whether or not something constitutes an infringing copy.
Whether AI training data falls within that or not is going to have to be decided in the courts or by some type of government action. It's clearly not an exact copy...the actual pixels in the original work aren't anywhere in the database. But is what is in there close enough to be considered a "derivative work"? I don't know, and neither do you.
Again, it's way too soon to be making blanket statements on the issue.
While it may differ in other countries, in the United States the purpose of intellectual property laws, as expressed in the Constitution, is to "To promote the Progress of Science and useful Arts by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries".
Enriching the copyright holders (on the rare occasions that actually occurs) is a secondary consequence, not the prime purpose.
Does AI "subvert" "promoting the Progress of Science and useful Arts"? I don't think so. Quite the contrary... I think it advances the progress of science and the useful arts, if anything.
It's pretty well established that a description of a copyrighted work is not protected, even an extremely-detailed description (see, for example, the way that Phoenix assigned one team to write an extremely-detailed description of the IBM PC BIOS, then gave that description to a second clean-room team that hadn't seen any of the actual source code. The second team then produced a clone of the BIOS that could be sold without paying IBM anything).
The data stored in these models seems more like a "description" rather than a "copy" to me -- though, of course, there's no guessing what a court or legislature will decide.
It does subvert that purpose to the extent that it makes some people no longer willing to share their works. The entire purpose of copyright is to encourage the sharing of works.
Absolutely, it may. That was my point that this is exactly to be observed and decided.
1. To showcase their work
2. To get feedback
Also, I don’t think it’s fair that we have to wait for the majority response before we can form an opinion of whether or not something is good for them. Do we have to ask millions of people if they like it if their health benefits, social security or left thumb gets removed before we can say that they definitely won’t like it?
I think it is pretty safe to say that artists don’t like having their whole livelihood overnight. (Inb4 get better jerbs)
this is a general problem in open source. And with obfusication in an AI it is just unfair. If AI / codepilot would reference the source in a nice way, then I woundn't mind.
This is absolutely not the reason and it's surprising to read this opinion stated as a fact.
First of all, not revealing multi-tracks doesn't prevent personal growth. Human ear is pretty good at discerning individual elements of a mix. It's all in the final product. If you can't learn from it or can't hear something, either it's an unimportant part of the mix, your ear is not good enough yet to distill any meaningful information from the balances of the mix in question, or it's just not a good mix.
Secondly, multi-tracks are not the source code. Putting it in relevant terms, it's just the assets. What's important are processes and artistic decisions made while processing them and combining them in their final form.
Not revealing stems/multi-tracks is not a matter of gatekeeping, it's a safety measure to prevent unauthorized use of individual composition elements and creation of "alternative" mixes which break artistic integrity of the piece.
Music is art (at least a considerable portion of it), and and artist needs as much control over the final product as possible. Mixing is a part of it.
If musicians' priorities were for the world to share in their musical indulgence not just in listening but also in like-minded creation, they would release as much of their mixing tools/backgrounds/stems/etc as possible. Since the capitalist framework is a foregone conclusion, they need to keep some secrets in their sauce to produce scarcity to justify a price for their labor so they can eat... Because the farmers and butchers all keep gates of their own kind, too :)
They do. For a price. When you know how to ask.
The thing is, if you really want to « like-minded create », you do not need, not even want their stems (you can find most of them anyway online, if only for learning).
If your thing is not « consumption » but creation, your own voice/process/way matters more (and is more fun) than starting from the studio separate tracks.
As for the music _business_ side of things, attention (hence scarcity management) on and availability/capacity to provide are the two main levers you have if you want to live from it. As with any other business actually.
I disputed a specific mistaken proposition, that "holding up industry secrets is deemed more important than fostering a community of personal growth". It's a mistake on several levels: there's no dichotomy as it's presented in this sentence, and the reasons for not providing multi-tracks and project files are different.
> If musicians' priorities were for the world to share in their musical indulgence not just in listening but also in like-minded creation
So? They have other priorities so this is not what they do. Teaching is an industry that's only an adjacent and derivative activity to the cultural sphere of arts. It's of no concern for the artist, unless they choose to capitalize on their skills/knowledge/experience.
I think you’re being a little too altruistic. When I publish code, it’s not about helping people grow. It’s so that others can see the code they are using and that if there is a bug or mistake, they can correct it. Hopefully they share that fix back, but it’s not required.
Only if Copilot was issuing compiled binaries. In this case the end result is the code.
When was this decided? I don't remember any licences or foundational essays saying this.
Copilot can provide very specific implementations of problems with "in the style of $NAME" prompts. Consider a case where you put your life's work as GPL licensed code, and someone can reproduce + adapt your highly optimized matrix multiplication code with a simple prompt, without license. Even if it's not a "textbook license violation", you're lifting one's GPL licensed function and landing it to your codebase without its license. If your code base is not licensed under the same GPL version (or later if the repo allows), it's both unethical and license breach at the same time. Adaptation of the code doesn't matter.
Same is true for more strict, source-available licenses. They are open source, but not open to be reused. What will happen if you put a function derived from a codebase with these strict licenses with or without knowledge? You're again in a dangerous grey area from both license and legal perspective.
The issue we discuss is neither straightforward nor simple to navigate. I left GitHub because of this, and may tag my repositories with this badge.
Open source means nothing if copyleft is taken out of the picture, and licenses are simply ignored.
This.
Since Copilot arrived I thought that most open source developers would be fine with it if github simply even tried to acknowledge the original licences in any way.
I personally would be ok with even a very indirect aggregate group based thing that should be no burden at all for github. They make a big list with everyone's name in it and call it the copilot contributors, and provide some kind of page for it, then when copilot spits out code, it includes a link to that page, and/or a user includes that in the credits/authors for their project like any other credited source.
No excuses about how impractical it would be to cite 3 other authors for every line of output.
But they don't even do that tiny bit. They don't try and fall short, they don't even try. But they still take the goods. The goods are already free, and yet they still manage to steal them.
Challenge accepted?
With the speed at which things are advancing....
I open source my code so that programmers and users can grow. I don't care if that programmer or user is a meatbag or a machine (or both! or neither!).
What I do care about is that the license terms of said code are respected. The vast majority of my code may be under something permissive like MIT or (lately) ISC, but I do make giving credit where credit's due a condition of using the code I've written, for good reason.
That's where tools like Copilot make a misstep: by ignoring the conditions I've placed on the use of my intellectual property. Plagiarism is plagiarism, regardless if it's a human or AI or dolphin or Martian or whatever doing it.
That's also where tools like Copilot differ from e.g. StableDiffusion. AI-generated art doesn't (usually) involve copying and pasting snippets of existing artwork into a new work the way Copilot has been demonstrated to do on multiple occasions.
(My other "problem" is that I can guarantee Microsoft will assert double-standards when it comes to Copilot infringing on e.g. the GPL v. Copilot infringing on Microsoft's own EULAs - and I really really really want to see that happen via someone tricking Copilot into ingesting the Windows source code and vomiting that into Copilot users' IDEs verbatim)
FTR I personally think that training commercial ML algorithms on bulk data produced by people that haven't released it to the public domain or licensed it for such use is ethically dubious (regardless of the legality) whether the data consists of code, art, metrics or anything else.
HN is actually one of the most homogeneous public communities I've seen.
And judging by how few downmodded posts are around, self censoring.
But then the next day a different post on the same topic could have an entirely opposite hivemind, often no discernible reason.
I'm not so sure this is a huge effect. I often express "incorrect opinions", and get some amount of downvoting or ridicule, but not usually to a significant degree. When I've had my comments dogpiled, I can see that it happened because of how I stated my opinion more than because of the opinion itself.
Sometimes the same story will be flag killed one day, front page the next.
Look in a thread where China is mentioned and you'll see the real face of many HN'ers - if they aren't removed before you see it. Luckily the MOD'ing is pretty good.
Unfortunately since most people aren’t experts on most things, this turns into a bunch of fugazi. People who are financially or academically successful in business or engineering or marketing think they understand politics too.
This is what happens with topics like China. I don’t care about China much. So I end up defending it because the negativity towards it by the Global North is staggeringly inconsistent.
This doesn't seem obvious to me at all. It seems not only plausible that this could be the case, but unsurprising.
Almost nobody here reads every story in the feed. Any given story on HN is going to be read by a collection of people that differs from story to story. And most of those people won't make any comment whatsoever. It seems unsurprising to me that a story on, say, capital punishment could get a large number of people who are passionate about that topic, and a story on pedophilia could get a large number of people who are passionate about that topic, and the majority of commenters in both could hold opposite opinions without even one of them being hypocritical.
I've seen this question posted on HN so many times in defence of Copilot but I have yet to encounter people who fit what you're describing: people in general raise a lot of issues with MidJourney/StableDiffusion/&c. too.
HN is particularly defensive when it comes to Copilot because Copilot deals with sourcecode which is central to the daily vocation of a large portion of HN's readership. HN focuses on Copilot because they're subject matter experts in what Copilot trains on.
Also - as mentioned in siblings - while criticisms of training material sourcing by various ML models is common, there are quite a few details that makes Copilot's sourcing that little bit worse.
> it's inevitable and artists need to up their game
If this is your answer it seems you didn't even bother to read the title (never mind the article). There are two issues with Generative AI to discuss: input and output. Your answer addresses output - i.e. fear that what is produced may compete with humans. This article, and the vast majority of HN criticism of Copilot exclusively addresses input: objection to copyrighted works being used to train models without author consent. These are two entirely separate discussions.
Is it, though? My impression is that HN opinions on possible copyright issues are generally mixed for both, and mostly dismissive about the job replacement stuff for both.
There's a miss match between producers of the content used to create the models and those using it. At the moment its mostly software developers who are using the output of Copilot, whilst its mostly non artists who are using MidJourney/StableDiffusion.
My effortless attempt: OMIJIB(Only My Information Job Is Brilliant)? O-My-Jib
I can’t handle hearing another person say “HN is a big place”. Sure. It’s also incredibly obvious how homogeneous it is and how HN essentially forces you to toe the line.
I’m a leftist in tech. HN is not the place where you can publicly behave as an assertive leftist otherwise would.
Source: over 20 years of working with their products and seeing the way they treat everyone else
There are probably not as many artists here, as there are coders. So the general opinion will be against all tools which will harm the coders ground, but more open for the benefits AI will bring for the laymen, when they must not fear harm from them.
There’s no real difference it is just wagon circling once a thing threatens somebody personally.
The tech community has been one of the big proponents of “data wants to be free” and “the entire idea of intellectual property is bs” for a long time and that’s suddenly reversed.