256 karma · joined August 1, 2019
Tidal's terms and conditions (https://tidal.com/terms) say that:
> “AI-Generated Content” means any audio content, inclusive of musical works and sound recordings, that is wholly or substantially generated by generative artificial intelligence, with limited or no direct human creative input beyond an initial text prompt or similar instruction. ... You acknowledge that AI detection technology may produce false positives or false negatives.
And:
> If you use TIDAL Upload, your Tracks may be scanned for the purpose of identifying whether the content is AI-Generated Content, and to label such content accordingly on the Tidal platform. You acknowledge that such scanning and labeling is performed on a best-efforts basis and that Tidal shall not be liable for any inaccuracies in AI detection or labeling. AI-Generated Content uploaded to Tidal is not eligible for monetization. If you believe your Tracks were erroneously tagged as AI-Generated, you can reach out to support@tidal.com.
> An algorave (from an algorithm and rave) is an event where people dance to music generated from algorithms, often using live coding techniques. Alex McLean of Slub and Nick Collins coined the word "algorave" in 2011, and the first event under such a name was organised in London, England. It has since become a movement, with algoraves taking place around the world.
That said, even if model training is fair use, model output can still be infringing. There would be a strong case, for example, if the end user guides the LLM to create works in a way that copies another work or mimics an author or artist's style. This case clearly isn't that. On the similarity at issue here, I haven't personally compared. I hope you're right.
The second part here is problematic, but fascinating: "I then started in an empty repository with no access to the old source tree, and explicitly instructed Claude not to base anything on LGPL/GPL-licensed code." Problem - Claude almost certainly was trained on the LGPL/GPL original code. It knows that is how to solve the problem. It's dubious whether Claude can ignore whatever imprints that original code made on its weights. If it COULD do that, that would be a pretty cool innovation in explainable AI. But AFAIK LLMs can't even reliably trace what data influenced the output for a query, see https://iftenney.github.io/projects/tda/, or even fully unlearn a piece of training data.
Is anyone working on this? I'd be very interested to discuss.
Some background - I'm a developer & IP lawyer - my undergrad thesis was "Copyright in the Digital Age" and discussed copyleft & FOSS. Been litigating in federal court since 2010 and training AI models since 2019, and am working on an AI for litigation platform. These are evolving issues in US courts.
BTW if you're on enterprise or a paid API plan, Anthropic indemnifies you if its outputs violate copyright. But if you're on free/pro/max, the terms state that YOU agree to indemnify THEM for copyright violation claims.[0]
[0] https://www.anthropic.com/legal/consumer-terms - see para. 11 ("YOU AGREE TO INDEMNIFY AND HOLD HARMLESS THE ANTHROPIC PARTIES FROM AND AGAINST ANY AND ALL LIABILITIES, CLAIMS, DAMAGES, EXPENSES (INCLUDING REASONABLE ATTORNEYS’ FEES AND COSTS), AND OTHER LOSSES ARISING OUT OF … YOUR ACCESS TO, USE OF, OR ALLEGED USE OF THE SERVICES ….")
I'm also skeptical of anything that claims to reliably detect AI writing. FWIW, I plugged the comment into Pangram Labs, which claims to be the most reliably and seems to have worked well before. It categorized the comment as 100% human written with medium confidence.
> Nor do we agree that AI training is inherently transformative because it is like human learning. To begin with, the analogy rests on a faulty premise, as fair use does not excuse all human acts done for the purpose of learning. A student could not rely on fair use to copy all the books at the library to facilitate personal education; rather, they would have to purchase or borrow a copy that was lawfully acquired, typically through a sale or license. Copyright law should not afford greater latitude for copying simply because it is done by a computer. Moreover, AI learning is different from human learning in ways that are material to the copyright analysis. Humans retain only imperfect impressions of the works they have experienced, filtered through their own unique personalities, histories, memories, and worldviews. Generative AI training involves the creation of perfect copies with the ability to analyze works nearly instantaneously. The result is a model that can create at superhuman speed and scale. In the words of Professor Robert Brauneis, “Generative model training transcends the human limitations that underlie the structure of the exclusive rights.”[0]
I disagree with the Copyright Office here, but ofc they're the Copyright Office and I could be wrong. More broadly I'm struggling with how to permit and incentivize creation of powerful generative models while not screwing creators in the process. There are startups and other efforts trying to address this through novel licensing, etc., but AFAIK there's no great solution. I'm also cautiously optimistic that there will be some decentralized and/or federated options. It's complicated indeed.
[0] https://www.copyright.gov/ai/Copyright-and-Artificial-Intell...
The headline misses the nuance that there will be a trial on Anthropic's gathering "pirated copies to create Anthropic's central library and the resulting damages." But the court got it right IMO that "the use of the [copyrighted] books at issue to train Claude and its precursors was exceedingly transformative and was a fair use..."
https://storage.courtlistener.com/recap/gov.uscourts.nysd.64...
> Sabine, this is amazing. You are, as usual, 100% right. The delayed choice quantum eraser is a prime example of over-mystification of quantum mechanics, even WITHIN the field of quantum mechanics! I (Matt) was guilty of embracing the quantum woo in that episode 5 years ago. Since then I've obsessed over this family of experiments and my thinking shifted quite a bit.
Assuming that they have a valid trademark, the issue becomes whether there is a likelihood of confusion between Perplexity and Perplexica. That is a fact-specific, multifactor test, which I’ll spare you. But there could be arguments both ways IMO
EDIT: trademark issues aside, cool project!
But let’s hope this stands! Back to my coffee…
The 1-28 pleading numbers on the side are annoying. They're specific to courts in California and a few other jurisdictions, and the rules of court require them. But many other courts don't have them, and they only help to cite specific lines within pages; eg "Complaint 5:4-9" means "Complaint at page 5, at lines 4 to 9". It's occasionally useful for court filings like this, but more useful for court/deposition transcripts of testimony to show precisely where a witness said something.
Related: I tried building an RNN to generate legal pleadings back around 2018/19 and gathered a bunch of docs like this from courts across the country as training data. Processing text with those pleading numbers was a pain, so I built a CNN to classify whether a document had pleading numbers or not, which affected downstream processing. Probably the wrong approach in a bunch of ways, but I was just learning.
Sorry to make this about AI, but it'd also be interesting whether such a non-profit makes its data fully open--i.e., for AI companies to scoop up--or has more restrictive terms that forbid AI "scooping" without a separate agreement. Lots of tricky issues, trade-offs, and interesting incentives involved. There could be alliances with orgs creating open source models, for example. If anyone is working on a nonprofit like this or just wants to chat about it, please reach out.
https://storage.courtlistener.com/recap/gov.uscourts.nysd.59...