AI training shouldn't erase authorship
justine.lol
justine.lol
This should in principle allow to determine how much each source influenced each output token of the LLM.
The problem is that this multiplies storage and compute time for tagged inference by the number of source tags, so it may be impractical to actually tag single documents or authors, but might be useful for very broad categories like "copyrighted" vs "non-copyrighted", "synthetic" vs "human generated", "photo" vs "drawing" vs "rendering", year range of publication, etc.
Alternatively: That's 100s of years of mentions of that quote it can pull from.
Describing living authors accurately is sensitive enough that I think it’s going to require human oversight via Wikipedia. Justine has a Wikipedia page, but it could use some expansion on the software side. A project to improve Wikipedia by crediting authors of notable software might be useful? And LLM’s will pick it up from there.
LLM’s pretend to be general purpose, but maybe optimizing for code autocomplete versus searching over a knowledge graph are two different things and might end up as different subsystems.
Or maybe they’re just different kinds of training data? Like, strip the author data from the code (for autocomplete) but then use it to generate author pages to train on.
Who exactly is the "you" in that claim?
Justine clearly does want their author name (accurately) included.
I'd argue that anyone who's used a license with attribution requirements has explicitly stated they want that.
OpenAI et al. don't want that, for at least two reasons I can think of, 1) because their "artificial intelligence" isn't capable of doing it accurately or without "hallucinating" incorrect authorship attributions, and 2) because of the tsunami of copyright lawsuits that would immediately drown them.
We have a collective tendency to forget that we're the authors of legality and tend to reverse our causal arrows. Reality should be reflected in the law, not be shaped by the law.
copyright looks to apply free market competition for authored works while protecting against those who wish to freeload. its even better than that, because freeloaders have the vast array of free stuff available and nobody would complain that you personally benefit from their use even if you don't contribute back.
but I do agree with you halfway, property rights are a social construct, but I disagree with the "invalid" part.
Technically it could be possible to have author metadata either prefixing or postfixing content, that way the model learns to generate from an author style or to predict the (likely) attribution of a snippet. Or they could create a BERT embedding model for author metadata and another for content, and train them with the CLIP method. So the model would learn to map text to author embeddings.
There are tons of data with author and date - books, journals, papers, media articles, blog posts, social network posts, github repos - they all provide a way to index ideas to their authors and in time, so we could potentially find who first invented an idea and who expanded on it later.
The negative effects could include disclosing authors of anonymous texts online, or disclosing the sources of influence of one author even when they keep them secret, or don't even realize themselves. Could be embarrassing to have your sources outed like that. You might find out your original ideas were invented elsewhere.
> .. if I ask something like Claude, "what sort of code has Justine Tunney wrote?" it hasn't got the faintest idea. Instead it thinks I'm a political activist, since it feels no guilt remembering that I attended a protest on Wall Street 13 years ago.
Is this because the code is so mashed together that it's impossible to say which bit was copied from which original source? Well, then, that's a very big problem and if a human did this we'd rightly label it plagiarism.