We already tried this with human communication and gave birth to the dystopian nightmare that is social media, why keep repeating our mistakes?
We already tried this with human communication and gave birth to the dystopian nightmare that is social media, why keep repeating our mistakes?
And personally, I think book recommendations are an absolutely underserved market, if I liked a book, having the ability to find more like it would be an absolute godsend for connecting authors with people who would be interested in their works, resulting in much more potential sales for them.
I can't count how many times have I discovered an absolutely great book on Amazon with like 50 reviews accidentally, as well as other, objectively less recommendable books that have nevertheless made an impression on me.
Discovering these books is sort of a hobby of mine, and is the exact kind of activity an LLM would be a great help with.
Going further, if there was an LLM that could be asked for book recommendations for your particular tastes, it could also identify markets for books not yet written, and would give a hint to authors on what sort of books to write to find an audience.
I discovered absolutely great books by moving slowly along the shelves of a library or a bookshop.
And you need to read bad books to understand the great ones.
As for your point about serendipity, torginus never said that he didn't wander book stores and libraries looking for books he wouldn't have been previously exposed to.
Based on the post, I'm sure they understood the basics of reading a variety of books, both good and bad -- there is no need to get judgemental.
I haven't read about the industry in years but isn't it the case that the job of "book recommendations" is essentially the publishers job? They unironically try to sell you more than a book. An algorithm would threaten their worth.
(There are, of course, other useful functions like publishing and the irreplaceable editors, but neither require the capital strength of marketing.)
This is manufactured, stretched, overhyped objections. I believe it's all as the OP suggests, because the word AI is in there. Not because anything illegal or immoral is going on. In fact it's a terribly useful tool, and once the mob cools off it'll likely return.
“Our algorithms are pretty much the same as human art criticism, so put down the pitchforks you unenlightened scum” is up there with telling them to eat (a Stable Diffusion generated picture of) cake.
It's actually really fitting to see that (mis-)quote used in the context of this outrage since from reading through the original vitriolic Twitter thread it's clear that many of the most outraged are incorrect about what the product does.
Just like the writers he talked to and got positive feedback? Everybody not agreeing with you represents "chauvinistic SV attitude"?
No, he didn’t say anything about them. People side against their interests all the time, finding a few writers that like this is trivial. Are those people the majority opinion on this or are we just trying to prove how wonderful this technology is?
Let's recap:
> I launched the prosecraft website in the summer of 2017, and I started showing it off to authors at writers conferences. The response was universally positive, and I incorporated the prosecraft analytic tools into the Shaxpir desktop application [...]
And he goes on mentioning that some authors even reached out to him to get their books added.
Unless you are accusing him of lying or unreasonably overstating the response he got ("universally positive"), for which I really don't see any indication, then a statement like "finding a few writers that like this is trivial" is not a good faith engagement with this topic/conversation.
“Everybody not agreeing with you represents "chauvinistic SV attitude"?”
…wasn’t very good faith either as it’s unclear whether the writers share the same belief as some tech people that AI and humans doing stuff are the same and use that idea to further a pro AI agenda as opposed to them just finding a useful tool to incorporate into their workflow regardless of the underlying technology or politics. Your response assumed the former and paints parent poster as wrong based on your assumption. Some writers liking the tool, just like some artists liking stable diffusion, doesn’t invalidate the original criticism or imply their ideology.
Indeed my experience jives with what he said. Many AI people I’ve seen comment are very much “adapt or die” when it comes to AI technology, suggesting that writers/artists must (even if begrudgingly) use these tools to stay competitive and see many datasets as fair game even when their authors are against its inclusion in said datasets, such as the author of this article.
this even didn't contain how developers decided to let people lose job. people is angry because they worried about losing job.
But I guess there would also be people up in arms about this.
So yes, I understand what a review is, thanks for the put-down, that certainly added something to the conversation.
I think we are in agreement - doing statistical analysis on written works is entirely a lesser thing than simple review, and is harmless.
Easy: humans are not machines. "X does it all the time, so I should be able to do it" is never a valid conclusion. It depends on the situation.
> In fact it's a terribly useful tool, and once the mob cools off it'll likely return.
Maybe this tool in particular does not "abuse" the books. Maybe this tool in particular is terribly useful. But you can't blame authors and artists for taking a stance against those new algorithms that provably have the potential to automatically "steal" from their work. You can believe that asking ChatGPT to "write a novel in the style of X" is not abusing the copyright, that's fine. And the authors can answer that they fear it has the potential to break their source of revenue to a point where they won't want to publish anything anymore. And they are entitled to it. And maybe someday we come up with licenses that prevent the use as training data (how in the world could one conclude today that "it is most definitely fair use", given that this is a very new way of using IP material?).
The idea that counting adverbs is steal their work to the point they won't want to publish anymore is clearly FUD. As my remark made clear.
I did not mean that, I am genuinely not sure if you rephrased my point to make it sound wrong or if you missed it.
My point was that, IMO, it does not matter to the other whether counting adverbs is stealing their work or not. Probably if you counted them manually they would be fine (and most likely they were fine before generative AI).
What matters to them is that generative AI is trained from their copyrighted material, and they fear it (I would, too).
The day people stop reading my blog because they can just ask ChatGPT and will get something generated (partly) from my material without any kind of attribution, I can promise you I will stop my blog.
And I understand that. It is not their job to learn how the black box works. What they see is that "machine learning models" (which they probably call "AI" now), which are complete black boxes to them (and that's justified: engineers who train them also don't know exactly what they do, but rather test their model on some dataset and judge it from there). And those black boxes are being trained from their copyrighted work and have the potential to generate a ton of money which they will never see.
You can go and say "you guys should learn how the technology works instead of complaining", but let's be honest: probably you are not an expert in AI yourself, and anyway why would the artists have to care? It is a totally legit question that they have: "Why can engineers take my copyrighted work, run it through an algorithm that does stuff no algorithm has done in history at a scale never seen before, make money out of it, and not even consider that maybe they are abusing my IP?".
Before dismissing the artists, you should try to understand their point of view.
If you have not learned the basics of how something works, you have no right for your opinion on it to be considered valid. Period.
Invalid opinions do harm to democracy and endanger our way of life.
That is so wrong it is actually dangerous. Do I need to understand how a nuclear bomb works for my opinion on it to be considered valid? Obviously not. I only need to understand the consequences of it. It does not matter at all how it works, if I am against the fact that it will kill a whole lot of people.
> Invalid opinions do harm to democracy and endanger our way of life.
And engineers have done much, much more to endanger most living animals (including humans) than authors and artists: technology is the reason for the mass extinction we are currently living, and the problems that are coming with climate change. Maybe it's important to start thinking about the consequences of what you do, not only the technicalities of how you do it. And maybe it's high time you start listening to people who are able to think about the consequences of what you do (maybe they understand that better than you do, ever thought of that?), even if they don't know how to do it.
If we work from the nuclear bomb analogy, you certainly don't need to be a nuclear physicist to protest nuclear bombs. You just need to have some a reasonably correct high level understanding of the impact of a nuclear bomb. But that's not what is happening here. This is more like storming the Belgian embassy to stop Belgium from using their nuclear arsenal to trigger a chain reaction in the atmosphere: totally detached from reality in every aspect.
As far as I can tell from your messages on this, you think that the harassment was entirely justified. Is that correct?
I don't think it is totally detached from reality. I believe that engineers are generally pretty bad at realizing the impact technology will have on society. There are many concerns with generative AI in general: it can potentially "break the Internet" (by finishing breaking search engines which already struggle with SEO), or maybe democracy, who knows? Copyright is one such problem.
> you think that the harassment was entirely justified. Is that correct?
I honestly don't know how far it went. What I saw in the article is a few authors who wrote online that they wanted their book removed from that software. Not sure if it is closer to harassment or to lobbying.
What I see, however, is many comments of engineers who don't see the problem with copyright and who don't seem to understand why non-engineers may be against this technology, or why one would even think about forbidding a technology ("but technology is neutral"). My point is just that those engineers should maybe take a step back and try to reflect on that "technology is neutral" belief.
Understanding the consequences of something is a PART of how it works. Since you understand that it can kill a whole lot of people then I'd say you have passed the incredibly low bar.
In this case most of the authors do not understand the consequences of the tool, they think it will generate convincing sound text that sounds like them or that it is serving pirated copies of their books (sourcing that from the original Twitter thread that I unfortunately read a lot of).
This doesn't seem like the thread to debate whether technology is a good thing, but I can't help but call this assertion ridiculous. Technology is responsible for almost every single good thing in the world today.
Because you do? That's my point: engineers believe that because they have some understanding of how machine learning works (and in my experience, usually it is very limited...), they can conclude that they understand the consequences of it. Simple example: the Facebook "like" function, that was supposed to be positive ("oh nice, I got likes"), and actually increases addiction and is mostly negative ("oh no, why did I not get likes?"). Clearly those who implemented the first likes had not realized what consequences they would have.
> Technology is responsible for almost every single good thing in the world today.
If you have a very limited view of the world, I guess it could be. I like trees, flowers, bees, birds, mountains, snow. Can you tell me which ones come from technology? Let me help you: most of them are threatened to die in this century because of technology. For most living species, every single improvement technology is bad news. To the point where it is now globally becoming bad news for humans, because it's quite likely that we will get into global instability, wars, and famines in the next few decades because of technology. Think about it when we start having billions climate refugees, and think about how you were dismissing opinions contradicting your beliefs based on the fact that you understand some implementation detail.
But let's even ignore the fact that the next few decades will most likely get pretty bad for us. It is true that right now, we live longer, we have more food (and obesity problems), and we can cure many diseases that we could not in the past. Does that mean we are happier? Happier than whom? Vikings? Ancient romans? Ancient greeks? That question seems closer to history and philosophy... why does your opinion count then? Are you historian/philosopher?
I feel like you miss the point of a law. You seem to read the law, and say "well, the law says X, new technology Y is compatible with it, so that's legal, everyone is happy". But that is wrong. The law reflects the society we want. Do we want a society that completely kills creative work because Big Tech found a loophole to launder their IP? I guess we all agree that we don't. It is not clear if LLM is that loophole, I agree. But you seriously have to take a step back and think about that. What if it does? Then we may have to redefine the meaning of "fair use".
Maybe this particular software was not a danger for those authors. But they don't know that. And given that most engineers talking about LLMs don't seem to remotely understand how one could be worried about it, I understand that they start speaking up wherever they can't. Because clearly it does not seem like those who build those systems give a damn about copyright holders.
There is no need to shoehorn that debate into this particular situation, and I see no merit in defending authors that had a knee jerk reaction to this project on the grounds that they have reasonable fears about other types of projects.
Engineers tend to globally think that LLMs are not really a problem for copyright holders. At least those who develop LLMs pretty clearly don't give a damn. And on top of that, it is in their interest to not be constrained by copyrights.
If this is my feeling (that engineers globally don't care about copyright holders), then it seems reasonable to me that non-engineers could feel the same. That sounds fair, doesn't it?
So those people start speaking up when they see a situation where they feel like "it is happening". And because they don't really know the technology, it is hard for them to know if this particular case is a problem or not. And they can't really trust engineers to tell them, because engineers built LLMs in the first place, and really it does not seem like they care about copyright holders.
Finally, engineers see this reaction from authors, and instead of trying to understand where they come from, they dismiss their opinion. Which probably will reinforce the feeling that engineers don't remotely understand the concerns of those people, and keep building their AI-powered laundering machines. Again, engineers working on those technologies in big companies have absolutely no interest in even considering that it is a problem. Because they get a big salary to help their big company get more profitable, even if it kills many jobs and is a net loss for society (because they benefit from that).
1) Some engineers (or more broadly, software developers) do not respect copyright
2) Therefore you reasonably are skeptical of projects related to material under copyright.
3) It is not always obvious if a project is respectful of copyright.
Now, applying these #1,#2,#3 you believe they justify the outrage for this particular project.
I disagree, because outrage combined with a lack of understanding (#3) is pretty much my definition of a knee-jerk reaction and vastly counterproductive to the interests of copyright holders because it will make the dismissiveness you predict a self-fulfilling prophecy.
No, I believe it explains it.
> it will make the dismissiveness you predict a self-fulfilling prophecy.
That's the thing: both parties need to listen to each other. The problem here is not this particular project, but the fact that we are not addressing the bigger concern which is LLMs.
IMHO, it is completely useless to try to solve this particular case, because it will happen over and over again. We need to address the LLM issue.
"Needs to stop." OK, you could be right on that one too. I don't think you are, but that's not the point.
Neither of those adds up to "it's currently illegal". (Whether it's actually illegal probably depends on the details of how he did what he did.)
Further, neither of those things adds up to "the howling mob should attack him until he stops". (Even if the "attacks" are purely online.) I am against "attack him with outrage dialed all the way up to 11 without actually understanding what his tool is and does". I am also against giving in to the outrage - it just shows the mob that baseless outrage attacks work.
You think it needs to stop? Fine. Persuade him that it needs to stop, and therefore that he should stop. Convince him - not with a mob screaming in outrage, but with reason.
https://metametricsinc.com/parents-and-students/lexile-for-p...
Not everything can be meaningfully quantified. Not everything needs to be.
Ok, so who decides what's OK to analyze or not? Is there some obvious moral line I fail to see, that everyone would immediately agree on?
It seems the project was about analyzing books, not about producing new books. How is that hurting the authors?
Which is what will happen if the authors don’t proactively stop it from happening. Look at how the music industry has evolved over time.
Let me rephrase your question: "how is it different to the current process, other than <the fact that it is different>?" :-). I would say that the answer lies in the question.
My point was that it is different: when humans read a book, they don't train a machine learning model. They can't read as many books as a machine, at the same speed, and they can't remember nearly as much as what a machine can.
Humans and computers are fundamentally different, and it matters. You can't conclude that because it works for one, it will fork for the other.
You seemed to be saying that the differences I listed (quicker and more specific feedback) were the only differences. Those are both positive.
I was saying that some people may think there are negative differences as well.
I am actually on the side that LLMs are a big problem for copyright, and I don't want my code and blog posts to be used in their training dataset without my consent. To me, at this scale, it's not fair use. IMO it's a bit like if Facebook said that it is fair use to leverage metadata about their users, because "someone who sees you in a public space talking to a friend knows that you are talking with that person, and it is the same for Facebook on social media". My problem is not that Facebook knows that I sent a message to a friend now, but rather that they know who writes to whom and when, at scale.
Similarly my problem is not that somebody could read my blog post, learn from it, and write another blog post. My problem is that LLMs automatically train on all written material they want on the Internet, at scale, and without acknowledging that all that material has a lot of value (and is copyrighted).
I think fair use should somehow consider the scale.
it is objective but potentially biased. and it could even be discriminating if the input for this tool isn't diverse enough. but these are the issues that can go wrong with any use of technology, and we have seen many examples of that happening. however i don't think that is problematic if writers use it to analyse their own texts in comparison. it is however a serious issue if publishers use it to decide what to accept
Nowadays writers can at least publish their books without the need of publishers and I think some like the help of the bad Silicon valley stuff that made writing, publishing and interacting with the readers easier.
I'm on your site if it's about automatic content creation and style copying but text analysis is not the real danger. Especially when the usefulness of such statistics isn't even given.
Except those are very likely to be metoo vampire novels. And lately LLM generated.
I'd move that on the contrary, the role of the publisher as a curator will only become more important in the future.
> It seems the project was about analyzing books, not about producing new books. How is that hurting the authors?
"Vivid books are really in this year, we're gonna have to ask that you aim for a Vividness(tm) of 85 or above."
"US books have 15% more adjectives, clearly this is proof of our superior detail-oriented work ethic!"
"What does the rise in Emotion(tm) have to say about the decline of society?"
Patterns rarely show themselves before we investigate.
In science they call this trap P-hacking. Even data "scientists" know to be wary of overfitting. We're really good at finding patterns, but few of them actually mean anything.
> In science they call this trap P-hacking. Even data "scientists" know to be wary of overfitting. We're really good at finding patterns, but few of them actually mean anything.
Quantifying things is not always p-hacking. When people do experiments on novel materials or structures they quantify the data, make readings and record them, and then look for patterns. For example measuring the electronic properties of a new novel nano structure or molecule.
When I think of p-hacking[1] I think of using the same static data and doing various data analysis over and over again until something potentially interesting is found and ignoring the risks of false positives as you do so.
Publishers rejecting manuscripts because "this years trend shows customers are looking for vividness in the 70+ percentile, your book is only at 55". Everything becoming the same style. If you thought Hemingway, Joyce or Nabokov had it bad with rejections, there'd be zero chance for actual innovative writing to break through the walls of The Algorithm.
Sure, but written words _can_ be meaningfully quantified. We have been doing that for thousands of years. Starting with numerology and other mystical/religious beliefs, poem metrics, stylometry, crypto analysis, stroke counting, to name a few.
> Not everything needs to be.
Why not?
I would argue that "Offensive" is either hyperbolic or you've used the wrong word.
> the idea that it's, in any way, shape, or form, useful is bafflingly laughable.
I don't know if it's useful because I never tried it. I might harbour my doubts but I'd like to find out. This is how I approach new things.
The later creates massive competition to human writers, the former is just an information for potential readers.
https://news.ycombinator.com/item?id=37042561
Pure text statistics won't do the same.
If it's gibberish you know you got scammed, LLM texts look convincing so you don't know for sure.
https://11points.com/11-amazing-fake-harry-potter-books-writ...
It's worse for new authors, they disappear between all the AI authors.
Publishers and readers will have to search a bigger haystack to find the needle.
The better an LLM can complete your joke the worse it is, for instance. Important to have a good Letterman-MacDonald quotient.
Usefulness is immaterial here.
Is he allowed to do this? Yes.
What's wrong with presenting a page count and word count, for example?
Anyone who is with the artists should pass a law. Moral outrage is not law.
Technology has to be protected from dumb people, or is it worth protecting dumb people from technology....
I am pretty confident they haven't. Sounds like you've set yourself up for a reverse "true scotsman" here ;)
If an AI tool was killed, I consider it a victory. That's because even if there are some small useful applications of AI, AI on the whole will certainly put most creatives out of business.
Instead, I propose the following: anyone who is interested in preventing AI from taking over their craft should join me in a coalition of ban AI from their own business. By placing a notice that your work is "100% AI FREE", you are doing something akin to the fair-trade/sustainably sourced sticker on chocolate or other food products: you are letting consumers know that your work was made by a human, so that they can support you.
If enough people get in on this, and pledge to support only those creators who don't use AI, then we can make AI an unprofitable venture and hopefully kill it forever!
I already put a 100% AI FREE badge on my YouTube channel, which means that I will never use AI for writing scripts, editing videos, producing images, etc. Moreover, I also pledge to support other creators who pledge never to use AI, by buying their products over others!
However, for practical purposes, a direct definition that encompasses every situation is not necessary, but can evolve. For now, I think we do not need a precise definition and we can start with the following: AI such as ChatGPT, LLMs, image generation tools like DALL-E and ohters, should be restricted.
As for YouTube's algorithm, I agree it is also dangerous. For now, I have restricted the use of direct content generation algorithms, in other words, all content can reasonably said to be human generated in terms of writing, composition, etc.
In other words: AI that makes any creative decision in making content should be banned. Other algorithms should be carefully debated.
But returning to the topic: even though AI has some benefits, I believe that AI in the long run will have negatives that FAR outweigh the positives, so I believe it still should be restricted.
As for translation, well, the AI transcription/translation sucks. I do attempt to put manual captions in my videos as much as I can though.
Also, what do you mean by "forces any sort of editing on my videos through AI". Do you mean like, changing the actual content of your videos?
What's your take on generative fill in Photoshop?
Some creative people are using these technologies, and while it is quite human guided NOW, at some point, the guidance that humans put into it will lessen. That's not to say that AI will ever produce a work like Dostoevsky --- maybe it won't, but it WILL be enough to eliminate most creative jobs, and reduce them to being at most being supervised by people who don't have much of a passion for creative works. And that's a shame, because it will remove the passion of creativity from society.
Generative fill: I don't use it, and that's part of my personal ban. It goes too far. I only use traditionl editing techniques in my photography that works with basically what is there.
Yes, you can say that photography has always been about manipulation, but basically, I have a personal line that I believe I can define sufficiently well, that is far behind the line of AI.