I mean think about it. Amazon had to stop publishing BOOKS because it can no longer separate the signal from the noise. The printing press was the birth of knowledge for the people and the LLM is the death.
I mean think about it. Amazon had to stop publishing BOOKS because it can no longer separate the signal from the noise. The printing press was the birth of knowledge for the people and the LLM is the death.
That's true, but it also allowed protestant "heretics" to propagate an idea that caused a permanent schism with the Catholic church, which led to centuries of wars that killed who-knows-how-many people, up to recent times with Northern Ireland.
(Or something like that, my history's fuzzy, but I think that's generally right?)
Not even close to white noise. White noise, in the context of the token space, looks like this:
auceverts exceptionthreat."<ablytypedicensYYY DominicGT portaelight\- titular Sebast Yellowstone.currentThreadrition-zoneocalyptic
which is literally the result of "I downloaded the list of tokens and asked ChatGPT to make a python script to concatenate 20 random ones".
No, the biggest problem with LLMs is that the best of them are simultaneously better than untrained humans and yet also nowhere near as good as trained humans — someone, don't remember who, described them as "mansplaining as a service", which I like, especially as it (sometimes) reminds me to be humble when expressing an opinion outside my domain of expertise, as it knows more than I do about everything I'm not already an expert at.
Specific example: I'm currently trying to use ChatGPT-3.5 to help me understand group theory, because the brilliant.org lessons on that are insufficient; unfortunately, while it knows infinitely more than I do about the subject, it is still so bad it might as well be guessing the multiple choice answers (if I let it, which I don't because that would be missing the point of using a MOOC like brilliant.org in the first place).
That's because they are trying very hard not to check what they are selling, hoping that their own users and a few ML algorithms can separate the signal from the noise for them. It seems to me that the approach is no longer working, and they should start doing it by themselves.
I get the most value out of asking for examples of things or asking for basic explanations or intuitions about things. And I get so much value from this that I really think the printing press is the most apt comparison.
So it is possible that LLMs will centralize the production and dissemination of knowledge, which is the opposite of what people think the printing press did. I hope I'm wrong and open models can challenge/overtake state of the art models developed by tech giants, that would be amazing.
Now it refuses, because OpenAI's morals apparently don't include spreading openly available knowledge about how to defend yourself.
Scary. I have also been using it to generate useful political critiques (given a particular theoretical tradition, some style notes, and specific articles to critique, it's actually excitingly good). What if OpenAI decides that's a threat? What reason do we have to think that a powerful institution would not take this course of action, in the cold light of history?
"Signal" would mean new data, which is by definition not possible via LLMs trained on publicly available content, since that means the data is already out there, or new and meaningful ideas or innovations beyond just combining existing material. I have not seen LLMs accomplish the latter. I consider it at least possible that they are capable of such a feat, but even then the relevant question would be how often they produce such things compared to just rearranging existing content. Is the proportion high enough that unleashing floods of AI-generated content everywhere would not lower the signal-to-noise ratio from the pre-AI situation?
Doesn't that make human content look bad in the first place?
If we can't distinguish a Python book written by a human engineer or by ChatGPT, how can we demonstrate objectively that the machine-generated one is so much worse?
I bet ChatGPT can come up with above-average content to teach Python.
We should teach beginners how to prompt engineer in the context of tech learning. I bet it's going to yield better results than gate-keeping book publishing.
Now that it’s possible to produce mediocrity at scale, that process breaks down. How is a beginner supposed to know whether the tutorial they’re reading is a legitimate tutorial that uses best practices, or an AI-generated tutorial that mashes together various bits of advice from whatever’s on the internet?
There are almost always trade-offs and choosing one option usually involves non-tech aspects as well.
Online tutorials freely available very rarely follow, let's say, "good practices".
They usually omit the most instructive parts, either because they're wrapped in a contrived example or simplify for accessibility purposes.
I don't think AI-generated tutorials will be particularly worse at this to be honest...
I've seen many developers using technologies without reading the official documentation. It's insane. They make mistakes and always blame the tech. It's ludicrous...
Humans, examining things, and putting a reputation that matters on the line to vouch for it.
The fact that Amazon doesn't want to have smart, contextually aware humans look at and evaluate everything people propose to offer up for sale on their storefront doesn't mean it can't be done. Same as how Google doesn't want to look at every piece of content uploaded to YouTube to figure out if it's suitable for kids, or includes harmful information. That's expensive, so they choose not to do it.
The LLM is the most powerful knowledge tool ever to exist. It is both a librarian in your pocket. It is an expert in everything, it has read everything, and can answer your specific questions on any conceivable topic.
Yes it has no concept of human value and the current generation hallucinates and/or is often wrong, but the responsibility for the output should be the user's, not the LLM's.
Do not let these tools be owned, crushed and controlled by the same people who are driving us towards WW3 and cooking the planet for cash. This is the most powerful knowledge tool ever. Democratize it.
Yeah, I mean, so can I, as long as you don't care whether the answers you receive are accurate or not. The LLM is just better at pretending it knows quantum mechanics than I am.
The best way to use an LLM for learning is to ask a question, assume it's getting things wrong, and use that to probe your knowledge which you can iteratively use to prove the LLM's knowledge. Human experts don't put up with that and are a much more limited resource.
The current iteration of the internet (more specifically social media) has used the same rationality for its existence but at a level, society has proven itself too irresponsible and/or lazy to think for itself but be fed by the machine. What makes you think LLMs are going to do anything but make the situation worse? If anything, they’re going to reenforce whatever biases were baked into the training material, of which is now legally dubious.
LLMs (Current-generation and UI/UX ones at least) will tell you all sorts of incorrect "facts" just because "these words go next to each other lots" with a great amount of gusto and implied authority.
Maybe they’re just orders of magnitude more useful at the beginning of a career, when it’s more important to digest and distill readily-available information than to come up with original solutions to edge cases or solve gnarly puzzles?
Maybe I also simply don’t write enough code anymore :)
Just yesterday, I asked if Typescript has the concept of a "late" type, similar to Dart, because I didn't want to annotate a type with "| null" when I knew it would be bound before it was used. Searching for info would have taken me much longer than asking the LLM, and the LLM was able to frame the answer from a Dart perspective.
I would say that that information neither "important to digest" nor "readily available."
For me, it's been able to give very good answers when they were within the first few Google results when searched for using the proper terms (but the value is in giving you these terms in the first place!).
For questions from my field, it's been wildly hallucinating and producing half-truths, outdated information, or complete nonsense. Which is also fair, because the documentation where the answers could be found is often proprietary, and even then it's either outdated or outright wrong half of the time :)
We can generate thoughts that are spatially coherent, time aware, validated for correctness and a whole bunch of other qualities that LLMs cannot do.
Why would LLMs be the model for human thought, when it does not come close to the thoughts humans can do every minute of every day?
Aren't we all just stochastic parrots, is the kind of question that requires answering an awful lot about the universe before you get to an answer.
But sometimes we're more than that: Some types of deep understanding aren't verbal or language-based, and I suspect that these are the ones that LLMs will have the hardest time getting good at. That's not to say that no AI will get there at all, but I think it'll need something fundamentally different from LLMs.
For what it's worth, I've personally changed my mind here: I used to think that the level of language proficiency that LLMs demonstrate easily would only be possible using an AGI. Apparently that's not the case.
It's all just a matter of perspective.
Yes, right now it looks like white noise, just like back then it looked like white noise which could drown out the religious texts. But we managed to get past it then and I'm sure we'll manage now.
If you actually meant something else, you should probably clarify.
It can be both true that right now predominantly low quality content emanates from LLMs and at some future time the highest quality material will come from those sources. Or perhaps even right now (the future is already here, just unevenly distributed).
If that was their reasoning, I tend agree. The equivalent of the Catholic Church in this metaphor is the presumption human-generated content's inherent superiority.
If we cannot distinguish, I'd argue they have similar value.
They must have. Otherwise, how can we demonstrate objectively the higher value in the human output?
Heck, we always did that since before GPT.
Good authors will continue to publish good content because they have a reputation to protect. They might use ChatGPT to increase productivity, but will surely and carefully review it before signing off.
If yes, well, there's the problem then. It's not AI, but the lack of guidance and research skills in support of the process of choosing a book.
There are lots of fake recipe books on amazon for instance. But how can you really be sure without trying the recipes? It might look like a recipe at first glance, but if its telling you to use the right ingredients in a subtly-wrong way, its hard to tell at first glance that you won't actually end up with edible food. Some examples are easy to point at, like the case of the recipe book that lists Zelda food items as ingredients, but they aren't always that obvious.
I saw someone giving programming advice on discord a few weeks ago. Advice that was blatantly copy/pasted from chat GPT in response to a very specific technical question. It looked like an answer at first glance, but the file type of the config file chat GPT provided wasn't correct, and on top of that it was just making up config options in attempt to solve the problem. I told the user this, they deleted their response and admitted it was from chatGPT. However, the user asking the question didn't know the intricacies of "what config options are available" and "what file types are valid configuration files". This could have wasted so much of their time, dealing with further errors about invalid config files, or options that did not exist.
I think the concern is that bad authors would game the reviews and lure audiences into bad books.
But aren't they already able to do so? Is it sustainable long term? If you spit out programming books with code that doesn't even run, people will post bad reviews, ask for refunds. These authors will burn their names.
It's not sustainable.
They make up an authors name. Publish a bunch of books on a subject. Publish a bunch of fake reviews. Dominate the search results for a specific popular search. They get people to buy their book.
Its not even book specific, its been happening with actual products all over amazon for years. People make up a company, sell cheap garbage, and make a profit. But with books, they can now make the cheap garbage look slightly convincing. And the cheap garbage is so cheap to produce in mass amounts that nobody can really sort through and easily figure out "which of these 10k books published today are real and which are made up by ai".
It takes time and money to produce cheap products at a factory. But once these scammers have the AI generation setup, they can just publish books on loop until someone ends up buying one. They might get found out eventually, and they will have to pretend to be a different author, and they just repeat the process.
The LLM allow DDoS attack by increasing the threshold needed to check the books for gibberish.
It’s not like this stream of low quality did not exist before, but the topic is hot and many grifters try LLMs to get a quick buck at the same time.
As an aside, the case you're thinking of was a novel, not a recipe book. Still embarrassing, but at least it was just a bit of set dressing, not instructions to the reader.
https://www.cnet.com/culture/zelda-breath-of-the-wild-recipe...
> I saw someone giving programming advice on discord a few weeks ago. Advice that was blatantly copy/pasted from chat GPT in response to a very specific technical question.
This, on the other hand, is a very real and a very serious problem. I've also seen users try to get ChatGPT to teach them a new programming language or environment (e.g. learning to use a game development framework) and ending up with some seriously incorrect ideas. Several patterns of failure I've seen are:
1) As you describe, language models will frequently hallucinate features. In some cases, they'll even fabricate excuses for why those features fail to work, or will apologize when called out on their error, then make up a different nonexistent feature.
2) Language models often confuse syntax or features from different programming languages, libraries, or paradigms. One example I've heard of recently is language models trying to use features from the C++ standard library or Boost when writing code targeted at Unreal Engine; this doesn't work, as UE has its own standard library.
3) The language model's body of "knowledge" tends to fall off outside of functionality commonly covered in tutorials. Writing a "hello world" program is no problem; proposing a design for (or, worse, an addition to) a large application is hopeless.
Hard disagree. I've used GPT-4 to write full optimizers from papers that were published long after the cutoff date that use concepts that simply didn't exist in the training corpus. Trivial modifications were done after to help with memory usage and whatnot, but more often than not if I provide it the appropriate text from a paper it'll spit something out that more or less works. I have enough knowledge in the field to verify the corectness.
Most recently I used GPT-4 to implement the paper Bayesian Flow Networks, a completely new concept that I recall from the comment section on HN people said "this is way too complicated for people who don't intimately know the field" to make any use of.
I don't mind it when people don't find use with LLMs for their particular problems, but I simply don't run into the vast majority of uselessness that people find, and it really makes me wonder how people are prompting to manage to find such difficulty with them.
ChatGPT as an autocompletion tool is fine, IMO. As well as generating alternative sentences. But anything longer than a paragraph falls back to the uncanny valley.
These pseudo-authors will get bad reviews, will lose money in refunds, burn their names.
It's not sustainable. Some will try, for sure, but they won't last long.
The equilibrium shifts to making it much harder to find good books, and that was already hard enough.
Choose a book from someone that has a hard earned reputation to protect.
Before the printing press two books cost around the same as a 2 story cottage.
Afterwards a couple books would be about a month of wages for a skilled worker.
That greatly limits ones ability to drown out anything with books.
Not for centuries. Due to the expense of the technology and the requirement in some locations for a royal patent to print books, the printing press just opened up knowledge a bit more from the Church and aristocracy to the bourgeoisie, but it did little for the masses until as late as the 1800s.
Maybe the self-publishing and BoD will decline in the long term due to ML white noise and publishers are a sign of quality again.
Prompt engineering is an example of this. A clever prompt by a domain expert can prime an LLM interaction to yield better information to the recipient in a way that the recipient themselves could not have produced on their own.
With the press, a greasy workman can churn out hundreds of copies an hour, for whichever charlatan or heretic palms him enough coin. The people are flooded with falsehoods by men whose only interest in writing is how many words they can fit on a page, and where to buy the cheapest ink.
The worst part is that it is impossible to distinguish the work of a real thinker from that of a cheap sophist, since they are all printed on the same rough paper, and serve equally well as tomorrow's kindling.