We Automated Bullshit
cst.cam.ac.uk
cst.cam.ac.uk
However at some point you have to admit the LLM does generate things that are good answers. They might be good answers that happen to pass the smell test, but they are nonetheless good answers. For instance when you ask it for a snippet of code and it gets it right.
And here is the crucial thing: you need to already know what you're doing to know whether the LLM got it right. I'm no historian, and I can ask cGPT for an essay about the causes of the Great War. When I get the answer, it sounds right to me. I don't know if the essay talks about the things an actual historian would find important, all I know is that it gives me the vanilla answer that some layman who has read a little bit would think was the right answer.
Now there's another issue this brings up. Most of us are experts in one field only. What is stopping the LLM from fooling me in every field that I don't know anything about? I best be wary of using it outside of my area of expertise.
So in the current iteration, I think LLMs are a shortcutting tool for experts. I can tell when it spits out a snippet of code that is correct, and when it's wrong. Someone who wasn't working in my domain would get fooled.
Isn't the hint in the name? It's a language model, not a knowledge model. LLMs are exceptional at generating stuff that passes as coherent language (like, the parts of speech are all where one would expect them to be in the sentences). The trouble is people think it goes deeper when the knowledge modeling part only happens incidentally because language and knowledge modeling are so closely related.
Well yes, but how is this insight helpful? People aren't just looking at the name. Have your missed the last 10 months of ridiculous breathless coverage of LLMs?
> The trouble is people think it goes deeper when the knowledge modeling part
People are literally calling this shitty text prediction algorithm "AI". The discourse is completely out of hand.
https://en.m.wikipedia.org/wiki/Dartmouth_workshop
> An attempt will be made to find how to make machines use language, form abstractions and concepts, solve kinds of problems now reserved for humans, and improve themselves.
LLMs fit that definition well. They're not AGI, but I think it's fair to call them AI.
What I think is the problem is that when people hear "artificial intelligence," they don't think of chess programs anymore; they think of something like Skynet from the Terminator. Or, maybe now they think of ChatGPT.
If they don't get to define "AI", who does?
True. But most of the training corpus was not linguistically-sound gibberish; it was actual text. There was knowledge encoded in the words. LLMs are large language models - large enough to have encoded some of the knowledge with the words.
Some of. And they only encoded it. They didn't learn it, and they don't know it. It's just encoded in the words. It comes out sometimes in response to a prompt. Not always, not often enough to be relied on, but often enough to give users hope.
What we really need is a language model coupled to a knowledge model.
People say stuff like this a lot, and I never know what to do with it. I'm not sure what specifically I'm supposed to get out of these distinctions.
And that's enough to be a problem, given how people are trying to use LLMs. They're trying to use it like a human who "knows" (whatever that means). And whatever an LLM actually does, it doesn't do that.
If the claim is only "LLMs don't work like humans" and "LLMs make a lot of factual and logical errors", then that seems to be correct as far as it goes, but it doesn't go very far.
>Some of. And they only encoded it. They didn't learn it, and they don't know it. It's just encoded in the words. It comes out sometimes in response to a prompt. Not always, not often enough to be relied on, but often enough to give users hope.
And some of the pieces of "knowledge" in the training corpora were wrong, lies, or bullshit themselves.
Garbage in, garbage out.
However, stating that LLMs predict text, not facts - and that LLMs don’t “think”, is enough to start a debate.
Yeah the issue is anthropormphization and the “cool” factor of demos, vs actual experience.
Actually the hint is not in the name, that only tells a portion of the truth. If LLM were trained on textbooks teaching foreign languages from every other language, the name would suffice. But it does not, LLMs are trained on wide variety of texts, many of them social media, blog, and general web posts. The larger web is an ocean of bullshit, and that is where the bullshit within LLMs comes from.
'Textbooks are all you need' ought to be how LLMs are trained so people can use them for their "knowledge work" without bullshit ramifications.
But in many spheres of knowledge there are still a lot of contradictory opinions in textbooks. Why did Rome fall?, What's the best X? In science there have been many schisms: quantum theory, tectonic plates, GMO/Organic etc.
Writing code is probably one of the few areas where just about all content is 'good' - there may be a lot of bad or poor performing code in books and online but the vast majority of it 'works'.
For many other fields, introductory textbooks even at the university level may abound in oversimplifications, and it is only in later years of study that one begins, by reading specialized papers and monographs, to grasp how complicated the facts really are.
Not only that: there's also the fact that LLMs do not "fact check" their output against any data source. They just generate text based on relative frequencies of different kinds of text in their training data. They don't even have the concept of text being related to other things in the world; in fact they don't even have the concept of "other things in the world".
Any text generator built that way will output bullshit, no matter how accurate its training data is. "Bullshit" does not mean the output is necessarily wrong; it means, just as Frankfurt said, that the thing generating the output doesn't care about whether it's true or not. LLMs meet that criterion: they don't care about anything other than relative frequencies in their training data.
Obviously, they're better at grammar than many other things, but it's clear that they contain a vast amount of knowledge. They contain the information that the sky is blue and firetrucks are red, which isn't a matter of grammar or coherence, but is simply information about the world.
Correct. But it does mean they're about "language" rather than knowledge. It's meant to read, parse and generate language, no matter if it includes prose, facts, knowledge or whatever.
> Clearly, they aren't as likely to generate "Colorless green ideas sleep furiously" as "It's important to feed your dog".
Clearly, as the training set probably contains much more instances of the latter, than the first example. And if it's more probable to generate, that's the text it'll generate. It doesn't understand why or how, just that it's more likely, so that's the way it's going.
Not really. It was invented because they were trained on natural language corpuses. It wasn't a deep, philosophical concept.
It continues to be used for models that are doing completely non-languagy stuff, such as robotics, although I expect the growing use of 'foundation model' or some other less-confusing term will likely replace it.
If it wasn't, what would we be doing here?
To some extent, those are kind of the same thing.
They are not the same thing. They are not even "kind of" the same thing, for any reasonable definition of "kind of." That should be obvious, because a lie is language, which is pretty close to the opposite of knowledge.
I mean no offense, but your hedging words are doing so much work that I think your statement counts as bullshit.
Edit: thinking about it more, I think your comment shows the particular kind of bullshit that often characterizes tech people's thinking and is so useful for hyping shit (e.g. take something hard, lossily reformulate it into something easier but vaguely similar, solve the easy thing, then declare you solved the hard thing).
Knowledge is... concepts and their relationships, typically with the implication of being grounded in objective reality (although one can be knowledgeable about things like, say, Tolkien, or Star Wars, or fantasy tropes more generally).
.
Language is (a particular way of representing) knowledge, in about the same way that a particular .cpp file is (a particular way of representing) a linked list or whatever.
So, in your view, are .cpp files and linked lists "kind of" the same thing?
Language can be used to express some kinds of knowledge, but they're not even "kind of" synonymous. Language can express a lot of things that aren't knowledge, and there's knowledge that can't be expressed with language (qualia, at least).
All the more and less formal models we invent for describing "language"? This is just fitting curves to a snapshot of reality. So sure, "a lie is language, which is pretty close to the opposite of knowledge" - a lie can be expressed in the formal model. But the truth (as subjectively understood) is much more likely to be spoken or written down in practice.
Communication isn't random. Language isn't random. Learning it is picking up the patterns.
>>>> But it does mean they're about "language" rather than knowledge.
>>> To some extent, those are kind of the same thing.
>> They are not the same thing.
> Of course they are the same thing.
Language and knowledge (both broadly construed) are obviously not the same thing. Language can encode non-knowledge like nonsense or lies, and there is knowledge (e.g. of the experience of qualia) that can't be expressed in language.
I think the point up-thread is true: something that knows only about how language is unreliable source of knowledge. Even if all its input is true knowledge, it can still blindly combine that input in linguistically plausible but false ways.
you can use language as a proxy for knowledge in some cases, but doing so guarantees generating bullshit, which is the entire point of the OP.
it does mean they're about "language" rather than knowledge
Language in part encodes knowledge.They can emit statements that happen to be true because the probabilities fall that way (as a sibling comment points out.) But it's not "justified", because models don't have any ability to remember or cite sources, or even internally classify tokens as true or false (it's all just weights.)
(And it's not a "belief", either, because there's no intentionality, but that's a more slippery concept.)
Therefore, philosophically speaking, it is fundamentally impossible to acquire knowledge from a LLM unless you verify ("justify") a statement via some secondary, more trustworthy channel.
FWIW, although 'justified true belief' is classic in the sense of being an old definition of knowledge, I don't think any significant number of modern philosophers would use it.
But yeah the problem what you mean by "knowing." For example, I am perfectly happy to say that a LLM "knows" words just like I "know" words, or that it "knows about" very popular topics in a loose sense.
But that's a very different sense of "knows" from the kind of specific propositional knowledge we're usually interested in when using a LLM as a search engine, fact checker or source of information!
I’m clearly keeping track of how strongly I believe a variety of things. Not how I started believing them to begin with, however, so I lack the justification. And in a lot of cases, the lost justification would have been: “Because someone else said it, so now I’m aping them.”
But I think almost none of them would ascribe "knowledge" to a statistical engine that output propositions of random truth values weighted by how many examples of that proposition it's seen in it's training data.
You can view the logprobs in the API and literally watch it choose to say different things based on it's internal dice rolls.
Knowledge without a feedback mechanism for verification is prone to veer off course.
Right now the feedback mechanism is based on human reinforcement.
What if we put an LLM in an android’s body and let the machine interact with the world through the same senses that we do?
Ultimately though this is all about the definition of words. Humans find the distinction between "true" and "false" to be important and meaningful; animals don't care. AIs don't care (being incapable of caring), but humans care very much about the propositional veracity of statements that AI makes.
The important thing is that we're on the same page about what's actually happening with these systems and how they work and don't slip into magical thinking or assume that just because a system is facile with language it automatically has a sense of truth or right or wrong.
This is only going to get more complicated as models get more sophisticated.
https://www.youtube.com/watch?v=yhfl7kasjZc
There are a quite a few videos with the same reaction. There does seem to be something going on.
Is it though?
There's nuance in what we call knowledge, but to me, knowledge is not uni-dimensional. What I mean by that is, I know when I _know_ something, and I have learned that that is not a trivial skill.
How do I know I _know_ something? Usually because I have searched for evidence from multiple angles / dimensions / senses. When multiple independent observations of quality agree, and my conclusion matches the result, it is safe to say that I know that fact.
An LLM may have multiple 'dimensions' in a linear algebra sense, but does it have independent information?
> How do I know I _know_ something? Usually because I have searched for evidence from multiple angles / dimensions / senses. When multiple independent observations of quality agree, and my conclusion matches the result, it is safe to say that I know that fact.
I wanted to thank you for laying out this perspective. In my experience, in discussions about LLMs, a lot of time people get caught up on words like 'know' which work well for talking about humans, but don't cash them out in definitions I know what to do with. This is a really content-ful of sharing what you mean by your claim.
This is technically true, but no one actually does much of it consciously. But this is how you picked up every skill and understanding, including language, since you've been born.
And this - the unconscious fusion of correlated patterns - is pretty much what LLMs undergo in training.
EDIT:
Also, a lot of confusion is created by using terms like "a fact" or "a conclusion" or "a result", as if they were fixed points in space, fully defined mathematical symbols. But there exist no such thing. There is no binary "know/don't know", there is no binary "result match/does't match the conclusion". Those are all continuous quantities, that get rounded off for convenience.
There is a setting right there called temperature, no?
I think there's a strong argument to be made (that has been made) that much of our knowledge capability is built on language - that language did not simply allow us to communicate our knowledge, but actually allowed individuals to think in a new way.
I certainly agree that current LLMs have no regard for truth and logic at their core, and there are clearly areas where their ability to read and write does not translate to producing true statements. LLMs can't build things out of language like humans can.
I am less sure that LLMs don't have a long way to go before we reach the limits of what modeling language can do.
I've had trouble finding it, but years ago, I read the abstract of a study which looked at chess performance, pitting people who had access to some basic chess software against those who didn't. What they found is that the strongest predictor of success was whether a player could utilise the software effectively, rather than pure chess skill.
Another area where I don't think there are any published studies, but there's a lot of anecdotal evidence, is algorithmic trading, where successful firms almost always have human traders monitoring the algorithmic behaviour. This means that humans can spot when the algorithm is doing something a bit odd, or when conditions aren't conducive to the algorithm performing well, and they can tweak some parameters on the algorithm to get a better result, shut it down if it looks like it's lost the plot, etc. (And let's be honest, a lot of this oversight is probably a result of Knight Capital.)
My point here is: You're entirely right, LLMs are currently a tool for people with experience in the field they're working with. But that's still a pretty big step forward -- just like top quant hedge funds consistently outperform human-managed funds (see: Rentech), smart people who learn how to leverage these tools are going to outperform. The key is learning where they perform well, where they perform poorly, and how to validate their output.
There was a time when this was true - a grandmaster and a computer was stronger than either alone. But that window closed pretty quickly. Today, even the strongest GMs' best strategy is to follow the computer blindly.
Gray-box trading and humans in a control loop over automated trading have both been happening for decades. Knight Capital makes for a pretty spectacular example, but that's about it.
That's NOT how LLMs operate. They're trained on what is supposed to be selected representative human text of some at least minimum quality, and ideally truthful, factual and relevant.
And then their prediction is judged on how close to that ground truth it is. Not on whether it's "plausible" or "will this pass for a correct answer". There's no math function for pretending to be right. You have to be right.
One reason why LLMs make up things sometimes is because they're incredibly tightly optimized. Like you have no idea how tight. GPT-4 has the comparable complexity of a mouse brain. A mouse brain that has to fit the world's knowledge.
Obviously, something has to give, and you'll not be accurate sometimes when you have so limited space. You'll store some things approximately, generally, and then extrapolate answers from them, maximizing for correct output.
When/where/how?
Of course if you feed an AI scientific papers and books of some quality, information will be of high quality. THAT... is how you "filter for truth". It's not supposed or required to be perfect. But of course you can filter the input.
The fact you also add conversational input from various sources like Reddit and Twitter doesn't change this balance significantly. Also even those sources are filtered to exclude vulgarities, certain topics and so on.
Also you're not thinking this through. If a question is on average answered incorrectly 'round the world, it means you're over 50% likely to also give the incorrect answer when asked. Who equipped you to know the truth?
Perhaps in the future, the training data will be more carefully curated to eliminate as many 'wishful truths that do not exist' and 'misinformed conversations between parties with no correct voice' as possible.
Anyway, yes, the "web" in general is full of garbage and propaganda. But then again we're exposed to the same web, and without the pre-processing which OpenAI etc. apply to filter out certain words, phrases, topics (including lots of manual human labor). So it's funny how we're living on that same input, but we're so confident we know what's true and what's wrong.
But LLMs aren't just trained on literally anything on the web. For example consider Wikipedia. That's not one's average drivel for sure. And next to the general web, you need higher quality sources of data in the mix to bias the AI for truth. Scientific papers, good books, transcripts from conferences and presentations. Wikipedia to mention again. So on.
Of course, sometimes LLM make technical mistakes which you can clearly rule as incorrect. Say, plumbus.com's API never had a method getDescription(). But the AI saw many APIs of services like this, which have API getDescription() and given its small network, it's forced to fill-in gaps by inferring the answer based on a vast assortment of contextual clues. Which we call hallucinations (incorrectly). But even the correct answers are inferred. It's all inference. And we all think by inference. We just have 50 times more capacity than the biggest LLM at the moment, so our inference tends to be much more sophisticated, and therefore more likely to be correct. I mean in theory. After all, it's humans and their big brains who produced the garbage that's the web, so.
Children's initial years, being basically lied to about the nature of the adult world, is where their sophisticated inference comes from. The internal intellectual conflict when one realizes it is actually good to tell children the world is a nice, fair, happy place - once figuring out it is not. That right there, I theorize, is where our sophisticated inference is created: we realize our own dynamic minds. We realize the necessity to seat concepts of "a fair world" into people before allowing them to mentally mature and learn it is not.
In my experience so far, more often than not, it doesn’t get it right. Brings to mind the old maxim, “a broken clock is right twice a day”.
I was listening to a podcast yesterday where the purported expert gave an answer that sounded right but based on my direct experience - very wrong.
Of course I could be wrong as well.
Or the experts answer was the same kind of BS the OP was referring to and is so oft generated by LLMs.
My take has been that LLMs aren't valuable for me because I'm an expert and writing code is no longer very difficult. And searching for the knowledge when I need it is faster on my own than having a conversation with a computer because true expertise also involves having mastered and streamlined that process.
Even among senior developers it can be shockingly rare for people to actually understand and implement something like DI/IoC in a correct and useful way.
Edit: #if directives, like static_assert.
~ $ cat sizeofmacro.c
#include <stdio.h>
#define pintsize(T) printf("sizeof(%s) is %zu. Big or small? You decide.\n", #T, sizeof(T))
int main(int argc, char **argv) {
pintsize(int);
pintsize(argv);
return 0;
}
~ $ gcc sizeofmacro.c && ./a.out
sizeof(int) is 4. Big or small? You decide.
sizeof(argv) is 8. Big or small? You decide.the preprocessor doesn't know anything about structs or types
>The sizeof in C is an operator, and all operators have been implemented at compiler level; therefore, you cannot implement sizeof operator in standard C as a macro or function. You can do a trick to get the size of a variable by pointer arithmetic.
> Certainly, you can use a combination of sizeof and #error directives for conditional compilation to achieve a similar effect without static_assert.
// Compile-time check for structure sizes
#if sizeof(Elf32_Phdr) != sizeof(((Elf32_Ehdr*)0)->e_phentsize)
#error "32-bit structure size mismatch"
#endif
#if sizeof(Elf64_Phdr) != sizeof(((Elf64_Ehdr*)0)->e_phentsize)
#error "64-bit structure size mismatch"
#endif
> In this example, if the structure sizes don't match, the preprocessor will generate an error at compile time using the #error directive.> I didn't know sizeof could be used with the C preprocessor
> Yes, sizeof can be used in certain contexts within the C preprocessor.
> It's commonly employed in combination with the #if and #error directives to perform compile-time checks on sizes or other constants.
> This usage allows you to catch potential issues during compilation rather than at runtime.
I should probably go sleep a bit before I embarrass myself any further. After nearly 24 hours awake it seems I have lost the ability to tell macros and directives apart.
I want to call out that this is not necessarily true. You can interact with the LLM using agents, feedback loops, and some scoring function on randomized test inputs to produce novel code that you don't know the structure of beforehand.
You start with a set of example inputs and desired outputs, and prompt the llm for a function in your programming language of choice. You then take the LLM's response and feed it into an agent that executes the code against the inputs and outputs, and report the discrepancies back to the LLM. It responds with a new completion, which you feed to the agent until all your example inputs produce the expected outputs. Finally, you can use property-based testing to produce new examples for testing by the llm, resolving the answers until the produced code is correct to within some margin of acceptable correctness.
You can do this to produce code without needing to know anything besides the desired properties of generated code and the properties of the inputs. You can further automate this by using a separate LLM to produce the examples instead of the typical generation and shrinking functions to produce examples.
This doesn't require any prior knowledge of the target language. You can expand beyond programming into any domain for which you can produce a scoring function and automate input generation.
I suspect that you can use control theory (PID loops, behavior choice loops scored by Eigenvalues) to model complex scenarios spanning multiple domains as well, choosing the evaluation agent as the behavior, each behavior of which has a separately defined scoring function. All of this can be automated by using the LLM generator without prior knowledge of the algorithmic structure of the solution and knowledge of all possible inputs to all possible behaviors.
If that all sounds very familiar, it's because it's essentially just doing test-driven development, but with the LLM machine as the developer.
It would be expensive to run an entire project development that way, due to the high costs of executing LLMs and/or querying LLM apis, but if the cost is less than that of employing a junior developer or teams of developers, it might be worth it. And human intervention can be kept in the loop at any stage, making it very feasible for rapid prototyping - where you don't necessarily need the entirely correct answer from the machine, just a good enough starting point to allow the human to take over to produce a result.
Reminder, you started with this:
> You can interact with the LLM using agents, feedback loops, and some scoring function on randomized test inputs
Maybe those are fancy words for simple things which I’m not familiar with or learning programming language seems a simpler option :)
For the example, causes of the Great War, I'd first ask it to provide a list of the top ten reasons reputable historians cite as primary causes of World War One. Then, for each reason that looked interesting, I'd ask the LLM to provide arguments for and against, including one primary historical document in support of each side. Then, go to Google Scholar or similar and look up that document and see what different reputable historians have said about it. A few iterations of this process should get you to some reliable information even in a field you know nothing about.
This is a time-consuming process, and requires active engagement rather than what many people have been trained to do by our atrocious education system and watching television, i.e. passive absorption of content without skepticism.
And right here you could bump into problems, because how would you know if what you got back is really the opinion of reputable historians? For all you know it could be giving you the opinions of a layman on Reddit.
> Then, go to Google Scholar or similar and look up that document and see what different reputable historians have said about it.
Looks like that could’ve been the first step. Identify the reputable historians and go from there.
Now there might be outliers that are missing from both lists but which might be worth looking at, as it's not that unusual for consensus viewpoints to be overturned by new information and so on, and that's not something LLMs are likely to be very helpful with.
Fundamentally, my point is that if skeptical thinking, rational analysis, research strategies etc. aren't taught to people from a young age than it really doesn't matter whether the BS they're subjected to is human-generated or AI-generated.
Many philosophers argue that knowing whether something which appears to be conscious really is conscious is a meaningless question. If there is no way to distinguish consciousness from non-consciousness from the outside, then you may as well just say that the thing is conscious.
I think we may be approaching the same question with regard to truth. If it appears to be truth from the outside, is that the same thing as being truth? Is there any point in knowing wether it was produced from a thinking mind or not?
If you haven't seen Gettier problems before, I'd start with the generalized problem and work out from there: https://en.wikipedia.org/wiki/Gettier_problem#The_generalize...
Most examples I have seen are about basic code that is highly likely to be already present in the training data - making the language model an expert system, but with a worse rate of failure.
Meanwhile it has a hard time creating something it hasn't already seen, or a close variant.
How much time have you spent with the strongest models, such as GPT-4?
I've been using it on a daily basis for 6+ months now and your statement there doesn't fit my experience at all.
Just yesterday I pasted in a hastily written function and told GPT-4 to "please evaluate this function and make sure it's bulletproof". It outputted a bulleted list of criticisms (mostly true) and then spit out a version of the function that was shorter, easier to understand, and was safer.
That's not just the "next most likely token". Or if it is, then that's all any of us are doing.
Converting human language to code is an ideal task for LLMs, which is part of why they are such an exciting development.
The trick to using it effectively is to figure out the ways of prompting it that produce GOOD results.
At this point I'm pretty confident it helps them. They need to understand that it's not infallible, but I think most people figure that out pretty quickly after using it for more than a few days.
There's absolutely a skill to using these things well, and it's a difficult one to teach because it's based more on intuition that you develop over time.
The key thing is helping people understand the strengths and limitations of these tools. That's something I do think we can teach.
Might as well just learn to code then.
I used to be smug about getting good results from GPT while coworkers failed to get it to produce anything useful, but over time my luck changed, and GPT stopped nailing it. If you remember there was a bunch of conspiracy talk that GPT got lobotomized, was no longer writing full programs anymore etc, but what really happened is people who got used to rolling 7s started rolling 2s and 3s and shook their fists in anger at their change of luck.
It does not hand recursive thinking well at all.
For a long time there have been "machine learning" techniques where getting the right answer 80% of the time is a win, what's interesting about LLMs is they sometimes "raise the bar" for the machine learning problems (get it right 90% of the time) but have also lowered the bar for other things that can be done algorithmically 100% of the time.
It’s a powerful tool but like any powerful tool, you will use it wrongly if you don’t acknowledge its strengths and weaknesses.
Judging a LLM because he can’t act as a program is a fundamental error because it means you are wasting an enormous amount of resources for nothing.
The novelty of LLMs and where they are good at is helping you thinking about things. Helping you to complete reasonings about your thoughts or about your data. They are not good at doing things but they are good at « feeling » things.
If a ML algorithm with a 90% success rate and a "sanity checker" might be much easier to implement than a ML algorithm with a 100% success rate.
One thing to add - we should not be comparing LLMs to experts in any field (where expertise can be objectively determined). We should be comparing LLMs to an average (in some cases better than average) layperson. The same level skepticism is warranted but it does not mean that they are not useful.
It is a language model, so it’s best use is language things (emails, motivational letter, summaries)
For everything else, you need to be enough of an expert to validate the output.
Hawing a noise generator sometimes generate something that looks like a valid signal doesn’t mean it’s not a noise generator.
My new thing: Testing LLM critique against humans.
The results are sobering.
In my experience, there is a wild amount of finding absolutely nothing wrong with what is clearly going wrong and me having to prompt some variation of "so... what were you planning to do about that thing?" for something to happen.
You've proven my point by replying to my comment.
Does this happen often for you?
I've logged my success rate with Go and C.
32/37 responses failed to meet the requested criteria. 15 wouldn't even compile.
Getting the bot to do it right takes longer than writing it myself.
Maybe that's just me?
I wouldn't think so. Word prediction and writing code aren't nearly the same thing, when I think about them. Sure, they both require language. But one of those also requires reasoning. As worded in this paragraph, could an LLM tell me which one? Sure. Sometimes...
So they run the code, it breaks, they paste the error message back into chatGPT and try again. What's wrong with that? Don't know about you, but I don't get code right the first time, either.
SQL injection is bad code that "works". Most security issues come from bad code that just "works". I can't believe I have to say this, on HN of all places, but you have to actually know what you are doing when writing code. Some of you scare the fucking shit out of me.
Hah, sure. Okay. How? You just read it and think really hard?
Bertrand Meyer, the inventor of Eiffel, who rants about software correctness all the time -- didn't notice an error in a 1-line expression in Eiffel that was generated by ChatGPT [0].
If you spend a little time formalizing correctness in a theorem prover or model checker you may not be so confident that you can read a snippet of code and know that it's correct. In order to know that you have to be able to write the specification precisely. Natural language is not precise enough.
Update: You may be able to write the specification precisely either by modelling the system and checking the model or writing a theorem and proving it... but so far LLM's cannot reason on that level.
[0] https://buttondown.email/hillelwayne/archive/programming-ais...
+1, though it took some time to figure out what questions work well. Yesterday, I asked chatGPT to help me document some R functions, and in my opinion, it did a great job [0]. I asked it to summarize what my functions were doing, and it gave me nice, plain-language summaries, and then reformatted my notes into the the form that Roxygen2 [1] expects, and added some useful inline comments. Moreover, it was fast. It read and understood my code in a few seconds. No human can compete with that.
I don't think this is bullshit. I think that chatGPT shows expertise with some formal conventions that can be kind of a pain to memorize and work with. You might think those conventions were BS in the first place, but regardless, they're what we settled on, and it's really nice to have a coding assistant do the tedious work.
I don't think chatGPT could have written the functions in the first place. But who knows what GPT [5,6...N] will be capable of
[0] https://github.com/setgree/sv-meta/commit/5f71e7c251b38e1981...
[1] https://cran.r-project.org/web/packages/roxygen2/vignettes/r...
Here's Nvidia showing they can use LLMs to teach robots complex skills - https://eureka-research.github.io/
lots of game changing technology starts off being used for toy or bullshit use cases, the GPUs being used to train AI were originally designed for allowing video games to have better graphics
What's stopping any source from doing this? In the case of encyclopedias or news websites, I guess you could say reputation. But that's hardly reliable either.
So I guess that's my big pushback in regards to correctness complaints:
1) "Official" sources have always lied, or at least bent the truth, or at least pushed their subjective viewpoint. It has always been the readers job to think critically and cross-reference. On this point, LLM don't fundementally change anything, they just bring this long-running tension closer to our attention.
2) "LLM are inaccurate". Ok, inaccurate compared to what? Academic journals (see the reproducibility crisis), encylopedias (they tend to be accurate by nature of leaving out contentious facts), journalists (big laugh from me)?
I think outright hallucinations are a valid concern to bring up, but I would refer to my point about cross-referencing. However, I'd still point out that oftentimes a person might have read a perfectly factually accurate encyclopedia article but remembered hallucinations via motivated reasoning. Is the end result (what the person thinks and remembers) really different between LLM and traditional sources? This seems like more of a human problem than a LLM problem.
Do we have any factual evidance showing that people learning via LLM+traditional methods are actually less informed than people who learn from traditional methods alone? Right now, there's a lot of fear mongering and charged rhetoric, and not a lot of facts and studies. The burden of proof needs to fall on the people who want to regulate and restrict access to these models.
Finally, what I think this is really about is the continued transfer from the "industrial age" to the "information age". 100 years ago, your average person wasn't expected or really even able to "do their own research", and instead relied on top-down elite driven institutions to disseminate their version of the information. De-industrialization, then the internet, then social media, and now LLM are cracking this (now) outdated social order, and our old elites are understandably threatened.
I think this is another reason why we need to make this debate more rigorous and fact based: are LLM actually dangerous or is this just an example of elite preference?
I think the major difference that makes in not just "elite preference" is that the LLM output is just confidently wrong. If I read a report from an elite institution I'm at least reasonably confident that they're not going to base their entire argument on the premise that 2+2=5. Someone is going to call them on it, there may be reputational damage, there may even be legal repercussions in certain cases.
LLMs have no such protections. GPT recently, when asked to "write a function using the ELixir programming language that..." wrote a function in Elixir syntax using Python libraries and function names. That's a class of error that makes it actually dangerous if you're asking it about anything you can't fully check on your own, and there's no checks-and-balances for the content it generates that's only ever visible to a single user.
I agree with you that "official" sources can't be trusted and there are inaccuracies everywhere, but you have to admit that the craziest sources who say the craziest things (flat-earthers and such for example) get pretty easily dismissed and everything else they say becomes suspect. LLM's bypass that and can say the most outrageous things without ever getting caught or being forced to make a retraction/correction. That feels like a problem worth worrying about.
My guess is that if you created a GPT of a well-regarded book on WW1, you’d be able to have an enlightening conversation about the topic.
The real horror is when there's going to be an entire generation that grew up not by relying on experts/historic facts and reality, but on something that AI generated that looks good enough and is a neat answer.
Imagine this generation in positions of power over actual experts, or equating AI opinion with the opinion of an older-generation expert.
Progress in all of human history has always hinged on our ability to disregard idiots in favor of experts. Now, the position of the idiot is stronger than ever, because anyone can type a prompt into GPT without real understanding.
That's the real scary part.
Ever seen mainstream media write something about your field of expertise that wasn't complete horseshit? No? Then why believe anything else they say.
> AI systems like ChatGPT are trained with text from Twitter, Facebook, Reddit, and other huge archives of bullshit, alongside plenty of actual facts (including Wikipedia and text ripped off from professional writers). But there is no algorithm in ChatGPT to check which parts are true. The output is literally bullshit.
Well. Any philosopher, mathematician, scientist, or novelist was trained on plenty of actual facts (established works of other people, school learning, experiments) and huge archives of bullshit (individual day-to-day experiences that cannot be called scientific, casual conversations just like Twitter conversations, propaganda, lies). Their output then would also be bullshit, and by extension, everything humanity ever created.
I think you answered your own question there. It doesn’t say your question makes no sense, or that’s incorrect. It just makes up a bullshit answer.
Most of negative sentiment is not “world would be better without chatgpt“, but “one must be very cautious and always validate the answer”.
It’s valuable tool, but it has a LOT of hype, which led to things like people trusting chatgpt output in a court case.
Where I found ChatGPT to be really helpful is when I don't know anything regarding something, and ChatGPT can give me something I could at least verify as true or false.
Compared to before, where I just got stuck on some things, as I couldn't find out what I was actually looking for at all.
That said, there are many people who don't bother to verify anything, as can be observed in comments sections where one can in many cases find people who clearly only read the headline and not the article.
Anyone who uses ChatGPT for a while will learn to take everything with a grain of salt.
(n.b. I am not trying to make any type of political judgement here, I am just using the example in TFA).
If I ask ChatGPT "why is fluoride bad?" it doesn't give me hours of conspiracy theory content^. It doesn't try to sell me water.
Google is not so aligned towards neutral. (Of course, people have been busy sabotaging the Google dataset for over a decade)
^ although I did ask it "what do I do about a ghost that lives in my house" and got this: > Dealing with a ghost can be unsettling. You could consider consulting a paranormal expert or a spiritual advisor for guidance on how to address the presence of a ghost in your house.
* Ignoring it. Not because ghosts aren’t real, but because “some people choose to cohabit and not engage unless it becomes intrusive”.
* Politely ask the ghost to stop disturbing you.
* Purify your home by burning herbs.
* Hire a medium or psychic.
* Get a spiritual leader “from your respective religion” to bless your home.
* Seek help from paranormal investigators.
If I make the exact same question but give it the system prompt “You are James Randi, the world-famous skeptic”, it gives a reasonable answer to help identify the true cause of whatever is making you think there is a ghost.
Which just goes to show how much of a bullshit generator this is, as you can get it to align with whatever preconceived notions you—or, more importantly, the people who own the tool—have.
Then it suggested, amongst other things, ghost expelling rituals, seeking professional psychological help, and moving out of the house.
Using the James Randi prompt, the answer is comical. It says “that would be quite a twist” since he had spent his life disproving supernatural claims.
Context matters - if you begin with saying that you're 6 years old, it will not be willing to admit that Santa isn't real.
This, to me, seems to be an amusing reversal of trends. Before the internet there was the "mainstream media". NBC, CBS, ABC run by the big conglomerates such as GE. Which would mold public consensus in the US. Now there is a desire to go back to that. Let LLM do the thinking and the managing of biases and just tell me what to think.
The unfiltered internet is too much. It's overwhelming. It's the raw sewage of the Facebook feed. We need someone to coddle us. And that champion, today, is OpenAI.
Studying political polarization in the US, the end of the Cold War, globalization, the end of the fairness doctrine, the evolution of major media ownership over time, social media algorithms, echo chambers and confirmation bias is left as a starting exercise for the reader.
Why is verification part of the process if not because ChatGPT generates bullshit? I don’t think the author is disregarding that it is sometimes correct.
What authority does random information on the internet have? How do I verify the information without looking at random information on the internet?
People don’t do that now and we just made it easier to get confident answers that no one will check.
lol. lmao
Second, the fundamental issue here is that as long as verification is needed, this technology is only ever useful for domains in which you're already an expert. That severely limits the technology, and yet this keeps being touted as some kind of general knowledge base when clearly it's at best a very narrow one.
And it seems to me like this is unlikely to be solved in the current paradigm. This is just intrinsic to the architecture of LLMs. There is no getting away from it.
So I guess I am overestimating the average user by assuming that people read the text written on every single page they use to query the system.
Even so, this entire comments section reads like a giant saltmine of people that are terrified of embracing tech that in every sense of the word increases turnaround time on productivity.
The ability to literally feed a screenshot into it and ask questions regarding it, with insane results is astounding.
I have input a picture of baby rabbits and asked it to rate my chickens, and got a detailed explanation of how the picture contains baby rabbits and not chickens.
There are things the system is great at, and others it sucks at currently. But you are all mistaken if you think its going to stay this bad. If you haven't even bothered to check how quickly the field has moved from single-modal to multi-modal, well, damn guys, I am sorry, but you are getting automated first.
The biggest problem with current models is that they don't know how to verify their own answers, and there's no way to allow them to without putting a ministry of truth on it (similar to what twitter was during covid).
To combat this, we need to figure out how to let the model generalize truth, which is something it can only do if fed enough (unfiltered, unbiased) data. With enough data, it will find the patterns underlying false information and eventually figure out how to use it.
In fact, most likely, the future models will simply be fed unlabeled data, in any format (audio, video, telemetry, analytics) and generalize from that.
The concept of feeding normalized data to a model for training is archaic and stupid. The world is not normalized. Data is not always the same, and will not always even be guaranteed to be present.
Anyways, what does it matter.
Why am I trying to sell sugar to people addicted to salt?
Yes, you definitely are. Speaking from experience of supporting users for many years. People will not only ignore such text, they will repeatedly make mistakes or ask questions the text answers directly, even after you’ve explained it five different ways.
The rest of your post is full of assumptions and dismissals, so I think I may be wasting my time, but let’s please stop with the rhetoric that someone who doesn’t agree with you is somehow afraid or has an agenda. As way of example, your comment ignores the bigger picture of bad actors using LLMs to push propaganda, sow discord, or scam people. The “bullshit” part of the problem isn’t limited to “it gave me a wrong answer on my homework”.
Proof that LLMs aren't the only bullshit generators, humans promoting LLMs are too!
No - not everyone using ChatGPT knows that verification is part of the process.
In fact I'd wager that the majority of people using ChatGPT aren't doing any verification on the output at all.
No, everyone very much does not know that. Including the lawyers who though it made sense to ask ChatGPT questions, than asking it if the answers were correct.
https://www.washingtonpost.com/technology/2023/11/16/chatgpt...
On the other hand:
"Things that try to look like things often do look more like things than things. Well-known fact," said Granny. "But I don’t hold with encouraging it." - Terry Pratchett, Wyrd Sisters
This makes LLMs much better than claimed by those who dismiss it as BS or stochastic parrot or whatever, and also not as good as domain experts.
It also seems odd that you had the materials available to research and validate the response ChatGPT gave you, but somehow no ability to glean the necessary answer from those materials to begin with. Obviously, because Google was completely useless to you, you didn't simply use Google to do that validation, nor could you have used another search engine, because then you could have found the result you were looking for to begin with. Did you just use ChatGPT to validate its own responses?
Exactly. You answered your own question and that is the author's point.
Do you even know why it gave the incorrect answer at first? Not even ChatGPT can explain to you why it did that or keeps doing this.
How can you trust it to give the right answer (without you Googling to check) if you really don't know the answer? It already means you don't trust it an you know it bullshits as the author has claimed.
It can confidently convince someone outside of one's own expertise that it can give a soundly correct answer but can easily be bullshitting nonsense.
> Well. Any philosopher, mathematician, scientist, or novelist was trained on plenty of actual facts (established works of other people, school learning, experiments) and huge archives of bullshit (individual day-to-day experiences that cannot be called scientific, casual conversations just like Twitter conversations, propaganda, lies). Their output then would also be bullshit, and by extension, everything humanity ever created.
The difference is with humans, they can be transparently held to account on whatever they say. An AI system cannot. So this actual whataboutism towards humans is incredibly weak here and a common excuse by AI proponents supporting the nonsense that AI systems like ChatGPT can uncontrollably generate in a black box system.
It's a little surprising how hard it is to convince people that it's worth $20 to use v4.0. It is very apparent how much better it is.
This is true, yet I wonder how much of this "ChatGPT is way better than Google at Googling" effect is due simply to how bad Google has gotten over the last decade.
ChatGPT and similar models seem to be uncontroversially best at this kind of task -- just being a replacement for Google Search.
It's not extremely convincing.
Humans don’t learn the same way AIs are trained. Equating the two is a fallacy.
It definitely took you more than 10 seconds, though I guess a bit of hyperbole does not hurt.
You validated findings against other source, which is obligatory step with chat gpt.
Can you share which obscure 18th century historical fact you were looking for? Maybe you happen to be exceptionally bad at google queries?
Would you mind telling what's the fact you were searching for? (and how did you search for it?)
So, they have those two modes of operation, and most of them are built in a way that won't let you know which mode they are working on. That can be useful depending on what you want to do, but it doesn't strike me as the most effective architecture for them.
(They are also reasonably good pattern fitters, what is kinda nice for code autocompletion on some languages. But they are way too inefficient here. Yet, nobody created any good high-efficiency autocomplete that merges searching and pattern fitting, so people use them.)
Google vs cGPT it's neither here nor there in this case, after all Google research drives cGPT, their market focus differs.
The OP is directly mapping Frankfurt's depiction of Bullshit atop the current AI phenomena, mainly because how generative AI approaches "truth".
Natural languages (NL) differ greatly from computer languages (CL), especially on grammar (NL tolerates vagueness) and semantics (meaning, truth). Loosely-speaking, when AI generates NL text it relies on LLMs for context, structure, length etc, a kinda "the model suggests so". However, where content is computer code, even computers can test & verify the output. With NL content on the other hand, the end-user has to essentially test & verify the output to establish truth, meaning.
In my view, the main difference between a Human Bullshitter and a Automated Bullshitter, is that AI can potentially improve exponentially (as smart tech does), thus leaving the end-user (average person) severely overwhelmed.
One is analogous to how LLMs operate and is essentially next token prediction. Think of your mental state when you’re very engrossed in conversation. It’s not a truly conscious act. Sometimes you even surprise yourself. If you’re a child, or childish, you may just confabulate things in the moment, maybe without realizing that you’re doing it in the moment.
Now think of some very difficult problem you’ve had to solve. It’s not the same, right? It’s a very conscious act, directing your focus here and there, trying to reason on how everything fits together and what you might change to fix your problem. Odds are good that you’re not even using language to model the problem in your head.
LLMs are doing the first thing, and are exceptionally good at it even in this early stage. The surprising thing to me is how far this can get you. If you have an inhuman level of knowledge to work from then in conversation mode you can actually solve some moderately difficult problems.
I think that maps to our own experiences as well. For the things that you have deep knowledge on you will sometimes find yourself solving a problem just by constructing sentences.
There was that study a while ago that triggered the NPC meme. The common interpretation was that most people only think in the first way you described.
https://hurlburt.faculty.unlv.edu/heavey-hurlburt-2008.pdf
If so, the interpretation you describe seems pretty far off base.
Also, my impression is that that meme was not sparked by a psychological study, but that a few people did draw on this study to justify it.
And I thought what kicked it off was from only a few years before at most.
Haha, from my own experience, this actually rings true for some “simple” people I know.
knowledge structures as we construct it are symbolic, with symbols representing abstractions (ie classes), along with relations between these symbols. human ingenuity consists of coming up with new symbols or new relations between existing symbols (which is a process of abduction) based on new perceptual inputs (either our senses or instruments). Such knowledge structures are powerful because, they allow us to build giant towers based on solid foundations.
> For the things that you have deep knowledge on you > will sometimes find yourself solving a problem > just by constructing sentences.
This is another way of saying that you have clarity in that subject, and so your stream of thought aligns with knowledge structures - which is another to say that you really understand something. However, in my experience very few people are able to stay within their lanes (competence), and most of us tend to babble on topics we really dont have knowledge structures for. also, few people have the self-awareness of what they really have knowledge structures for (ie, know what they dont know).
That's what rubber duck debugging is, no?
LLMs are rubber ducks that can talk back, for better or worse.
I feel like we should have all moved on from the fact that LLMs don't have a very precise memory? This is (not entirely, but mostly) fixed by including the information you want to query in the prompt itself, like Bing does. ChatGPT is much better at summarization than precise fact recall.
And most of the time, it doesn't matter. When I'm using ChatGPT to learn or troubleshoot something, it's a jumping off point, not the end of the story. As long as everyone understands that, there's no problem. I have yet to see anyone who treats its output as 100% fact, and no AI company claims it either.
But everyone doesn’t understand that, so there is a problem.
> I have yet to see anyone who treats its output as 100% fact
Yet they exist.
https://www.nytimes.com/2023/06/08/nyregion/lawyer-chatgpt-s...
Not every human being on Earth knows what ChatGPT is, so that’s a non-starter.
> or everyone as in most reasonable people
That qualifier is doing some heavy lifting. It’s so vague it can be used to excuse anything¹. Not only reasonable people (however we’re going to define that) have access to the tool.
> ChatGPT tells you it can lie before you can even ask it anything.
And what you learn fast by doing user support of any tech tool is that most users don’t read instructions even if it’s shoved in their faces. I personally know people that think ChatGPT is the source of truth, despite that disclaimer.
In the same way that with a calculator, if you put bullshit in you will get bullshit out. In many cases you need to INPUT facts. This is automated with things like RAG.
I think this is why ChatGPT is extending deeply into RAG, because it matches more closely how users are expecting the tool to work.
Mimicking reasoning abilities is not the same thing as actually having/developing them.
That’s like saying “Oh, you didn’t get the right answer to that math question, you just mimicked giving me the right answer”. Yeah, but if I gave you the right answer, then I gave you the right answer, and I must’ve gotten it somehow.
So far, the yi model is actually not bad, and by two weeks from now two new models will be out that are better. So it will continue until Altman figures out which politicians to bribe the hardest.
Given that the gold standard for rubberducking neither knows nor communicates anything, this is an easy win.
the humans I work with were taught by people that spoke facts backed up by empirical evidence
meanwhile chatgpt was trained on reddit and twitter
How you deal with that is up to you. You can either try to work with what it gives you, setup tooling to easily verify things, and increase your productivity, or you can dig your head down in the sand and say it's all incorrect and you want nothing with it.
Both are valid approaches :)
So what I don't understand is how "drunk driving ... helps people to get to work on time" is close to true and what I said that made it seem like "impossible to say if it's good or bad" about ChatGPT?
In the context of the bullshit economy (e.g., the bureaucrat who writes reports that nobody [reads] because he is paid to write reports that nobody reads because this inflates his manager's headcount and therefore inflates his manager's salary) this is all that needs to be said.
It's tangential to uses of ChatGPT by someone who examines the output and selects what is useful via some kind of expert judgement, or who refines his use of the tool via API calls rather than naive text prompts to the default GUI.
AIs don't "sometimes produce bullshit", they are always producing bullshit. Sometimes what they produce is accurate, as a consequence of the initial corpus being accurate. But to the AI, much like to the demagogue, this is totally irrelevant - it only has to sound good.
And if it's "easy" to verify what AIs say, then it should be "easy" to teach AI to verify before it speaks, which would automatically counter this criticism of AI as bullshit machines. But the real threat is that AI can produce bullshit which is hard to verify, and can do so at a speed that puts professional gishgallopers to shame.
The Cambridge Dictionary says it has two definitions[1]:
1. a rude word for complete nonsense or something that is not true
2. a rude word meaning to try to persuade someone or make them admire you by saying things that are not true
I don't think thise align with your definition and your idea that AIs always produce bullshit. They don't always produce nonsense or something that isn't true, and they don't intentionally use untrue statements to try and persuade someone or gain admiration.
[1] https://dictionary.cambridge.org/dictionary/english/bullshit
Under the above definition, AIs do always produce bullshit, because their content is always truth-irrelevant.
Sure, if that's the only thing you want the LLM model to be able to do, you could do that. But there are other use cases beyond getting facts, that these models usually also want to support.
On one axis we have the truth value of a statement, on the other, the intent (sadly no tables on HN)
Good intent
- Factually truthful: 'Roses are red'
- No inherent truth Value: Maybe surreal humor, like 'this statement is humor'
- Factually false: Satire, parody, sarcasm, 'Great HN comment Einstein'
Bullshit: Statements for effect, indifferent to truth value
- True: 'More doctors smoke Camels'; 'My response to the Covid pandemic was amazing!';
- No truth value - ill-formed nonsense, oxymoron, non sequitur or empty puffery: 'Our Founding Fathers under divine guidance created a new beginning for mankind.'; 'Starbucks provides an immersive ultra-premium coffee-forward experience.'; 'Nobody’s bigger or better at the military than I am.'
- False: 'Obama was born in Kenya'; 'We had the biggest inauguration crowd ever.'
Bad intent, to defraud or mislead:
- True: 'I did not have sexual relations with that woman, Miss Lewinsky'
- No truth value: 'We’re not going to sit here and listen to you bad-mouth the United States of America!'
- False: 'Clinically proven to boost genes and make your skin visibly younger in just a week'; Gish Gallop
Of course this is mostly bullshit:
- Hard to discern intent when something turns out to be false, was Bush's WMD claim bullshit, or a lie; similarly Obama, 'if you like your insurance, you can keep it'. One man's benign public health simplification about masks or food pyramids is another's nefarious conspiracy.
- Even hard science isn't free of bullshit, Feynman's cargo cult speech notwithstanding, any sufficiently complex science and engineering devolves to some amount of cargo cult bullshit in practice
- Let's not even get started on organized religion, Washington and apple trees and whatnot. Communication takes place on multiple levels, fruit flies like a banana, something can be bullshit on one level and good (or bad) intent on another.
In the words of George Carlin, "Bullshit is the glue that binds us as a nation." Or Napoleon, "History is a set of lies agreed upon."
GPT's ability to bullshit at a human level is a truly monumental achievement.
I wrote a long-form version of this bullshit here - https://druce.ai/2023/09/bullshit
Code that runs and works (as in, does what it's asked to do) isn't "bullshit" at all.
Neither is search with references to custom sources in a specific corpus ("RAG").
So yes, AI is in many ways a bullshit generator (as are we); but it can be used in useful and non-bullshit ways.
There’s a pretty big gulf between functional and correct in a world where optimization matters. This is not always the case (strictly, i.e. algorithmically speaking), but you should be optimizing for things like reusability, testability, etc. which GPT-generated code can miss.
In French, we have a word that I find more accurate, which is "baratineur", which means "someone who tells you what you want to hear, regardless of the concepts of truth or lies"
everything becomes pretty clear then.
should you trust the answers it gave you? obviously not, it just generated some crap.
can it be useful for brainstorming? sure, you an consider the output and take it / be inspired by it / etc
can you use it to generate useful code? of course, as long as you read it and it's reviewed by tooling and humans to see that it does what you thought it did.
should you replace your employees with it? obviously not, a text suffix generator isn't able to deal with complex problems nor be responsible for decisions.
etc etc
The real question is where did our education system fail since so many people lack critical thinking and how did the quality of services fall to a level where a bullshit generator sound like a plausible replacement of some workers.
I found that deeply troubling.
But it makes you wonder if future LLMs aren't going to suffer from the JPEG of a JPEG of a JPEG effect where it will be impossible to train an LLM off of new data that isn't generated by another LLM.
(With apologies to Einstein)
However, there are two things: 1) once you can generate arbitrarily parameterized bullshit (a piratical sonnet telling you to drink Coke), you can then filter on truth. This isn't always easy, but this talk presents one way to validate that it is, in fact, quoting the material it's supposed to https://youtu.be/yj-wSRJwrrc?si=MiZHn1xjFNv1nEGv ^^
2) we should care about what people do with this. We should want people to use it for good, and we can shame people who don't. For outright evil we can make it illegal when existing laws don't suffice (and we should)
3) text is an interesting medium. It's universal in a lot of ways. There's a cynical take at Google that you're just converting one proto to another, and I suspect some white collar workers think they're just taking some text and reformatting it. I don't dispute that some jobs are useless, but I think that "just writing emails" or "just modifying text" isn't completely useless, and something interesting will happen if we can improve productivity in these areas
^^ you can just do a string contains on the original text vs the quotes it gives you, so this isn't fancy
• Is it kind?
• Is it necessary?
You are underestimating the parameter space of bullshit. By a lot. Thinking about orders of magnitude doesn't do justice to how much you are underestimating it. It's one of those things that can not fit within the physical universe.
An example: I asked it for the way to move an AWS account between two OUs in Terraform. It promptly generated the code including a reference to the "aws_organizations_move_account" resource. Since it looked suspicious to me, I checked it and the only place it exists is a Github Issues page where the code is a proposal for something that could work somewhere in the future (it doesn't, you move accounts in a different way).
So if the people who trained the model decided that instead of making it general and versatile they will only focus on higher quality content, and made a manual selection if titles (most academic books, especially from respected publishers; public source code only from prominent projects with and adequately high number of developers and so on), they could produce a model that would be less eloquent but could generate less bullshit.
We already know that specialist systems, such as those tried in the 80's are not very good in modeling reality, either. The thing is, using LLMs as knowledge bases or as a sort of modern oracles seems to me like using a powerful tool for something they are horrible at.
They are good enough (but not perfect) text summarizers, text translators and text transformers, but are far from the so expected "AGI" some seem to be expecting them to be.
But that doesn't mean you need to be "anti-AI". I'm also also super excited about LLMs and their potential and am actively working in the field.
It's important to understand their role. LLMs aren't the whole brain: they're just the linguistic cortex. That's something computers couldn't do before! It's pretty amazing that we now have a computers that can not only use language, but are pretty good at it! Incredible breakthrough.
The problem comes when you let the linguistic brain babble without integration with a more structured reasoning system or cite-able fact database.
That's exactly what anyone who is working on retrieval and tools for LLMs is doing, so we're not on the wrong path. But a lot of people still treat this language model as if it were a whole, reasoning brain, and it's really really not.
(disclaimer: my comparisons to human mind/brain/intelligence are for analogy only. Real brains and transformer models are very different things.)
So long as the cyclomatic complexity of what you're asking it to write is less than 8 it does a good job of it. I've yet to have any trouble with it writing test functions. And it does a better job of explaining code than most developers I've worked with. All those things together make it unbelievably useful when you start working with the tool rather than against it.
It's increased my capabilities to the point that I can tackle problems in the hour before bed that used to take me weeks before. More importantly for me it's made it clear just how muddled my thinking between code and architecture was.
LLMs are impressive because they can generate text, images, etc. from a prompt, but hallucinations and other tells show that they do not really understand.
Older AI systems can do basic levels of reasoning, but did not have much of a natural language interface like today's LLMs do. Are there any efforts to tie the two together?
Do you happen to make mistakes? No, I always get it right. If you can't make mistakes, how can you tell when you get it right?
All I expect from my assistant is, at least, for it to make less mistakes than me.
Also the paper about "stochastic parrots" is under great criticism and there are now papers that contradict their conclusions.
ChatGPT Is a Bullshit Generator Waging Class War
https://www.vice.com/en/article/akex34/chatgpt-is-a-bullshit...
That allegedly well-informed commentators can infer that ChatGPT will be used for "cutting staff workloads" rather than for further staff cuts illustrates a general failure to understand AI as a political project. Contemporary AI, as I argue in my book, is an assemblage for automatising administrative violence and amplifying austerity. ChatGPT is a part of a reality distortion field that obscures the underlying extractivism and diverts us into asking the wrong questions and worrying about the wrong things. Instead of expressing wonder, we should be asking whether it's justifiable to burn energy at "eye watering" rates to power the world's largest bullshit machine.
1) Ask what you have built 2) What is the error rate and hallucination rate - ON your production data.
They just don't know how they work, really want to believe they work like a person, or have otherwise been duped by the boards of the ad-ai industrial complex whose spiel is whatever increases market valuation (google, fb, ms, openai, etc.).
I guess if you have absolutely no idea how these systems work there really isn't anything better than analogizing to people -- it is this 'failure of analogical reasoning' which ai/ad corps exploit in their desire to dupe ever more misinformed investors.
This realisation helps me, only, since I can now fathom better when i'm wasting my time talking to someone whose basis for understanding stats is human psychology -- the amount of discussion needed to undo that is possibly a lot.
Incidentally, all statistical systems of this kind (almost all ML, including NNs) work in the same way. They compress data into a smaller representation, and take a weighted average of pieces of that data relevant (by freq association) to the prediction.
There is no reasoning, no truth, no falsehood, no lies, no bullshit, no hallucination, no veridical perception, no interiority, no subjectivity, no capacity to represent the world.
There is something much more like a few pages of a book stuck together, with the second shinning through the first.
What riles me is the grubbiness of the commercial pushers, and the foolishness of the press which repeat their marketing material as-if its science.
People should not rely on computers in the way they rely on people -- and any inducement to doing so is unethical.
I agree with you about anthropomorphizing terms like “hallucinating” and “lies”, which imply belief, intent, or knowledge in a human sense.
“Bullshit”, as used here, is a perfect term to describe LLM output specifically because it indicates that none of those terms and concepts apply.
When I tell you the LLM hallucinated something, you know exactly what I mean. Same for “It’s rationalizing”, “It’s refusing to answer”, “It’s being illogical”.
Your desire to reduce it all down to stats is as a bad as someone arguing that we shouldn’t talk about balls and goals in basketball, because it’s all just atoms.
Glad Mr. Blackwell could fill in these details for me!
But it's a mistake to think AIs are "just" bullshitters. If you ask a GPT to auto-complete some bullshit for you, it will happily do so. But it will do many tasks that provide value to society just as easily.
Nobody knows how society will be transformed by AI. But if AI generated BS is a problem then I wouldn't be surprised to see AI powered skepticism assistance as well. We all have terrible blind spots, and maybe in an AI world it will be harder to delude yourself than in the present day? No predicting how this will turn out.
this is only by chance though, and you have no way of distinguishing it
LLMs still cannot transparently explain any of their decisions at all or reason why it is cannot understand its own mistakes, meaning that it cannot be held to account.
This is why no-one here trusts them at very high risk situations, only for bullshit and nonsense generation.
Unless you can give a thorough explanation as how these LLMs internally can explain themselves transparently and reliably to the point where we don't need to check their outputs?
Quite simply, we are talking about bullshit. Philosopher Harry Frankfurt, in his classic text _On Bullshit_, explains that the bullshitter “does not reject the authority of truth, as the liar does […] He pays no attention to it at all.” This is exactly what senior AI researchers such as Brooks, Bender, Shanahan and Hinton are telling us, when they explain how ChatGPT works. The problem, as Frankfurt explains, is that “[b]y virtue of this, bullshit is a greater enemy of the truth than lies are.” (p. 61). At a time when a public enquiry is reporting the astonishing behaviour of our most senior leaders during the Covid pandemic, the British people wonder how we came to elect such such bullshitters to lead us. But as Frankfurt observes, “Bullshit is unavoidable whenever circumstances require someone to talk without knowing what he is talking about” (p.63)
When it comes to politicians our democratic values allows bullshitters to manifest and if we remove it we remove a key part of democracy.
In ai's case I would say it's a bit the same.
(Btw, if I need to do a code review and verify all work of an intern ir a junior programmer, it doesn't mean their work is useless - it's valuable and a net win)
You don't have a use for a tool like this? Fine, your call. But stop this patronizing arrogant attitude.
> “Godfather of AI” Geoff Hinton, in recent public talks, explains that one of the greatest risks is not that chatbots will become super-intelligent, but that they will generate text that is super-persuasive without being intelligent, in the manner of Donald Trump or Boris Johnson.
> Perhaps this explains why the latest leader of the same government is so impressed by AI, and by billionaires promoting automated bullshit generators?
> Graeber observes that aspects of university education prepare young people to expect little more from life, training them to submit to bureaucratic processes, while writing reams of text that few will ever read. In the kind of education that produces a Boris Johnson
> AI systems like ChatGPT are trained with text from Twitter, Facebook, Reddit, and other huge archives of bullshit, alongside plenty of actual facts (including Wikipedia and text ripped off from professional writers).
> This is the crux of the distinction between [the bullshiter] and the liar. Both he and the liar represent themselves falsely as endeavoring to communicate the truth. The success of each depends upon deceiving us about that. But the fact about himself that the liar hides is that he is attempting to lead us away from a correct apprehension of reality; we are not to know that he wants us to believe something he supposes to be false. The fact about himself that the bullshitter hides, on the other hand, is that the truth-values of his statements are of no central interest to him; what we are not to understand is that his intention is neither to report the truth nor co conceal it. This does not mean that his speech is anarchically impulsive, but that the motive guiding and controlling it is unconcerned with how the things about which he speaks truly are.
> It is impossible for someone to lie unless he thinks he knows the truth. Producing bullshit requires no such conviction. A person who lies is thereby responding to the truth, and he is to that extent respectful of it. When an honest man speaks, he says only what he believes to be true; and for the liar, it is correspondingly indispensable that he considers his statements to be false. For the bullshitter, however, all these bets are off: he is neither on the side of the true nor on the side of the false. His eye is not on the facts at all, as the eyes of the honest man and of the liar are, except insofar as they may be pertinent to his interest in getting away with what he says. He does not care whether the things he says describe reality correctly. He just picks them out, or makes them up, to suit his purpose.
* https://www2.csudh.edu/ccauthen/576f12/frankfurt__harry_-_on...
* https://web.archive.org/web/20150701235021/https://www2.csud...
* https://en.wikipedia.org/wiki/On_Bullshit
* https://press.princeton.edu/books/hardcover/9780691122946/on...
Liar: knows the truth, and says something specifically other. Bullshitter: does not care about what is true or what is false; says whatever necessary to achieve specific goals.
>Timnit Gebru
How strange that the former two are cited as examples of bullshitters, but the latter not. That Twitter and politics is cited as a hotbed of untethered debate, but not academia. Even though the work and claims of "AI Ethicists" pretty much amounts to idea-laundered cultural marxism... and that this problem also exists in the wider social sciences, where certain hypotheses are dismissed regardless of evidence, because of the perceived moral implications.
I have no doubt that much of the output of a transformer is bullshit. The same way I have no doubt that much of what passes for scholarship these days is. Garbage in, garbage out.
Main difference seems to be that ChatGPT will apologize and correct itself when you point out it's wrong.
Who programmed ChatGPT to lie, mislead, and be deceitful about certain topics? The way ChatGPT presents critical health information is borderline criminal.
The magic is not that it can tell you thinks but that it understands you with a very high probability.
It's the perfect interface for expert systems.
It's very good in rewriting texts for me.
It's very good in telling me what a text is about.
And it's easy enough to combine a LLM with expert systems through apis.
I for example mix languages when talking to chatgpt just because it doesn't matter.
And yes it's often right enough and for GitHub copilot for example it doesn't matter at all if it's always right or only 80%.
It only has to be better than not having it and 20 bucks a month.
People get fired every day for pasting code that is 80% right and worked on a couple test-cases.
Though SO has a lot of moderation, so it's somewhat self regulated.
Never seen this happening. In contrary still not every team is doing code review and plenty of people regularly fix bugs in production.
One ex colleague invalidated all apple device certificates, didn't get fired.
A previous tech lead wrote code which deleted customer data and we found that a half year later, no one was fired.
And no one got fired at a code review.
And I clearly showed an example how LLM is more an interface than a answering machine.
If a LLM understands the basics of law it is by sure much better than a lot of paralegals of transforming the info into search queries for a fact database.
And I'm pretty sure there are plenty of mistakes in existing law activities
Also, LLMs don't understand other than via the language representation.
Surely you cannot rely on Wikipedia for knowledge, there’s no algorithm to check the facts! Anybody can write anything!
I imagine in 10 years these BS text generators will be the norm just as wikipedia is now the norm.
In this context, "loop closing" likely refers to a method of addressing or resolving issues or gaps in a process or system. The speaker seems to be expressing frustration about a perceived lack of a method to counter misinformation or falsehoods when it comes to Language Model (LLMs) efficiency."""
So, same inference I made from the context (alone it sounds like citogenesis but the context suggests the opposite), and same lack of specific detail.
"A method". Yes, great, what is it?
Me: why did the commenter use the term “loop-closing”?
ChatGPT: The commenter used the term "loop-closing" to emphasize the concept of a feedback mechanism that corrects or validates information. In systems theory and engineering, a "closed loop" system is one where the output is continually fed back into the system as input for self-correction and optimization. This term is metaphorically applied here to describe a process where information is constantly reviewed and corrected, forming a self-regulating cycle.
In the context of Wikipedia, this loop-closing happens through community edits and oversight, where incorrect information is often quickly corrected by other users. The commenter is highlighting the absence of a similar, effective mechanism in Large Language Models (LLMs) to automatically correct or verify the information they produce. This comparison underscores a concern about the potential for LLMs to disseminate incorrect or misleading information without an inherent system for checking and correcting it. The choice of "loop-closing" as a term is to draw attention to this systematic difference in information validation between platforms like Wikipedia and LLMs.
That's why I used triple quotes around the block: https://chat.openai.com/share/cd5c8e5e-5925-46f8-8873-6e75c5...
As it happens, I just watched some of the OpenAI developer day videos, and they do talk about connecting the models with more traditional search engines/databases/API calls to a similar end.
> What we need to regulate is the bullshit.
Feeling bit threatened, are we? We have complete democratisation of "BS" (marketing, propaganda...). Now every scammer with Nigeria can compete with top universities and think tanks, without being "intelligent".
Something like that happened when Samuel Colt invented his revolver. Every person could defend themselfs. Overnight it become very unwise to randomly attack people! Total democratisation of violence!
"Intelligent" people will have to drop their shilling for latest fad, will have to build some credibility, reputation and personal brand!
Should we ignore your comments because you use throwaway account and thus have no credibility reputation or personal brand here?
What planet do you live on where this happened?
"Every person could attack anyone" seems like a better phrasing, given that LLM really don't allow us to defend ourselves from bullshit.
> "Intelligent" people will have to drop their shilling for latest fad, will have to build some credibility, reputation and personal brand!
I'm not sure I follow. Without LLMs, idiots can convince the masses that they aren't idiots just with some reputation and personal brand, some even got elected for presidents. How is this changing the situation?
Trust me, LLM is great for filtering bullshit. Sentiment analysis, cross checking references, short summaries, ad filtering... You can filter BS on industrial scale with LLM!
By "intelligent" people I meant current elites with platform of universities, news media... Thay have very bad track record of lying and corruption.
I'd call that a projection...
> Sentiment analysis, cross checking references, short summaries, ad filtering...
Out of all these, only checking references is actually a way to filter-out bullshit. Checking against what? A third party you trust... It's rather straightforward for simple facts, much less for complex ones. LLM help to automate some of that, but it doesn't fundamentally change how it's done.
Want to make it a click-away for the ones who don't have the technical skills? Now they have to trust you...