LLMs Don't Know What They Don't Know–and That's a Problem
blog.scottlogic.com
blog.scottlogic.com
I know this is very cynical of me, but I am becoming very convinced that this is actually the biggest draw to AI for a lot of people
They don't want to use their brains
People talk about how it can summarize long text for them. This is framed as being a time saver, but I'm positive for a lot of people they just don't want to read the full text
It can generate images for people who don't want to learn how to draw, or save them money by not hiring artists
They also talk about democratizing art- are they using LLMs' probably vast corpus of art feedback to improve their own work? Well, no.
LLMs generate text, not knowledge. They are great for parsing human culture… but not good at thinking.
Yes, those can be very much the same thing and it does not mean that people don't want to use their brains. If I encounter a long article and have serious doubts as to whether it will be worth my time, reading a summary first helps a lot.
In fact, now that we're on the subject: _Proper_ journalism is _supposed_ to provide the key points at the beginning of the text (as opposed to some silly information free nonsense meant to 'set the scene' or 'hook you in'). Don't bury the lead/lede. See https://en.wikipedia.org/wiki/Inverted_pyramid_(journalism)
Half of the value of reading things via HN is seeing the top comment be a summary of the key points of the linked content. It's very, very valuable in this day and age where good writing has taken a back seat to making money.
Still, drunk , educated uncle Steve is pretty handy sometimes.
The difference is, the hammers and screwdrivers perform a single task, and have been designed and optimised for that specific task.
LLMs are much more versatile and capable of performing a wide range of tasks. Yet, at the same time, their capabilities are ill defined.
> LLMs are much more versatile and capable of performing a wide range of tasks. Yet, at the same time, their capabilities are ill defined.
That’s my point, I want to skip to the part where we know what LLMs are good for, what they are bad for, and just consider them another tool at our disposal. We’re still in the phase of throwing shit at the wall to see what sticks, and it is exhausting more often than not.
GOOD: Language parsing.
BAD: Information retrieval.
We are now seeing the LLM is used to parse the question and retrieve information from elsewhere.
Before you would ask the LLM who the president of the US was and the LLM would autocomplete. Now the LLM constructs a query through a tool and searches the internet for an answer.
It parsed the entire internet to have enough data to learn about language, but you don't necessarily want to depend on what it learned, other than to parse the syntax of the user.
Totally agree with that.
Similarly LLMs are a thing that turned out to be useful and we end up looking foru usecases for it.
Similar to the YC analogy of the company that discovers a brick and they have to find out useful ways to use it: To put out fires, to hit people in the head, etc..
Well they’re most certainly not the right tool to fasten screws with :P
That seems problematic, too.
I wish hallucination benchmarks were far more popular.
I notice it a lot when coding with static typed languages, you paste your code and it will tell you that you were "deceived" very quickly.
It gives extremely confident but wrong answers, I've found it to just be way more convincing.
An LLM might be seen as a kind of very elaborate linguistic hoax (at least as far as knowledge and intelligence are concerned).
And I like LLMs, don't get me wrong. I'm not a hater.
To knowingly advance something false as true, is a hoax.
You know?
i got 3 wrong answers in a row (that i could easily confirm were wrong by compiling)
then the 4th worked. it was much faster than reading the jvm spec about wildcard generic subtyping relation (something ive read before but couldn't quote) and it taught me something i didn't know even though it was wrong
I was thinking more like you don't know that you don't know poop.js because you didn't know about it until I said it just now. Otherwise you'd have continued blissful in your unaware-ness of poop.. AKA you didn't know that you didn't know it.
LLM's aren't an all knowing power, much like ourselves, but we still take the opinions and ideas of others as true to some extent.
If you are using LLM's and taking their outputs as complete truths or working products, then you're not using them correctly to begin with. You need to exercise a degree of professional and technical skepticism with their outputs.
Luckily LLM's are moving into the arena of being able to reason with themselves and test their assumptions before giving us an answer.
LLM's can push me in the wrong direction just as much as an answer to a problem on a forum.
So, A.I. would make an excellent politician or used car salesman then... ;)
In fact, I’m much more surprised at just how capable their are of such a wide range of task, given that they have just ‘learnt from the internet’!
Could have also been the fact that my custom GPT instructions included stuff like “ALWAYS clarify something if you don’t understand. Do not assume!”
Most of the comments on this thread are needlessly pessimistic
"If you say please LLMs think you are a grandma". Well then don't say you are a grandma. At this point we have a rough idea of what these things are, what their limitations are, people are using them to great effect in very different areas, their objective is usually to hack the LLM into doing useful stuff, while the article writers are hacking the LLM into doing stuff that is wrong.
If a group of guys is making applications with an LLM and another dude is making shit applications with the LLM, am I supposed to be surprised at the latter instead of the former? Anyone can do an LLM do weird shit, the skill and area of interest is in the former.