Building a ChatGPT-enhanced Python REPL
isthisit.nz
isthisit.nz
> I think that the LLM thinks that it cannot generate code for the simple task. Unlike every other computer API in existence we can prompt the LLM to tell us why it responded in the way it does.
I agree the LLM is going to produce text that is going to be narratively consistent with the assumptions of the question. But we already know LLMs are giant bullshitters. Is there any reason to think they are doing actual introspection and description of internal state? Or are they just going to give plausible-sounding words?
> I agree the LLM is going to produce text that is going to be narratively consistent with the assumptions of the question. But we already know LLMs are giant bullshitters. Is there any reason to think they are doing actual introspection and description of internal state? Or are they just going to give plausible-sounding words?
This sounds like something we could test by inverting the LLM-provided reason of why it can/cannot do something and playing that back through the model.
The example I gave in the blog was, unprompted, the LLM thought it should not generate code that it thought had already been executed. If we take it at face value, I could add a line in the code generation prompt to instruct it to generate code even if it "thinks" it has already executed it. We then assert the outcome.
Not very scientific, but might give us some insight when trialed across thousands of prompts.
Someone needs to call out the double standard applied to AI.
If machines that bullshit are philosophically indistinguishable from intelligent humans, how is it that humans that bullshit are distinguishable from intelligent humans?
I believe vba616 is trying to point out that if we consider AI systems that
generate plausible-sounding responses to be "bullshitting," then we should
also acknowledge that humans who do the same are not necessarily
distinguishable from "intelligent" humans. In other words, both AI systems
and humans can produce seemingly coherent responses without necessarily
having a deep understanding of the subject matter.
It raises the question of whether we should hold AI systems to the same
standard as humans when it comes to evaluating their output or
understanding. If we accept that humans can sometimes provide plausible yet
shallow answers, should we not also accept that AI systems, like LLMs, can
do the same?We get served bullshit and fiction-presented-as-fact every day from humans. To make matters worse, alot of these lies are not even the result of an error. A lot of them are deliberate, to further some sort of agenda.
Yes, AI can fantasize about facts. Yes AI isn't perfect. But at least, when it does spew misinformation, it is virtually always the result of an error. When an AI lies, it's because of errors, not to sell us something, or get our vote, or talk itself out of something.
> If machines that bullshit are philosophically indistinguishable from intelligent humans, how is it that humans that bullshit are distinguishable from intelligent humans?
Finally someone gets it. A brand new state of the art AI was trained on the entire digital corpus of human knowledge. Said AI then sometimes bullshits people, and people are surprised..?
Why hold AIs to the same standards as humans. We already have humans, and can't really change them. Let's apply standards to AIs that make them useful tools.
Indeed, I think what a lot of people are doing with LLM output is the same thing customers of paranormal scams do. Faced with artfully constructed ambiguity, they fill in the blanks themselves, not even noticing how much they're supplying the work. Some bit go well and some poorly; the successes are proof that the psychic has paranormal powers, and the rest are quickly forgotten about.
If rationalizations work as well as human ones do then treat them as you would human ones.
But you didn't address my point. Do you believe that explanations produced via tarot or zodiac should be treated with that same level of deference? That explanations from ouija boards are things we "might as well go with"?
Well, if we use the ReAct model, the flow of action involves the LLM narrating its reasoning and then acting on it. Which can still be bullshit, of course, but its at least a different thing than a post hoc rationalization when questioned. (Of course, we know that when questioned about past actions, humans often construct post hoc rationalizations, too.)
(This was a great post btw - highly recommend if you’re just perusing the comments)
The reason being: All software engineers are trained to read and write code, aka. precise formal languages. All feedback from systems about errors in that are also precise and formal languages.
What will likely happen, is that we will use anthropomorphic language more and more to generate code, especially code that is easy to describe but just a boring hassle to type (aka. boilerplate or things like simple unit tests).
https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7305066/
(1)
https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7305066/
(1)
The state that lead to their previous generation was GBs of floating point calculations which were discarded as soon as the output had been generated.
If you ask them why they said something they will generate a brand new explanation based on the previous text they generated, but any relationship to their actual decision process will be entirely coincidental.
Update: I just noticed that the article uses this:
> If you cannot return executable python code return set the reason why in the description and return no code.
That DOES work - getting it to generate justification during a single response isn't affected by the lost state issue I'm describing here.
I feel like the answer to the question “give me a possible reason that a person/llm would give this answer given this situation” is all that’s ever really possible!
Each time it just generates one more token until it hits a stop token, using the previous input. It shouldn't matter if it created a token or you did.
I'm not saying you can't call that an explanation if you want to use that word, I'm trying to make a distinction between the generated output and an explanation of that output. All of this is the former. The latter is not possible unless you analyse the model as it is running, and we don't really have a good understanding of that process yet.
There was a cool paper where researchers analysed a finetuned model playing othello, and they found some neurons that clearly represented the game state- if they changed those neuron values, the network responded correctly, even if that game position was impossible to get to normally [1]
I'd argue work like that is getting us closer to getting explanations for llm behaviour.
>It cannot even distinguish between text it generated and text you added unless you tell it that.
Many people cannot distinguish reality from fiction invented by the brain. This doesn't mean much. If i could insert visual sense into you anytime i wanted, you'd struggle to tell it apart at best.
People can't really recreate mental states. It's all post hoc rationalization. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3196841/
and the split brain experiments tell us the brain will happily fabricate those rationalizations to be something entirely unrelated. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7305066/
But as a matter of technical implementation I can tell you: there is no state. You can't simply say "we don't know if theres state" and be done with it. You individually might be unsure. Check the paper "attention is all you need", which introduced transformers, if you'd like to see for yourself [1].
The state is all in the text, that is why I said it's important that llms can't even distinguish between text you wrote and it wrote- how can you talk about an internal state of the llm if I can create that state by my input of text? That's just input at that point, not state irrespective of input. The same would go for your inserted visual sense, I suppose.
Is there something like that yet?
Neptyne spreadsheet has a GPT-enhanced python repl as well: https://neptyne.com/
I think we will see demand for it, but I also think it completely misses the nature of LLMs / for what they are most useful.
Basically trying to shoehorn "old" patterns into a new paradigm.