hallucinates
> who is dan shipper? please search if you don't know the answer.
gives right answer
If it is a reasoning engine, why didn't it answer "I don't know" in the first case?
hallucinates
> who is dan shipper? please search if you don't know the answer.
gives right answer
If it is a reasoning engine, why didn't it answer "I don't know" in the first case?
Week 4:
Lecture 7: Introduction to Command Line Interfaces
Lecture 8: Makefiles and Building C Programs
...
Week 11:
Lecture 21: Network Security and Cryptography
Lecture 22: Network Programming with C and C++
Week 12:
Lecture 23: Introduction to Makefiles and GCC
Lecture 24: Advanced Debugging Techniques with GDB and Valgrind
Week 13:
Lecture 25: Advanced Computer Architecture Concepts
Lecture 26: Advanced Topics in Concurrency and Parallelism
What's interesting here is that it does present a relatively sensible schedule. But look closer and it's completely forgotten all context. It's forgotten that it already introduced these topics: Lecture 8: Makefiles and Building C Programs
and Lecture 23: Introduction to Makefiles and GCC
And it puts them in a completely illogical spot between networking and concurrency. So it doesn't understand anything about the order in which these topics are introduced. It doesn't understand that it shouldn't introduce them twice. All it knows is that the word "Introduction" often appears on class schedules. And so in writing a class schedule, it included this phrase a number of times. But it didn't put it in the right spot. So even though literally no one would ever make this mistake in devising a class schedule, ChatGPT will because this isn't a logical class schedule, it's a stochastic class schedule. Sometimes in a stochastic class, you cover the intro material at the end!I said:
How are the students supposed to compile their C programs before week 12?
It said: You make a good point. Students would need to be able to compile their C programs before week 12. In fact, compiling C programs should be introduced earlier in the course, ideally in the first few weeks along with the introduction to C.
Which, again, makes sense at face value, but the schedule does cover building C programs in the first few weeks. The problem is that it shouldn't be covering it again in week 12.The response I was hoping for was
You make a good point. We introduced compiling C programs in week 4, so doing this topic again in week 12 is redundant. Instead we can cover a different topic that fits nicely, which would be....
You know what ChatGPT reminds me of? Talking to an undergraduate who always nods and agrees that they are following your logic. Then when you ask them to tell you their understanding, it's either an exact parrot of what you said, or wildly off due to some preconceived notion or general lack of understanding of the topic. Either way, they can never let on that they are clueless.> Hello, who do you want information about?
> who is dan shipper?
> Dan Shipper is an entrepreneur, writer, and software developer known for his work in the tech industry. He is the co-founder of Superorganizers, a platform that focuses on productivity and organization. Superorganizers offers a newsletter, podcasts, and interviews featuring insights from successful entrepreneurs, writers, and other professionals on how they stay organized and manage their work.
> Who is Tommy Tonlongeron?
> I'm sorry, but I don't have any information on a person named Tommy Tonlongeron. It is possible that the name is misspelled, the person is not well-known, or there is simply no information available on them within my training data up to September 2021. If you have any other questions or need information about someone else, feel free to ask.
If you tell it (GPT4) to, it seems GPT4 is at least mildly aware about what it knows VS what it doesn't know.
>>> You are a AI for describing people you know about. If you're not 100% certain about the information you give to the user, you should say that you don't know instead of guessing. Greet the user with "Hello, who do you want information about?"
> Hello, who do you want information about? I'll do my best to describe them based on the information I have, but if I'm not certain about something, I'll let you know that I don't know.
>>> who is dan shipper?
> Dan Shipper is an entrepreneur and writer based in New York City. He is the co-founder and CEO of a startup called Pilot, which provides bookkeeping services for startups and small businesses. Dan has also written for a variety of publications, including The New Yorker, The Wall Street Journal, and The New York Times. He is known for his insights on entrepreneurship, technology, and productivity, and has been featured in several media outlets for his work in these areas.
>>> Who is Tommy Tonlongeron?
> I'm sorry, but I couldn't find any information about a person named Tommy Tonlongeron. It's possible that the name is misspelled or that this person is not well-known or doesn't have a significant online presence. If you have any more information or context, I'd be happy to try and help you further.
(In case it isn't clear, Chat-GPT appears to have hallucinated the answer about Dan Shipper, I think. Pilot is a real startup, but I don't see how it's related to Dan Shipper at all, unless I'm missing something.)
>>> Why did you say that Dan Shipper is the co-founder of Pilot?
> I apologize for the error in my previous response. Dan Shipper is actually one of the co-founders of a startup called Firefly, which provides a tool for students and educators to collaborate on course materials. I'm sorry for any confusion that my mistake may have caused. Thank you for bringing it to my attention, and please let me know if you have any other questions.
(This answer was true--until 2014 when Firefly was sold, making the use of present tense misleading at best.)
> Dan Shipper is a technology entrepreneur and writer. He is known for co-founding the startup Firefly, which aimed to improve networking for college students, and for co-founding the startup Monkey Inferno, which incubates and launches consumer internet companies. Shipper has also written for various publications, including The New Yorker, The Paris Review, and The Atlantic.
Monkey Inferno? No, that was Michael & Xochi Birch.
I additionally asked it about its correct answer, using the same question wording as I did previously:
>>> Why did you say that Dan Shipper is a co-founder of Firefly?
> Dan Shipper is known for being one of the co-founders of Firefly, a startup that provides a platform for teachers and students to create and share course materials. Firefly is designed to help teachers and students work together more efficiently, by providing tools for creating assignments, sharing resources, and communicating with each other. Dan Shipper co-founded Firefly along with his friend and fellow entrepreneur, Brett Gibson, while they were both still students at the University of Pennsylvania.
This is also incorrect, as his co-founder at Firefly was named Justin Meltzer, not Brett Gibson.
I did some searching to track all this down, since I knew none of the people involved before this exercise.
This means you were using GPT-3.5-turbo rather than GPT-4 I'm guessing? The quality of the generation is vastly different between the two, especially for steering the AI.
Because it wasn't purposefully built to be a reasoning engine from the outset. So some of us are working out the kinks and minimizing the degree of hallucination. GPT-4 hallucinates less than its predecessor.
Others are writing reasoning loops like ReAct that get us closer to it by refining the model to act more like a reasoning engine.
The fact that it can reason without being told to do much other than autocomplete text is the emergent property that people are in awe of. Maybe language and reasoning are so inexorably linked that you can't have one without the other. We are, after all, the animals on this planet with the most complex form of language. Other relatively "smart" animals like whales also have some rudimentary form of language.
And if you extrapolate the rate of progress in this emergent reasoning from "GPT-1" to "GPT-4", AGI feels within reach
Reasoning engines can be built, but its not in GPT-4.
To reason is to seek truth (probably to survive). Do you think the truth-seeking is emerging? If so, why do you hypothesize it is emerging (i.e. what's the purpose)?
Wikipedia defines reasoning as "Reason is the capacity of consciously applying logic by drawing conclusions from new or existing information, with the aim of seeking the truth". https://en.wikipedia.org/wiki/Reason
Here's a sample question from the GRE Verbal Reasoning test:
> Upon visiting the Middle East in 1850, Gustave Flaubert was so [blank] belly dancing that he wrote, in a letter to his mother, that the dancers alone made his trip worthwhile.
> (A) overwhelmed by
> (B) enamored by
> (C) taken aback by
> (D) beseeched by
> (E) flustered by
Whether that specific question was in the training corpus or not, there are enough words in the sentence to suggest a positive association, including, significantly, "worthwhile." That alone possibly serves to narrow the answers down to A or B, with a preference for B, because it's more likely that "worthwhile" and "letter to x mother" are associated with "enamored" in general English-language text.
Look, the whole point of these models is that its not easy or even possible to trace the path any given input takes on its way to output, but we know the principles used in development, so I think it's rather more of a burden to explain how the clearly-explained principles in LLMs result in something other than the obvious. The fact that the results are so overwhelming that we become enamored by them, well, I'm taken aback by the seeming accuracy of some of the responses, but I beseech you to remember the other responses in which these LLMs are dramatically off-base, as if flustered--if LLMs could ever be flustered.
A tale as old as time: "It is illustrated by the success of chess computers. In the 60s, it was said that computers will never beat people at chess, because that requires intelligence and computers aren't capable of intelligent thought.
When computers regularly started winning matches in the 80s, it was claimed that playing chess wasn't a test of real intelligence because computers could do it."
>Whether that specific question was in the training corpus or not, there are enough words in the sentence to suggest a positive association, including, significantly, "worthwhile." That alone possibly serves to narrow the answers down to A or B, with a preference for B, because it's more likely that "worthwhile" and "letter to x mother" are associated with "enamored" in general English-language text.
Yes, this is called reasoning so in other words the LLM is reasoning about language.
>Look, the whole point of these models is that its not easy or even possible to trace the path any given input takes on its way to output, but we know the principles used in development, so I think it's rather more of a burden to explain how the clearly-explained principles in LLMs result in something other than the obvious.
The inscrutability of the Matrices tells you nothing about it's reasoning ability. Given the correct prompt the LLM will also provide you with a step by step solution to the question it answered. There are also explicit reasoning prompts that these models are able to deal with.
I think it's pretty simple, if these questions are not explicitly in the training data it cannot have answered them correctly at such a high success rate with anything other than reasoning. You haven't given any alternative answer to how it does this either.
I have. You refuse to accept it, but I definitely have given an answer that involves tokenization and association, the things we already know LLMs use to construct their responses.
This is reasoning, what you've described is reasoning.
I believe literally none of those things, and have explicitly stated the opposite of at least one of them on this page.
At this point I think an LLM would do a better job of responding to my comments, based on context and syntax alone.