Suddenly it became something 'real' then.
(This is purely talking about the public popularity of GPT)
Me? I made a few comments like a scared luddite when ChatGPT solved two of my outstanding engineering problems instantly.
I got better. But this is exactly right. The world in general now knows about AI and ML. It’s a pivot point.
When something scares a seasoned engineer for a minute, and anyone can now make use of this… write it down in your diary as a moment in history
ChatGPT is currently best at things programmers would think about. You’re correct about spatial reasoning. But try stuff like this:
“Write a python program that calculates the static forces on a cantilevered ledge 15 feet long, with a support beam”
Haha it took the longest I’ve ever seen. You may have a point. It’s really good at writing code though.
Caution. I tried my example with matlab instead of python, and I think I may have set a server rack on fire ;)
The delta of GPT3 -> ChatGPT is from the expanded context and control the model offers through fine tuning. Eg read the instructgpt paper to see the path on the way to ChatGPT.
I had wired up GPT3 to a Twilio phone number and made something basically like ChatGPT months before ChatGPT was released -- me and my friends texted it all the time to get information, similar to how people use ChatGPT. The prompt to get decent performance is super simple. Just something like:
The following is a transcript between a human and a helpful AI assistant.
The AI assistant is knowledgeable about most facts of the world and provides concise answers to questions.
Transcript:
{splice in the last 30 messages of the conversation}
The next thing the assistant says is:
Over time I did upgrade the prompt a bit to improve performance for specific kinds of queries, but nothing crazy.Cost me $10-20/mo to run for the low/moderate use by me and a few friends.
Interestingly, for people who didn't know its limitations / how to break it, it was basically passing the turing test. ChatGPT is inhumanly wordy, whereas GPT3 can actually be much more concise when prompted to do so. If, instead of prompting it that it is an AI assistant, you prompt it that it is a close friend with XYZ personality traits, it does a very good job of carrying on a light SMS conversation.
A couple years ago a friend and I trained GPT-2 on our WhatsApp chat history. GPT-2 was more primitive, but it still managed to capture the gist of our personalities and interests, which was equal parts amusing and embarrassing.
We'd have it generate random chats, or ask it questions to see what simulated versions of ourselves would say.
Even now that it's improved and free to use its actual practical usability is marginal at best given the rate of blatantly wrong info being spewed with 105% confidence at the moment.
There are some approaches. For example in this paper they say truth has a certain logical consistency that is lacking in hallucinations and deception. So they find this latent direction that indicates truth in a frozen LLM. This actually works better than asking the model to self evaluate by text generation, or training with RLHF.
"Discovering Latent Knowledge in Language Models Without Supervision" https://arxiv.org/abs/2212.03827
There's also a video with the first author: "Making LLMs Say The Truth" https://www.youtube.com/watch?v=XSQ495wpWXs&t=1515s
Btw, I think this is one of the deepest discussions about LLM hallucinations and alignment I ever saw. Worth a watch, even if it is a bit long. Not every day something like this comes long.
It makes you wonder what other abstract concepts current models may have had to learn to get as good as they are. If they're doing a good job of modelling when someone is speaking the truth, then what else have they learnt about us?
How complete of a "world model" can you learn purely in a passive way by consuming whatever online text is available to train on, or maybe by consuming all existent written material were it to be digitized? At some point I'm sure you need to be able to interact with the world to test hypothesis etc, but how far can predictive "intelligence" go without that?
1. Twitter thread with examples: https://twitter.com/sjwhitmore/status/1601254826947784705
2. Tweet/screenshot + Colab notebook:https://twitter.com/aman_madaan/status/1599549721030246401, https://tinyurl.com/codex-chat-gpt
The second tweet is mine.