Using GPT-3 to explain how code works
simonwillison.net
simonwillison.net
- GPT-3 can produce simple SVG diagrams (https://twitter.com/goodside/status/1545883023719239689) from an English description, or incrementally modify a given SVG diagram (https://twitter.com/goodside/status/1545888869207400454).
- GPT-3 can produce diagrams using Graphviz dot notation (https://twitter.com/goodside/status/1545109709035900928), including diagrams that incorporate real-world context (https://twitter.com/goodside/status/1545112039147847682) not given in the prompt.
- It can emit and understand HTML and Markdown, including table markup — e.g., "Produce an HTML table giving the members of the Beatles and their instruments"
- It can accept/emit multiple named files, with arbitrary content, simply by embedding those files as values in a larger JSON/YAML document. Here, (https://twitter.com/goodside/status/1545848754883919872) I demonstrate completing the file tree of a trivial Python package given a README.md.
- It can understand the raw JSON file contents of .ipynb files, successfully following directions given to it in markdown cells, emitting multiple code cells, etc.
While all this is fun, as many people have pointed out the model has a bad habit of "hallucinating" — producing highly specific, wrong output. Avoiding this (https://twitter.com/goodside/status/1545793388871651330) is more of an art than a science. Be highly skeptical of any output you get from GPT-3, and search for verification methods where you can.
Prompt:
Perform a code review of the following module, identifying any bugs:
```
#!/usr/bin/env python3
def fib(n):
"""Return n-th Fibonacci number."""
if n <= 2:
return 1
else:
return fib(n-2) + fib(n+1)
```
GPT-3 completion: The bug in this module is that the n+1 parameter in the return statement should be n-1.
But don't expect miracles on real-world files, because GPT-3 struggles with long-term dependencies in text (e.g., remembering what happened ~4000 characters ago). Perform a code review of the following module, identifying any bugs:
```
#!/usr/bin/env python3
def factorial(x):
"""This is a recursive function
to find the factorial of an integer""" if x == 1:
return 1
else:
# recursive call to the function
return (x * factorial(x+1))
```
A: The bug in this code is that the function is attempting to find the factorial of an integer by recursively calling itself with an incremented value. This will cause an infinite loop.But if you unindent and then also perform the n+1 error it only will see the unindent issue.
Perform a code review of the following module, identifying any bugs:
```
#!/usr/bin/env python3
def factorial(x):
"""This is a recursive function
to find the factorial of an integer""" if x == 1:
return 1
else:
# recursive call to the function
return (x * factorial(x+1))
```
A: The code will not run because there is a syntax error on line 6. The indentation is incorrect. The return statement should not be indented.Two uses come to mind: 1. Alternative variable name suggestions. I am writing code, and I try to figure out what to name a variable that about a data frame which is my performance data which takes the maximum of random samples and batch size leaving data size and thread count, and all I can come up with is unbatched_speed or something. I can think of a lot of scenarios where a context sensitive thesaurus / autocomplete would help. If it's obvious to an AI, maybe it'll be obvious to a junior dev?
2. I expect there are a lot of cases where not technical people could understand code, but get intimidated because it looks like code. An AI to say what the code for the business logic that they care about, but using prose.
Maybe a way to do the same with formulas. I always get hung up on the odd notation and the symbols that aren’t defined anywhere.
I'd like to see it get good at continuing poetry.. It's pretty awful at that, from my own experiments.
It even does pretty good with Verilog, which I assumed it would have issues with, admittedly some of this is just "google-able", but even so. It can even create a reasonable attempt at a testbench (even though it doesn't quite get it right), which is crazy to me.
*Explain this code:
module my_module(a,c,y); input a,c; output y; assign y= c ? a : 1'bz; endmodule*
This code is a Verilog module that defines three inputs (a, c, y) and one output (y). The output y is assigned to the value of a if c is true, otherwise it is assigned to 1'bz.
*What is 1'bz ?*
1'bz is a Verilog constant that represents a bit value of 'z', which is used to indicate a high impedance state.
*Why is high impedance used ?*
High impedance is used in digital circuits to allow multiple drivers to share a single bus without contention. When two or more drivers attempt to drive the same bus at the same time, a contention can occur and the bus voltage can become undefined. By using high impedance, drivers can avoid contention and allow the bus to remain stable.
*Give me a testbench for this module*
module testbench; reg a; reg c; wire y;
my_module dut(a, c, y);
initial begin
a = 1'b0;
c = 1'b0;
#1;
a = 1'b1;
c = 1'b1;
#1;
a = 1'bz;
c = 1'b1;
#1;
a = 1'bx;
c = 1'b1;
#1;
$stop;
end
endmodule*Create the same testbench but add validation of the output.*
module testbench; reg a; reg c; wire y;
my_module dut(a, c, y);
initial begin
a = 1'b0;
c = 1'b0;
#1;
a = 1'b1;
c = 1'b1;
#1;
a = 1'bz;
c = 1'b1;
#1;
a = 1'bx;
c = 1'b1;
#1;
$stop;
end
always @(*) begin
if (c == 1'b0)
assert(y == 1'b0);
else if (c == 1'b1)
assert(y == a);
else if (c == 1'bz)
assert(y == 1'bz);
else if (c == 1'bx)
assert(y == 1'bx);
end
endmoduleThis is incorrect.
(?xm) enables the VERBOSE (X) and MULTILINE (M) flags.
From the docstring:
M MULTILINE "^" matches the beginning of lines (after a newline)
as well as the string.
"$" matches the end of lines (before a newline) as well
as the end of the string.
X VERBOSE Ignore whitespace and comments for nicer looking RE's.Works with mermaid js
https://twitter.com/Bsunter/status/1535062559069376512
Great for making study guides
https://twitter.com/Bsunter/status/1535730704138444801
I also wrote a plugin to interact with gpt3 in the logseq text editor and some examples here.
https://github.com/briansunter/logseq-plugin-gpt3-openai
I have a few posts on my blog
https://briansunter.com/blog/gtp3-openai-logseq-notetaking/
https://briansunter.com/graph/#/page/how%20to%20use%20gpt3%2...
In any case, my question here is: how do you not go mad in this techno-religious environment where people (seemingly desperately and earnestly) equivocate correlations in written text with thinking?
I feel at once overwhelmed, despairing, and a sense of need to point out that it isn't true. It's the most obscene form of wishful thinking and superstition.
HNers seem capable of identifying pseudoscience when its an article in nutrition (namely, eg., that correlations in the effects of human behaviour are not explanations of it)... but loose all sense when "AI" becomes the buzzword.
Correlations in *TEXT* are not models of language, intelligence, thinking. This is pseudoscience. And the most extreme and superstitious form: thinking that correlations are meaningful because we find them meaningful.
A person using language uses it in response to an environment, so "Pass me the salt" is said on the occasion that one needs salt. The "language model", `language = person(...)` includes the environment as input (and many other things).
Correlations in pixel patterns of the night sky are not a model of gravity. Correlations in text are not a model of language. Correlations in the detritus of human behaviour are not models of intelligence.
But to my question: how does one cope in an environment of this kind of wishful thinking?
I think I answered it myself... I have been seeing people's fervour to attribute intelligence to AI as either charlatanism or misunderstanding... but I think, now, it's actually wishful thinking (and hence essentially religious).
Trading in the usual superstitious foundations of religions: "AI works in mysterious ways", "correlations in symbols are meaningful", etc.
I still don't know why people wish that AI were intelligent; but at least now I have a clearer sense of it. There's some utopian faith behind it which creates this odd urgency to "require" machines to really be intelligent. I'd be interested in knowing why this wish exists.
I can never quite put my finger on how irreligious tech people end up in these tech religions. I guess there's some attachment to there being "a heaven somewhere", which with tech, will be heaven-on-earth.
As human beings, we slowly evolve from one generation to the next but our tools can evolve much more quickly and it can give us in a sense, superpowers. Cars were invented and allowed us to travel great distances while Self-Driving cars (which will require some sort of AI) will reduce car accidents and provide mobility for everyone.
AI has the potential to greatly improve the lives of a vast number of people and so it might be looked-upon as some sort of "saviour" and that has some religious connotations but most people still would treat it as just a tool but possibly a life-changing one.
We can not unequivocally explain what thinking is (in a human sense), and what makes it vastly different from Cephalopod thinking or a ridiculously large NN 'thinking'.
Until we reach a body of knowledge in Neuroscience that can give such an explanation on is it that makes human thinking somehow special (other than scale) then it is equally religious to assume that human consciousness or human creativity are somehow unique (in nature) and irreproducible (with technology).
This is an argument from ignorance: you dont know what caused the universe, so maybe god did.
Well, we do know much more than enough to regard any use of NNs today as having nothing to do with intelligence; and likewise, that any model of correlations in symptoms of human social activity (text, books, etc.) as not relevant to modelling/simulating intelligence.
A neural network model is just a set of averages over training data, which in this case, are simply recording correlations in word frequency. There's no neural "Neural network", it's just some mean()s.
We could claim the human brain is just 100 million neurons arranged in a mesh were each neuron has their own 'activation potential'.
Neither your or my statement are wrong. But again until we can pinpoint what (human) intelligence is, then claiming squid or neural networks are unable to be intelligent agents is just as emotional as the people you criticize.
You actually can't because, unlike neural networks (we know exactly what they do), we don't have anything even close to an exhaustive view of the functioning of the human brain.
It doesnt matter what algorithm in this family we choose, they all just produce correlative models of the data given. So if we give a NN a library of books, it produces a function whose structure is just correlations of words in those books.
To believe that the structure of this function has anything to do with language, the brain, or anything else is pseudoscience. It isnt close to science, it isnt an open or a philosophical question, it's homoeopathy.
Language use in humans is not a function of correlations of word frequencies. We know this, as much as we know where mars is.
No human is shown a trillion books and arrives at an internal model of language corresponding to word frequencies in those books. We know that it would be impossible to do this, since words refer to the world -- and correlations between word frequencies dont.
Your comment here defending this is equivalent to a homoepath saying "we dont know everything about chemistry". We dont know everything, but we do know that diluting with water doesnt make substances more efficaious.
And we do know that the correlational structure of text in books captured with any algorithm whatsoever, has nothing to do with modelling language. We know any such system cannot use language, as at least, it has no model of how words refer to the environment.
However:
> Language use in humans is not a function of correlations of word frequencies.
I would argue that this is untrue. For example, people pick up on common expressions (and other behaviors) of those around them. If I hang out with people that say "catch my drift" a lot, I will probably also start using "catch my drift" more. I find it very plausible that "my internal model of language" is in many (but not all) aspects driven by the word correlations I observe around me.
I do wonder about the absoluteness of your statements though. Would a bed-ridden person that has no sense of touch/sight/smell etc. except hearing (i.e. only able to interact with the world via language) be "not human" to you? If not, how is "the environment" different for it and GPT-3 (other than context window)?
When I say "not a function of" i mean, explicitly, that the model does not have a term which depends on these correlations. Just as F=GMm/r^2 does not have a term for "image pixel weights".
Real Language is what wrote all those books, as much as gravity is what made the night sky. Pixels and word frequencies are symptoms, a few of an infinity of symptoms -- distant effects, whose correlations are not models of their causes.
What makes you sure that your thinking process is different?
What's the difference between a model and complicated correlations?
The most basic, most broken down, most simplified answer is what is called "poverty of stimulus".
A NN has an abundance of stimuli, orders of magnitude greater than any biological being could ever possibly take in. Finding enough correlations (enough for productivity) in such abundance is not actually that surprising. What is surprising is kids being competent at speech after less than 2000 days of stimuli, often really low-quality stimuli.
With that said, my guess is that we'll probably have to take a few more hints from the biology of the brain before we are able to achieve "human level intelligence".
A contemporary CPU and/or GPU is 'a bunch of connections that are pre-encoded' - just not in DNA, but in silicon.
No, CPUs/GPUs are not relevant if we are interested in the speed of learning in relation to the quantity of stimuli processed. You could even compute an ANN's learning algorithm with a pen and paper and it wouldn't change its "learning speed" within that definition. Pre-trained weights would.
Regardless, this head start is probably not sufficient to explain the disparity between the brain's capacity to learn and modern ANNs. ANNs are probably just not a very good approximation of how the brain works, for now at least.
For me the real question is how well can GPT-3 identify pseudoscience.
How can you not be impressed by something that would have been considered inconceivable just a few years ago?
At no point since digital computers were invented was it inconceivable that we could continue prompt phrases by searching through similar phrases, near such prompts, in digitized text.
It was thought computationally infeasible, in just the same way that a harddrive would never be TBs, and so wouldnt be big enough.
But there's no magic here, and nothing suprising. It's presented by charaltans who have money to make as an achivement that it isnt. This isnt intelligence, and has nothing to do with it.
It can complete phrases by having seen TBs of similar ones no more suprisingly than google can. It is as limited as that process implies, namely, radically limited.
Likewise, it was never inconceivable that one could 'play' chess by searching through billions of chess permutations and selecting a simple high-value path. It was thought infeasible given, say, an abbacus to do it with.
We have never thought totally dumb solutions to these problems were unimaginable. They were just so dumb, that the only route to using them would be to somehow run a Ghz processor. Well, since we now have those, dumb partial solutions to dumb problems are possible.
Hard disagree here. Google can't do what GPT-3 does and it has access to a lot more than 350gb (size of GPT-3 model). I doubt you literally believe that.
> We have never thought totally dumb solutions to these problems were unimaginable. They were just so dumb, that the only route to using them would be to somehow run a Ghz processor. Well, since we now have those, dumb partial solutions to dumb problems are possible.
Did the GPT-3 architecture exist 10 years ago? Could an ANN from 10 years ago have produced comparable results, given access to the same hardware? I doubt so.
It basically seems that in your worldview, anything below "human level intelligence" is just unimpressive? Seems quite a high bar.
The mathematics of NNs are Victorian, it's just regression. The NN architecture is largely irrelevant: that only determines how easily the regression converges. The strategy is the same: just compress the data into a correlative statistical model, and estimate using statistical associations.
That's the wrong category of solution, and as a category, it's anchient. The hardware of modern digital computers is impressive, NNs are an absurd joke -- as is any correlative model of data -- a magic trick carried entirely by the performance of CPUs and HDDs.
Newton did not predict the existence of an undiscovered planet by finding correlations in the stars of the night sky. He first produced an actual theory of gravity.
Intelligence is theorising, ie., it is an explanatory process. Correlations in the frequencies of characters in books (we wrote!) is the basis for a neat trick.
There's a tremendous amount of pseudoscience in the claim that these are models of anything at all. Inasmuch as correlations in patterns of pixels in the night sky is not a model of gravity, and provides no means to "predict" anything other than pixel patterns of the night sky.
Back to my point, your bar is pretty high. Artificial intelligence to the level you describe will basically be a "world as you know it is now obsolete" type discovery. It will likely go down as the most impactful discovery of all time. Most of us have a lower bar for what counts as impressive.
I am, yes, anti-wonder; I do think people get hopped-up on wonder and that lowers their threshold for reasonable belief and they end up in (religious) absurdities.
My issue here, in these terms, is that the "wonder" associated with the magic trick of modern ML is preducing a gross cognitive bias in which people believe this is an example of intelligence; or otherwise, an example like us in some significant sense. This leads to pseudoscience, charlatanism and a religious techno-utopian faith.
At the heart of it is a wonder-scam in the vein of all miracle-based shysterism. The fool looks at the miracle, is impressed out of all reason, and concludes miracles exist.
For sure, I dont really trade in wonder and I have never found it to do anything other than profoundly distort people's beliefs.
It's like someone in crypto talking about a future crypto-based financial system whilst at the same time being unable to processes basically any transactions, stabalise any currency, ensure any regulation, deal with any fraud, etc.
Or likewise, an MLM health company talking about cures for cancer all the while hocking vitamin supplements.
The question of what properties animals have which enable them to implement intelligence isn't relevant to whether correlative statistical techniques which capture word frequencies are "intelligence" -- inasmuch as vitamin 'science' isnt relevant to curing cancer.
This gross level of chalatanism, wishful-thinking, scamming, delusion and out-right fraud is overwhelming.
Try telling that to Sutskever or Karpathy (or indeed a bunch of "rationalists" on twitter, at least a subset of which works in AI) because both of them have made the claim you assert you haven't seen.
I'm sceptical of AI writing programs for people, but I can easily imagine such AI guiding/unblocking non-programmers enough for them to actually program.
GPT-3 doesn’t actually know anything about anything at all. It’s a huge pattern generator. You can’t trust anything it says, because all it does is group words together into convincing looking shapes based on text that it’s seen before.
Once again, I’m reminded that tools like GPT-3 should be classified in the “bicycles for the mind” category. You still have to know how to pedal!
They’re fantastic tools for thinking, but to actually use their output effectively requires VERY deep knowledge—both of the subject matter in question, and of the way that the AI tools themselves work.
Otherwise it's just a variant of Russell's teapot.
GPT-3 has more textual training data than any human could read in a thousand lifetimes, but we base our judgment on much more than that.
That's a bit much to ask of a fiction generator, though. It doesn't know it's not okay to make things up. All its training is about making things up whether it knows the answer or not.
Sounds very human to me.
Imagine if we gave it more power and funding. Maybe redirect all of Congress's salary to it?
Either way, it's not just scaling up, it's changing the algorithm.
Intent: We can have a goal which directs our thinking. This changes what we come up with much more flexibly than experience, even if experience is often needed for good results.
Temperment: We can be angry, tired, excited, relaxed, etc. while thinking. This isn't always good, but it is a way we differ.
Self-awareness: A very difficult term to define, but we can (hopefully) take a step back and discard our current thought if we realize we are falling into one of our usual bad patterns of thinking.
The goals we want to achieve are tied to what has been going on throughout our lives thus far. We don't invent our own values out of a vacuum, we tend to adopt them from other people, and rank them depending on how our nervous system works.
Self-awareness is the ability to reason about the process of reasoning, plus, in your example, a recollection of past examples of incorrect reasoning that disrupted the process of achieving our goals.
So basically current models need a memory and the ability to change their own weights in response to new data to even start approximating humans, but that looks like a doable task, in principle.
That may be the cornerstone of "intelligence", but that's only one part of it. I think another one is making or finding relations between the parts of the whole.
Life is the continuous adjustment of internal relations to external relations [1]
Once you have internalized relations of the real world, you can start to hack them in your head, run simulations of the results and take action.
[1] https://www.thoughtco.com/famous-education-quotations-herber...
Maybe a confidence level for a given explanation, along with some sort of "here's where you can go to learn more" would be useful? No idea if language models would be good at that kind of meta reasoning.
I really enjoy how well it gives wrong answers. The internet being polluted with GPT3-like text models is going to send us back to the time when verifying anything online as truthful is difficult, if it hasn't already.
Models more advanced that GPT-3, such as LaMDA, have entire subsystems specifically dedicated to “grounding” output in truthful information. Hallucination is at least partially a solved problem, but the methods haven’t disseminated broadly yet.
I also wonder if it's GPT-3 itself arguing as the text is kind of strange. :)
I'd hate to see just the programmers put out of work by GPT-3.
Unsure if they use GPT-3 specifically, but the core idea is the same.
the same querying works but it will save you a load on gpt-3 cost https://text-generator.io/playground?text=SQLite+create+a+pa...
codex is technically free if you have access but soon it wont be so keep text-generator.io in mind there too as it can generate code from descriptions etc too at a wildly more competitive price.
Had to use a bit of the "Create Table" syntax, but it did a create table if not exists which is nice... it missed the content field though
Once these systems scale up to attention over long pieces of program text (instead of just a few thousand tokens) most of software development will be near minimum wage, almost skilled labor.
You didn't mind when taxi drivers became obsolete. "I'm not a taxi driver, so I don't care!" We'll see how you feel when your economic value approaches zero.
Another example is phone support. AI is already being used there, for example, in some banks (in Russia) you are first greeted by a robot if you call, and only if it cannot answer your question it will let you talk to the human.
For the next 10 years at least. Beyond that who knows. If we get strong AI then we have bigger things to worry about than one sector's jobs.
* Create movies, music and books on-demand
* Design interiors, buildings, electric circuits, web sites, build pipelines, software architectures, etc.
Software devs may be one of the last industries to be obsoleted, but it'll happen. About the only thing left will be the executive function of decision making, but even that's likely to devolve into:
"AI, where's a gap in the market?"
"Jonny, there are gaps in blah, blah, blah. Products that do x, y, z are in demand"
"OK AI, design and print me a prototype"
"Printing in progress..."
I feel like this could be one of the biggest technological changes I'll see in my lifetime but there isn't a lot of hype around it like say VR/AR and Web3. If this does end up exponentially improving, life is going to feel like magic.
I don't know how often OpenAI build new versions of the core GPT-3 language model.
Jokes aside, I wonder when we’ll see an official language linter (go vet, rust-clippy, etc.) that is an ML black box instead of a traditional handcrafted deterministic rule-based engine.
I then decided to ask it about Creator controversies which it gave surprisingly accurate and knowledgeable answers that pretty much mirror my own thoughts about a complex controversy that spanned over two years.
It's kind of crazy, a little worrying with just how much the AI could gather from a few nasty reddit comment threads.
Is there any chance that this code snippet was sampled from repositories that contained associated commentary/discussion, guiding GPT-3 to produce similar explanations?
Or is this genuinely an explanation that it can produce without context local to the code in question?
And what number of people are able to determine the answer to my first question?
There is a chance.
To get a better answer, you have to specify what exactly you mean by 'sampled' and by 'without context local to the code in question'.
Imagine pair programming with an intern who memorized your database schema. You don't need to go back and forth between schema and your code, you can keep your focus at the cursor. It's great for automating the boring stuff.
I would love if Github provided an API for Copilot. Imagine integrating Copilot with your shell.
Another area of huge interest is legacy code, or low level languages like assembler, COBOL, and some C stuff. Also almost everything without proper documentation.
Well played.
That would be huge
Because I think that’s both true and entirely unrealistic. Like, how would you quantify who the best experts are, and how would you verify that you’re not training on wrong data from those best experts? Maybe they had an off day or a specific agenda.
Even browsing through verified SO answers there are still answers that are outdated and answers that are just plain wrong. The greater the experience needed to answer a specific question, the more unlikely it is that there is a clear and correct answer given.
Programming while biking with ai is like relying on a really buggy early stage auto-drive ai system to take you places down winding, slippery roads while you sleep at the wheel. It might get you there most of the time, but I’d might bend you around a tree too
GPT3 isn't available; GPT2 is, as well as BLOOM which is (supposedly) as good.
You can try them online, and you can download them. Downloading them isn't as easy as downloading as app, though: the models are pretty large and need some expertise and lots of compute to run. Though Hugging Face does offer to set you up with cloud providers.
(I haven't actually try any of this, so I don't know how easy it is in practice. I've only tried the "playgrounds" that are available with some models)