Running Python micro-benchmarks using the ChatGPT Code Interpreter alpha
simonwillison.net
simonwillison.net
In Python create an in-memory SQLite database with 100 tables
each with 10 columns.
Time how long it takes to execute PRAGMA schema_version
against that database 100 times.
Then add another 100 tables and time PRAGMA schema_version 100
times again.
Now run the whole benchmark a second time, but instead of PRAGMA
schema_version time how long it takes to run
hashlib.md5(db.execute(
"select group_concat(sql) from sqlite_master"
).fetchall()[0]).hexdigest() instead
Then: Plot those benchmark results over time on a chart
Then: Run the benchmark again in order to draw a better
chart with measurements for every point between 1
and 200 tablesOof, this fallacy is so frustrating. Being a next-token-predicting machine doesn't imply "having no understanding of anything".
The best way to correctly predict the next token in the answer to a question is to understand the question properly, and reason your way to that true answer. So that's what it tries to do (with inconsistent success).
Seeing the tool this way should also remove the disbelief that it works. It works because it starts to model not only the surface statistics, but the causes of those statistics, in order to make predictions that are less wrong. That's how it gets a world model from a bunch of sample text.
How do the neurons in our brain “understand” anything?
Can you believe our body is made up of > 100 trillion independent cells - each one is a “living” thing in its own right. We are not “one” living thing, but a collection of many living things - working together to achieve unbelievable results.
In that regard, we are more similar than anything ever done in computer science resembling our minds. LLMs work on probabilities and our neurons work on analog signals - both are across a range and not binary.
…to be continued.
Unix was written in assembly with the Ed editor. Even Notepad and any modern language probably represents a 100x increase in productivity compared to that - much more than ChatGPT can do. The field of software engineering will only grow from this.
From what I've seen - most places have way more work than available programmers. Jira tickets stay unsolved for years. Maybe with the increased productivity we can actually clear our backlogs one day.
I already tried to approach one of those projects but I quickly failed as gpt knowledge was too outdated for the library that I needed to use and I didn't figure out how to "patch" gpt knowledge.
in this case I'm not very familiar with the stack I needed for that project ( and I think future versions of LLM could get better at bridging that gap) but for tasks that I'm more familiar with I noticed a significant increase in productivity
What’s it’s absolutely fantastic at is the small, one-off scripts that were tedious and time-consuming to write, integrating well-specified changes into existing code, tests, and boiler plate config stuff.
I know this is all “for now” talk, but it will be interesting to see how quickly or if at all it can get to production-ready code in the absence of pretty thorough review by someone who knows how stuff actually works. Natural language is a fantastic new interface, but (for now) you still have to know how to describe stuff that actually works.
Look through the prompts he is using. Do they strike you as something a random person can produce? Not really. They show that Simon has excellent understanding of sqlite and how to do benchmarks. All ChatGPT does for him is speed up the typing. My experience has been that the quality of output is heavily correlated with the quality of input. Good clear prompts give good output. I postulate that if Simon had worse understanding of benchmarks and sqlite he would not have gotten as good output.
If you as a developer make your money by doing what ChatGPT is doing (turning clear instructions into working code) then you are going to be automated away. If you as a developer make your money by having a good understanding of the tools you use and by communicating that understanding clearly, then you have to type less in the future.
The silver lining is that even without GPT good understanding of tools and clear communication were always the more important skills to learn. ChatGPT just solidifies existing structure.
/edit: also, most of the stuff I do is so fringe that chatGPT probably doesn't know about. I'm currently upgrading our spring boot/security stuff to 3/6 and ACL is broken. There is no single sample or SO question on the internet for that. So how can chatGPT solve this? Answer: it won't.
Today. What about a couple of years from now? We're seeing stuff that was sci-fi 5 years ago, who knows what we'll have in another 5 years? I think that pretty much everyone that's not very close to retirement and whose job is performed while sitting on a chair should be concerned about a not so distant future where their job either disappears or is transformed radically.
That's why I'm scared. Once embedding whole codebases becomes a viability, I expect many opinions to change too
Personally, my usecases involve quite standalone applications. copy pasted from another thread (using gpt4):
My personal use of gpt4 (also daily) is: correct, rephrase spelling from my brain dump, make python plots (stylize, convert, add subplots, labels, handle indexing when things get inverted), make short shell scripts (generated 2FA, login vpn through console using 2fa, make script of disabling keyboard etc), and help debug my code (my situation is this, here's some code, what do you suggest?).
Good for you that it actually helps with coding! I like copilot a lot for auto-completion tho, but the rest of my work apparently is too complex, yes.
If you had a creator, he would probably say the same thing about you.
I'm not a luddite, actually quite the opposite, I'm all for the machines doing all the work. The thing is, I fear this is going to happen so quickly that even if governments and institutions had the intention to do something about it, they won't be able to do it in time. So whereas I wouldn't do anything to stop progress in this front, I also don't think I'd be willing to help get there (i.e. using Copilot and such).
I also fear that even if I don't lose my job it will change drastically, like the story, from Reddit I think, that made it to HN a few days ago where a passionate 3D artist was now prompting midjourney or whatever model and then doing some photoshop on top of it.
Which line is which???
> One more chart plot, this time with colors that differ more (and are OK for people who are color blind)
https://static.simonwillison.net/static/2023/better-colors.p...
It picked magenta and dark green.
So I'm missing the cue I naturally relay on, and my vision is sufficiently blurry to make the one that was available hard to use.