Tips for better coding with ChatGPT
nature.com
nature.com
"Oh shit, something huge is happening and we might not be intelligent enough to see the ramifications, but here's a humble attempt"
rather than
"Ha, the techies have done something they think is impressive again. It's certainly interesting, but as usual they're exaggerating it and failing to think from a nuanced human perspective. However us journalists are trained in that sort of thing, so we can help out here, and we know you readers will be all too familiar with the way those techies can only think like computers lol."
How much more productive do you feel you are coding with an LLM? As another HN user said, to me it's like a talking dog -- incredible yet useless.
There has been some work done on generating code in a feedback cycle to winnow out hallucinations and it seems to work fairly well [1]- I think 99% of the challenges LLM face are primarily related to a lack of constraint, optimization, agency, and solver feedback. As they get integrated into a system with the ability to inform and constrain and guide using classic AI techniques their true value will be attainable. But they’re pretty useful even today.
N.b., I’m a 32 year veteran distinguished engineer level at FAANG and adjacent firms that programs daily.
That's if you're explicitly asking it questions. Have you sat down and coded with Copilot offering suggestions as you went? It's honestly incredibly helpful, especially when leveraging a new language or stack.
I almost always begin my tasks with ChatGPT, by asking it for a framework. "Design a x request that gets me results in form of this JSON, and now display it on y table according to this criteria"
It gives me framework, I adjust and expand upon it. Works wonderfully.
I even have a 'bot' running on one of the IRC channels that was almost 100% self written by ChatGPT.
The world right now seems to be divided more into “have GPT-4 API/ don’t have GPT-4 API”.
What is interesting is having access to GPUs (cloud or real hardware). The world is divided in "has access to enough GPU power/don't has access to enough GPU power".
My prediction is that this will get worse for a while.
Get out of town.
I meant 40B Falcon is better than 65B LLaMA in many benchmarks and we will see even better models released with time.
What is wrong with that?
Competitors aren't even at GPT 3.5.
I’ve been looking into this. Nothing definitive yet, but my hunch is that current LLM’s struggle because they lack curiosity. They will answer your question, however vague, with the first most obvious answer they can think of.
This is great, if you’re a junior team member. Super talented, very eager. But a more senior engineer approaches the problem differently. They ask more questions than provide answers. They’ll happily spend the first 25min of a 30min meeting asking heaps upon oodles of really dumb sounding questions.
Then the last 5min, that’s the magic. They now have a solution perfectly tailored to the problem at hand, with all sorts of edge cases either explored or eliminated through questioning. The questions they kept asking weren’t dumb after all, they were pruning a decision/option tree in their head of all possible solutions until they landed on the most optimal solution for a set of known constraints. With further options to dig and improve.
I think you can make an LLM do this (i’m trying), but it’s very very slow still.
If I had to make a bet, I suspect Google and Meta will royally lose this generation: Apple will dominate personal access to LLMs, Microsoft will dominate business access to LLMs -- and honestly Palantir for the first time terrifies me. The current generation of LLMs are not about automation as such, but helping individuals accomplish self-directed goals (as possible in language and code).
This is because curiosity requires free energy, hence it is very expensive when you're limited on very expensive compute.
This is what tree of thought is attempting to simulate in some ways. Build a set of multiple questions around the original question and then build on and prune that list based on a 'show your work' set of steps, and then keep iterating.
Humans naturally solve the halting problem when thinking about things... we work on something long enough without a break and we'll pass out. Maybe when we wake, eat, and go to work we'll stop working on the same problem. But an LLM never sleeps. In theory with TOT and no time limit, you could find out your AutoGPT spent 10 million in computing resources contemplating navel lint. So, there are a number of unsolved problems there.
What really becomes concerning is if Nvidia achieves its goals of speeding up training/inference by 1 million times in the next few years, and if the amount of compute we produce increases by a few million times. You and me simply can't use hundreds of minds thinking for years straight and machines could.
I don't have much opinion on this though, i just know i see that description of GPT4 a lot. I use GPT4 daily, but i think i don't really challenge it most days, as i've never trusted it enough to do so hah.
What worries me is that I have seen people here more often commenting on the author's life, their academic or professional attributes instead of paying attention to the article.
I find that a bit exhausting; beyond simple use-cases, I've found it easier to just do it myself.
Based on my own attempts to have ChatGPT generate code for me, it's better to NOT trust it. It tends to hilucinate even with simple requirements.
GPT: Alice; GPT-2: Bob; GPT-3: Carol; GPT-3.5: Mountain Carol; GPT-4: Dallas; …; GPT-19: Sydney
You mean other than the fact that it costs money?
Okay, the tip with anthropomorphisation is okay, and reminding people that you can't believe in LLMs output is always good... but, like, I'd assume an average Nature reader is smart enough to not need tips like "iterate" or "embrace change".
Regarding Malcolm Gladwell, do you mean the "Blink" book?
Here are some of my thoughts on how to code with chatgpt. Many overlap with topics covered in the Nature article.
- It really does help to think of chatgpt as a junior coder. I mentioned this in the writeup about my first AI coding project [2]. A bunch of things in this list would also be helpful when working with a junior dev.
- GPT isn't good at code architecture. It will repeat code and take the shortest, laziest path to coding up your request. You need to walk it through changes like you might with a junior dev. Ask for a refactor to prepare, then ask for the actual change. Spend the time to ask for code quality/structure improvements.
- Don't copy-and-paste between a chat session and your files. Use a tool like Copilot [3] or aider [0] that will give GPT your code to edit and apply its changes automatically. This is critical as it removes friction and lets you iterate quickly on code with GPT.
- GPT is amazing at generating fresh new self-contained code. It takes more skill and better tools to work with it to edit existing code.
- Break down a big change and ask for a series of smaller, self contained steps. This is a smart way to code solo, but really helps GPT succeed at more complex code modifications.
- GPT has a wide breadth of coding knowledge, and impeccable command of syntax... but it is sometimes overconfident about details. So it will usually choose the right library for a task, but sometimes make an error about the details of the api/method/params. Paste doc snippets or error messages into the chat and it will fix those bugs.
- GPT is really strong at churning out good boilerplate and roughing in a solution. This reduces the activation energy required to start larger changes or changes which will require unfamiliar libraries/packages/languages. GPT can can quickly prepare the ground for you to do the interesting work.
- If you can, use gpt-4 (not gpt-3.5-turbo) since it can successfully generate larger, more complicated code changes.
[0] https://github.com/paul-gauthier/aider[1] https://news.ycombinator.com/item?id=36020809
I've also been working with the APIs directly for about that time and I only get more depressed daily reading the horrible takes about GPT's problems based only off of ChatGPT (and hearing about how even many companies try to use the product). GPT has forever changed how I interact with computers.
I was using CodeGPT but this looks better. And it's exciting to have those command line options to easily insert other contextual input like command outputs.
Can I ask about your plans? CodeGPT planned to add semantic embeddings of code, but I haven't seen any updates for quite a time. The idea being that, like walking an Abstract Syntax Tree of the code, you could include relevant functions and function definitions(/classes/constants/etc), to help improve output quality.
Alongside what you note with module/library versions and pasting in doc snippets, I have often found is necessary to include package.json files to nudge the results towards the right API calls. I think semantic code search might push it finally to being able to work even on comprehensive, multi-service style repos.
For MVC or services architected solutions, I think this could be more helpful than including whole files - though it may be overkill - do you have any thoughts on doing this?
Yes, aider already has features to provide GPT with "code context" to let it edit larger, more complex codebases. I wrote up some notes about these features:
https://aider.chat/docs/ctags.html
You might be especially interested in the "future work" section near the end. I have actually shipped some of these ideas into the tool already. I need to find some time to write up the details.
Great to see, thanks again for developing it.
These are all good tips and very much align with my experience over the months, but breaking a complex task into as small, self-contained and logical steps as possible is truly going to make the biggest difference for most people.
Once I understood that, just like previous versions (GPT-3), ChatGPT has no internal memory beyond what has been typed, that changed how I interacted with the model for the better. Depending on the task, I am either providing simple steps or asking the model to write simple steps based on my task layout, then I provide input on those, ask for refinement of said steps with a focus on code generation and remove anything superfluous/not focused on code creation, the output quality became significantly more consistent and usable.
I can understand why, for a lot of more experienced developers, this can seem tedious and fully get why this leads a lot of observers to feel that, considering the effort required to set these models up for success, they might as well code themselves, especially considering this setup process can eat into the precious 25 prompts limit, which I hit consistently.
However, I feel that once these models have become more efficient, a lot of this outlining may be handled in the background.
The same goes for checking code errors in my eyes. For the same reasons (no internal memory), there often can be errors in provided code, yet asking for the model to check for errors generally resolves those. In most cases, this works without providing compile errors, though those of course do improve the response.
I try my hardest not to anthropomorphize these models, so pardon this comparison, but in fairness, even with our human wetware memory the average dev is rarely able to write flawless code without the need for any revisions on their first try.
If we get a more efficient model of comparable quality to GPT-4, adding a line to the frontend to request a second pass on all code based request may not be unreasonable and could, in my experience, yield more consistently usable results.
Considering how quickly OpenAI released gpt-3.5-turbo, I am hopeful that we will see such a development soon, though I currently do not have personal access to the GPT-4 API, so maybe that could be made more accessible first.
https://github.com/knowsuchagency/llmo
I'm really excited to try out aider, thanks for making it!
Many orgs have a big enough codebase that they can improve productivity with existing patterns.
I can't remember what the other platform options are, but for VSCode you need to use a preview version, which has meant I haven't had much time to test it yet as I'm full on with the stable release.
But as an example I can select a block of code and ask it to generate tests, explain it or various other stuff. I haven't tried it further yet to see to what extent is provides value beyond just saving copying and pasting and establishing a context in ChatNormal.
It's not clear either that it is 4 as you requested; it may be the default 3.5, I've had a very brief look but couldn't see the spec.
Update, according to the Verge, its 4:
https://www.theverge.com/2023/3/22/23651456/github-copilot-x...
Maybe I am just getting old but the idea of using a non-deterministic tool that can hardly be reasoned about that will straight up hallucinate facts for any professional work sounds insane to me.
Yes, I do see the value for Junior Devs as I am sure it can drastically increase their output in the short term but aren't they shooting themselves in the leg in the long run? That might sound elitist but at the end of the day, you will need to learn to read technical documentation anyway and once you understand it there is no need for ChatGPT.
Yes, if you are constantly hopping from one framework of the month to the next big thing, you just don't have the time to learn anything in depth and then again ChatGPT can help. But do you really want to live like this? Instead of band aid solutions we might want to push for less hype-driven and more pragmatic development styles that allow us the time to learn our frameworks.
80% of the time, the suggestions are either too generic for our codebase and so has to be patched up with our variable names, methods, etc. About half the time it completely misses the mark/intent of what is being written, and offers an auto-complete for a different problem.
20% of the time, it actually acts like a full line or full section autocomplete. Areas I've found it particularly helpful are writing tests. For example I added a new integration yesterday that touched a bunch of interface files for those integrations. Copilot was able to pick up on that and auto-generate a 90% functional test with ~30 lines, including setup and assertions.
I'm not convinced it's saving me time yet since I have to correct it/ignore it most of the time. When it works, it's cool. And like another comment said, typing speed has never been a bottleneck for me (record of 155 wpm on typing tests, I probably sit around 120-130 for average day-to-day typing).
For example, we plan to use GPT-4 for generating text that ends up being seen and read by our customers. (Can't disclose the details, as it's near to core of our business model.) We need a _lot_ of text, and its in style that is time-consuming to think and write by hand. However, all the generated data is checked for correctness by humans. I think that these tips are excellent for designing the generating process so that the end result is of as high quality as it could reasonably be expected to be.
Edit: Oh, you mean coding work as for generating code for developers? Not in the sense of using it as a component of a data-generating/transforming system? In that case, I pretty much agree.
Edit2: Oof, I just realized that I mistook this whole thread for another thread that was posted earlier on Hacker News. So I thought we were talking about engineering systems that has GPT-4 as a subsystem: https://platform.openai.com/docs/guides/gpt-best-practices
Text generation with human proofreading/editing is definitely a valid use case. Maybe not for top-tier prose but yeah if you need large amounts of decent quality text it is the way to go. Definitely makes some new business ideas viable.
It's good at generating simple things that are 60% to 100% correct.
It's good at summarizing high level concepts or summarizing how multiple high level concepts relate with 60% to 100% accuracy
It's ok at helping you think about things to troubleshoot when you have an error you're not sure to do with.
To be honest I would say it's actually not going to be good for junior devs because they don't have the skills to properly fact check, quickly. But for more senior folks it's very easy to immediately see what's wrong, ignore those parts or ask it for clarification, and use what's right.
If you mostly know what you're doing, it can be very helpful to get immediate feedback, niche examples targeted exactly to what you're working on, or summaries of whatever high level concept you need a little more clarification on.
Much much faster than spelunking through Google.
As a concrete example I just this morning had it summarize Python's asyncio Queues versus Tasks, what they are, a few examples of using them, and when you would want to use one versus the other. All in just a few minutes.
But yes, I have seen it make up its own Powershell cmdlets (my main usage)
So cannot be used unless you already know.
I feel exactly as you do and I've reviewed code, found all sorts of weirdness and then realised that the developer used ChatGPT to do it and probably to do its unit tests. It all worked within itself but clearly didn't work in the real world at all! :-)
Ultimately, I cannot really be bothered to use ChatGPT. It might be my backwardness but I can be bothered to use an IDE when it helps (instead of VIM) so I don't think I'm a hopeless fundamentalist. Why has my brain written it off? I don't know but I know I am lazy and I do things that make life easier.
Getting a simple example roughly tailored to my needs, which I can use as a launching point, is often extremely useful. For example, I encounter scientific libraries with poor documentation and I could spend half my day reading through it or searching stack overflow for a good explanation…or, I can ask ChatGPT4 for a specific implementation using a specific library, then ask it to explain anything I don’t understand.
These tools aren’t solving any real problems for me - they are replacing slower and less responsive systems I already relied on for research and development.
This sounds like a decent enough description of the human brain.
Don't get me wrong: I'm not for a moment suggesting that what we have today is anything approaching "General Artificial Intelligence", and I am extremely concerned & worried about the inevitable massive damage AI is going to do to our world, but I do think it's a little funny that many people's specific objections to it amount to "it can do what people do". Yes. That's the idea.
The main concern with AI usage is not that it will write bad buggy code or lie: we already do that ourselves plenty, so that's not a novel skill in the professional arena. The main concern is that we can scale that stupidity.
But that's a problem of scale: individual usage isn't really going to do you notable damage on an individual level.
> Yes, if you are constantly hopping from one framework of the month to the next big thing, you just don't have the time to learn anything in depth and then again ChatGPT can help.
I actually think the opposite is true. Having used ChatGPT quite a lot for work (by mandate - I wouldn't have chosen to either but am glad now that I had to), I've found it's really very good at generating bad code and being confidently wrong (again, much like people): if you ask it about a subject you're not deeply knowledgeable about, it's going to lead you astray. The most astute way to use it is actually in going the last mile on something you're already very confident in, so that you can correct it as needed.
Granted, us humans are doing that now. Just many orders of magnitude slower. Probably slow enough that we can find/fix the important stuff.
The other aspect that terrifies me is the potential for nation state entities with deep pockets to inject vulnerabilities. What would it be worth to the NSA to seed the future with programs they could exploit?
Granted, us humans are doing that now. Just many orders of magnitude slower. Probably slow enough that we can find/fix the important stuff.
Yes, there is missing/unclear language specifications and compiler bugs and sometimes you get surprised by certain optimizations but you can reason about compilers well enough for practical purposes. Plus it's not economically viable anymore to do without them, sadly.
Real world example: Hey chatgpt, write me some Camel routes to take an input off of a JMS queue, wire it up to the camel salesforce getSObject endpoint, use these headers for the lookup, and write some unit tests.
Actual result: Code that is 99% correct, but messed up a couple of parameters and used a slightly wrong approach to mocking some things in the unit test. Having used all of the frameworks extensively, it took ~5 minutes to fix and certainly saved time. Partially because I'm used to reviewing code from people who make similar mistakes.
I can't imagine that working out as well if it was a truly junior engineer (not just a junior in name only) who didn't understand what was going on with the code very well.
And this is just the beginning. It is hard to imagine not using this for the bulk of the work within a few years.
I believe he is very familiar with C, R, Python, Julia, and other languages.