Pairing With GPT-4
fly.io
fly.io
It reminds me of working with a junior developer, you need to understand how to communicate with them effectively.
That has taught me some things I didn’t know and helped me formalize and understand concepts better (I think).
Also, on a more funny note, it’s the best ffmpeg interface I’ve ever used :)
I asked it to help me extract certain values from a spreadsheet. It gave me a complicated and overly long formula using FIND and MID that didn't even work.
GPT-3.5, otoh, gave me a neat little REGEXEXTRACT formula that worked perfectly.
When I pointed out to GPT-4 that it can just use regex instead, it "apologized" and again rewrote a formula that didn't work.
you can write prompt to clean unstructured text into structured text. that becomes the work area.
instead of asking gpt "write me some code in pyton that does x" you can just tell gtp "do x with the structured data in the work area and present a structured result"
gpt becomes the computation engine
granted it's super expensive and quite slow, today
but with a few plugins to take care of mechanical things and the token memory space becoming increasingly bigger, writing the programming step directly on there may be the standard. as soon as vector storage and reasoning steps gets well integrated, it's going explode.
One thing I could see is a "self updating" script, that adjusts its logic on the fly when new data are found.
Unfortunately the gpt api is non deterministic, probably due to stupid censorship requirements.
Having used it for the last week or so to help me translate Python -> Rust, I'd say GPT4 is on par with even senior engineers for anything built in. The key areas it falls down on are anything involving org/repo knowledge or open source where things can be out of date.
There's a few teams working on intelligent org/repo/opensource context fetching, which I think will lead to another 10X improvement.
I agree, I avoided saying that in my original comment to avoid nit picking. The depth of knowledge is incredible. It’ll just make junior level mistakes while writing senior level code.
> The key areas it falls down on are anything involving org/repo knowledge or open source where things can be out of date.
With GPT-4, I’ll paste snippets of code after the prompt so it has that context or even documentation of the API I need it to work with. The web client is still limited in the amount of tokens it’ll accept in the prompt, but that is expected to increase at some point
More intelligently putting it in a loop & other helpers for it (which we're now seeing worked on more) should massively improve performance.
There’s still a lot of need for a human to be involved, buts it’s not hard to see how this will make individual developers more productive, especially as LLMs mature, plug themselves into interpreters, and become more capable.
You thought GUIs were bloated and slow before? Just you wait!
https://gist.github.com/int19h/4f5b98bcb9fab124d308efc19e530...
Take a close look at the SQL queries it generated to answer the question. On one hand, it's very impressive that it can write a complex query like this as part of its reasoning in response to a question like "pick a bride for your son":
SELECT name, age, clan_name FROM (SELECT * FROM people WHERE name IN ("Mesui", "Abagai", "Yana", "Korte", "Bolat", "Chagun", "Alijin", "Unagen", "Sechen")) AS potential_brides INNER JOIN clans ON potential_brides.clan_name = clans.name WHERE clans.tier = (SELECT MAX(tier) FROM clans WHERE kingdom_name = (SELECT kingdom_name FROM clans WHERE name="Oburit"))
On the other hand, this is clearly far from the optimal way to write such a query. But thing is - it works, and if there's no way to get it to write more performant queries, the problem can be "solved" by throwing more hardware at it.So, yes, I think you're right. The more code out there will be written by AI, the more it will look like this - working, but suboptimal and harder to debug.
Not to mention that humans will stop learning how to program, since they won't need to as the AIs will be doing all the coding for them.
There'll be a general de-skilling of humans, as they offload the doing of virtually everything to robots and AIs.
Then Idiocracy will have truly arrived, only it'll be even worse.
I think with LLMs we will start to see code with better and more descriptive comments, since there is significant advantage to be found when the code has as much information and context as possible.
Human programmers often hoard that for job security or out of laziness.
zipinfo -l foo.ipa "Symbols/*" | awk '{t+=$6}END{print t/1024/1024}'
I work with developers who are not CLI-fluent and wondered what they'd do. And then I realized, they might ask ChatGPT. So I asked it: "How do I figure out how much space a subdirectory of files takes up within a zip file?"It took a lot of prompting to guide it to the correct answer:
The weird part to me is how sure it is of itself even when it's completely wrong. (Not too different from humans I work with, self included, but still...) e.g.
> "The reason I moved the subdir filtering to 'awk' is that the 'zipinfo' command doesn't support filtering by subdirectory like the 'unzip' command does."
Where does it come up with that assertion and why is it so sure? Then you just tell it it's wrong and it apologizes and moves on.
Now, the addendum here is that if either of us had read the zipinfo man page more closely, we would've realized it's even simpler:
zipinfo -lt foo.ipa "Symbols/*" | tail -1
---I also tried to use it to translate about 10 lines of Ruby that was just mapping a bit of json from one shape to another to jq. It really went off into the weeds on that one and initially told me it was impossible because the Ruby was opening a file and jq has no I/O. So I told it to ignore that part and assume stdin and stdout and it got closer, but still came up with something pretty weird. The core of the jq was pretty close, but for example, it did this to iterate over an object's values:
to_entries | .value.foo
Instead of just using the object value iterator: [].foo
---One of my pet peeves is developers who never bother to learn idiomatic code for a particular language, and both of these turn out to be that sort of thing. I guess it's no worse than folks who copy/paste from StackOverflow, but I kind of hoped ChatGPT would generate things that are closer to idiomatic and more based on reference documentation but I haven't seen that to be the case yet.
---
Has anyone here tried using it for code reviews instead of writing code in the first place? How does it do with that?
It's always "Sure thing!" or "Yeah I can do that!", and it responds so quickly that you equate it to when a human responds instantly with a strong affirmative - you assume they're really confident and you're getting the right result.
If it took 5-10s on each query, had no positive framing and provided some sort of confidence score for each answer, I think we'd have different expectations of what it can do.
As for ChatGPT i have jokingly thought to always tell ChatGPT “that didn’t work” until it settles on something it is really confident on (no more alternatives) then try that.
Which- maybe that's just the thing, maybe that's what people are using this for right now
(This could change once it gets to the point of stamping out entire projects, but for now)
As far as debugging, I find that I can just say "are there any bugs in that code?" and it does a great job at finding them if they're there (even though it's what it just gave me)
The first one was to write a python script to watch a list of files given as cmd line arguments and plot their output anytime one of the files changed. It wrote a 100 line python script with nice argument parsing to include several options (like title and axes labels). It has one tiny bug in it that took a couple of minutes to fix. When I pointed out the bug it was able to fix it itself (had to do with comparing relative to abs file paths). If I wrote the script myself I would not have made something general purpose and it would have taken maybe 30 minutes to do.
The second example required no fixing and appears bug free. I asked it to write a python function to take a time trace, calculate and plot the FFT, then apply FFT low pass filtering, and then also plot the filtered time signal vs the original. This is all straight forward numpy code but I don't work with FFTs often and would have had to lookup a bunch of different API docs. Way faster.
I have also had it write some c macros for me, since complex C macros can be hard to escape properly and I'm not comfortable with the syntax. Its 100% successful there.
This is the thing: you don't need to be exact. Even very vague gets a long way.
I'm a member of the management team for a company, and everyone was very excited for ChatGPT. As someone with a technical background, the caveats were a lot more obvious to me from the start. Treating it as an incredibly fast, eager junior developer who is desperate to please and submits code without testing it has made it a lot more usable.
Really the only prompts I've gotten decent results for were asking it to write a one-off script that interacted only with a single well-documented API like Github's. It even gets that wrong sometimes by just imagining non-existent fields on Github's APIs.
I can see it being useful one day, but I still find Copilot to be much more usable even in its v1 incarnation.
Which is right?
The other day, I had a subtitles file with slightly mismatched timestamps. GPT wrote me a Python script to fix them that got 90% of the way there (and, in particular, the code had all the API calls that I needed to get it to 100%, even though this is the first time I've heard of the library it used). The whole thing took less time than finding and installing the app that would do it for me.
The catch is that you need to have a pretty good "gut feel" understanding of its limitations to figure out whether it's going to be a time saver or not before you sink too much time into making it do something right. But it is a skill that can be learned from experience (for a particular model, anyway), and I suspect that the ability to decide what to delegate and what to do yourself will be one of the key differences between junior and senior devs going forward.
The good news is 80% of what it spat out was usable. I gave up getting it to try to give me the last 20% and figured it out myself.
One thing I’ve found helpful in those cases it to tell it that it doesn’t know what it’s talking about and ask it for good keywords to search on Google.
Someday I’d love if it just gave me the relevant links itself…
Being limited the way it is - not being able to use a project's greater context, which is millions of tokens generally - is a severely limiting factor that makes it unsuitable for serious work as well. Working around this (figuring out what dependencies to inject into the prompt) is as much work as actually coding.
I can see that it's passed some sort of threshold with the greater population. It is great generating filler text. Just what it can do right now should have many applications.
How would ChatGPT help you update button texts? Isn't that just changing a text field? ChatGPT can't search your code for where the button is, and I don't see how it would help you change the text.
I used to get more non-existent APIs but not as much with GPT 4 recently.
I spent few hours asking to do few things with a well documented public but uncommon api. It sent me on a wild goose chase of functionality that API never had..(contacted developer to make sure I am not crazy). It wrote some stuff that looked VERY good, except for the fact that it interacted with things that were never there. Would have been nice if they were..
It seem to exercise wishful thinking. You can ask it to implement functionality and it will just imagine that API/system has stuff needed.
Major problem is that I for one have no clue how it is "thinking" maybe over time we will develop better understanding of it.
I dont use copilot anymore (at least for now). Just chatgpt as an alternative to google/SO
Abbreviated prompt:
dateOfBirth: date().required().typeError('Invalid Date')
})The dateOfBirth validation is incorrect because it accepts values like "11". What is the correct validation to ensure that a user has inputted a valid date?
GPTs first go:
**** dateOfBirth: date() .required() .typeError('Invalid Date') .min(new Date(1900, 0, 1), 'Date must be after 1900') .max(new Date(), 'Date cannot be in the future'),
In this example, the min method sets a minimum date of January 1, 1900, and the max method sets a maximum date of the current date. You can adjust the minimum and maximum dates as needed to fit your specific use case. ****
Me:
It still treats "11" as valid input
GPT second attempt:
***** dateOfBirth: date() .required() .typeError('Invalid Date') .test('valid-date', 'Invalid date', (value) => { if (!value) { return false; } const date = new Date(value); return !isNaN(date.getTime()); }), *****
Me:
The parameter 'value' passed into test is a date and not a string
GPT third attempt:
******
dateOfBirth: date()
.required()
.typeError('Invalid Date')
.test('valid-date', 'Invalid date', (value) => {
return value instanceof Date && value.isValid();
}),
******Me:
There is no isValid function on Date objects
GPT fourth attempt:
********
dateOfBirth: date()
.required()
.typeError('Invalid Date')
.test('valid-date', 'Invalid date', (value) => {
return value instanceof Date && !isNaN(value.getTime());
}),
********Me:
That still accepts 11 as a valid date.
GPT fifth attempt:
*********
dateOfBirth: string()
.required()
.test('valid-date', 'Invalid date', (value) => {
const dateRegex = /^(?:(?:31(\/|-|\.)(?:0?[13578]|1[02]))\1|(?:(?:29|30)(\/|-|\.)(?:0?[1,3-9]|1[0-2])\2))(?:(?:1[6-9]|[2-9]\d)?\d{2})$|^(?:29(\/|-|\.)0?2\3(?:(?:1[6-9]|[2-9]\d)?(?:0[48]|[2468][048]|[13579][26])|(?:(?:16|[2468][048]|[3579][26])00)))$|^(?:0?[1-9]|1\d|2[0-8])(\/|-|\.)(?:(?:0?[1-9])|(?:1[0-2]))\4(?:(?:1[6-9]|[2-9]\d)?\d{2})$/;
if (!dateRegex.test(value)) {
return false;
}
const date = new Date(value);
return date instanceof Date && !isNaN(date.getTime());
}),
********Which may or may not be correct. That regex is bonkers and given the errors in the previous responses I don't trust it.
For reference, I used ChatGPT to build a simple ChatGPT-clone web interface. The JavaScript it generated is literally 100% correct (not 100% good. By correct I mean it does what I describe in prompt, without syntax errors or calling non-existing functions. Code can be both correct and bad, obviously.)
I also tried to use ChatGPT to generate some Houdini vex code. It sometimes gave me a useful scaffold, but mostly just pure garbage.
As long as it can't read my entire codebase, understand it and help me with it - that's absolute horseshit. I don't spend much time writing a bunch of new code, I'm spending most of it trying to understand the heaps of legacy code in my company and then make some small tweaks/additions.
The day it can understand big code repos will truly be the end for human made code I think.
Weren't we also sure self-driving would get better rapidly?
Co-pilot has been handy for generating structs and boilerplate code.
Where I think these models are currently interesting is in discovering conventions - libraries - approaches. Like an outline, which we know is probably broken - but enough for us to squint at and then write our own version.
Kind of like browsing an OSS repository to see "how it works" for inspiration.
Someone asked ChatGPT to write an OpenFaaS function for TTS - he needed to prompt it with the template name, and even then it got the handler format wrong. But it did show us a popular library in Python. To find the same with Google - we'd have to perhaps scan through 4-5 articles.
https://twitter.com/Julian_Riss/status/1641157076092108806?s...
On the other hand, while ChatGPT is useful for some definitional boilerplate and some possibly correct listicle type of stuff, as an editor, it also guides you in a certain direction you may not want to go. I'm also used to being somewhat respectful of what humans have written, so long as it's not incorrect, even if it's not the way I would have written something.
This is a pretty cogent insight, I much prefer having a draft of, for example, an email and then completely rewriting it, rather than starting an email from scratch. I'm not sure why, but it makes it seem much more approachable to me. That being said, I have found GPT-4 code not useful for anything but boilerplates or small functions. It often gets large blocks of code so wrong that it is just much faster to write it from scratch.
What I have found useful so far is asking for a commented outline of what I'm asking for with just variable names filled in, but no implementation. It often gets that correct, and then I do the implementation correct.
And that's what sourcegraph.com is for
A small part of me was hoping more experienced folks who read this would pickup on some level of tediousness, so it looks like that landed for some folks! There is def some tedium in GPT models.
In practice, I’ll ask GPT-4 once about syntax errors. If it keeps getting them wrong I assume it doesn’t have the context it needs to solve the problem so I’ll either change up the prompt or move on to something else.
Productivity booster? Sure! End of programming jobs? eh, far from it. At the very least you need the knowledge to tell if what it is producing is good or garbage. I haven't managed to make ChatGPT tell me that it doesn't know how to do something, instead producing BS code over and over.
There were a few cases where GPT/chat interface was better but I'd say it's 95/5%.
I'm hoping copilot x will be more streamlined. If it is it's easily >100$/month
A CoPilot comment can't do that.
I had one case where I needed state management in a really simple react app. It created an entire reducer for me, from scratch. Handled all the boilerplate cases, I'd expect.
I've also had cases where it struggles to built a from button.
I.e. it does well at writing the body of a function that retrieves data from SQL and returns a particular type, and it does poorly at parsing strings of varying formats.
I'm not super surprised it does well at reducers and poorly at form buttons.
GPT really shines as an uncomplaining sidekick and I'm not sure I believe the doom and gloom about comoditizing professions yet.
I'm skeptical someone with no domain knowledge can use any of these emerging AI tools to replace someone with domain knowledge + chatGPT. The latter will just breeze through a lot of grunt work and enable more time and inspiration needed for even more value adding work.
But then again, there's plenty of examples of technology evolving and grunt work evolving in response
Juniors will be learning from GPT as well, and they might get proficient faster than the seniors did.
First I said, that might be true, we just don’t know. (Couldn’t sugarcoat things). But a better way right now to think about this is, imaging having a one on one senior programmer mentor that’s pairing with you all day long. You will learn much faster.
Similarly I’m sure younger attorneys will benefit from having unlimited access to a more senior “analyst” of contract and legal language to spot things they miss, and can learn much faster.
I expect it to have problems that are hard to spot. However, that's:
1) a useful skill to practice for code review
2) Usually faster than starting from scratch
I'm using this not to replace my work, but to enhance it. It's like having a mid-level developer sitting next to me that knows enough about everything to push me in the right direction.
I also prefer to learn by doing, and chatgpt gives me enough support in the places I’m lacking to make It much easier for me to learn about what’s going on. Infinitely patient, and just a few keystrokes away whenever I need it.
While I used to do the following, I don't do what the article says anymore, asking what the error is and asking it to fix it, since I could usually just fix it myself, and only when I can't do I ask ChatGPT. Now, I just use ChatGPT to write my boilerplate and basic structure of the project myself, after which I code on my own. This still saves me huge amounts of time.
I don't think it's particularly helpful to try to have ChatGPT do everything end to end, at least currently (this might change in 5 to 10 years). You simply get frustrated that it doesn't solve all your errors for you. It's still better for iterating on ideas ("what if I do X, show me the code. Now how about Y?") and not having to code something out yourself but still being able to see how to approach problems. Like anything, it's a tool, not a panacea.
It's interesting to see the somewhat conflicting suggestions of this book and the argument that something like ChatGPT / AI can be leveraged to solve the simple problems. We kind of discussed it a bit in the book club. For example, few programmers today worry about assembly / very low level solutions...it's just kind of back of mind / expected to work. I wonder if that will be true for programmers in the future, in relation to the foundation of a particular language. And if so, will the increased usage of AI be a positive or negative thing for the workers in the industry.?
From what I can tell, management / C-Suite seems to be pretty open to adoption and it's the developers that seem to be a bit more reluctant. Maybe that reluctance is coming from a place of concern about self preservation or maybe it's a valid concern about things like AI hiding / solving some of the fundamentals that allow one to learn the foundations of something...foundational knowledge that might make the more useful for solving the more complex problems.
I haven't really used AI...so maybe this observation is coming from a place of ignorance. It'll be interesting to see how things work out.
1. The first round of plugins that will connect CGPT to the cloud platforms in order to run code directly. Because there’s no need to copy/paste on my own machine. With cloud costs so cheap, why wouldn’t I just ask Chat to spin up running code over and over until it works? Token limits will continue to fall, as will AI costs.
2. The round of plugins that translate my questions/search into prompts that can do #1 for me. Because I don’t need a quarter-inch drill (running code), I need a quarter-inch hole (an answer to a question). Now we’re not only in the world of serverless, we’re codeless. (Can’t wait to hear people complain “but there’s code somewhere” while the business goes bust.)
3. The services that run autonomously, asking questions and gathering answers on my behalf. Because it benefits companies to know my preferences, take action to acquire goods for me, measure the results that I share, make inferences, and perform the next step without me asking, sometimes without me even knowing. As long as I’m happy and keep paying the bills. Or I should say, as long as my bank AI keeps paying for me. Now we’re in the agent cloud.
4. Then…
Or maybe I just expect too much...
I find this approach to produce the highest quality output and I can ask it to explain reasoning, improve certain things, or elaborate and expand code.
In my case, instead of a Ruby program, I asked for an outliner application similar to Workflowy with clickable bullets to be implemented in CodeMirror v6. It understands what "outliner" means and can put "clickable bullets" in the right context. However, while it gives "plausible" answers, they are all wrong. Wrong API version (often the examples use the pre-v5 version of CodeMirror). Often confuses CodeMirror with Prosemirror. Imports non-existent CodeMirror packages that are clearly imaginary, and has considerable difficulty explaining them when asked. It completely doesn't understand the concept of Decoration objects. And many other problems associated with delivering uninterpretable code.
I had better luck with Cody, who is more familiar with CodeMirror's latest API.
But as others have pointed out - GPT is great if you know exactly what you're looking for and are using it to save time typing up bunch of stuff. Great for boilerplate stuff and doing mundane stuff like writing unit test code.
To date I’ve only seen screenshots of LLM conversations, which usually get chopped up between screens and require a lot of effort to read.
The direction I chose for this article was to include most of the contents of the LLM dialog directly after a header and my prompts in a quote block. It made the structure of my article feel a bit awkward and long-winded, but we’re so early in the game I felt like I needed to include much of the dialog.
I look forward to a day where we reference LLMs with the ease of tweets, other websites, citations, etc.
I too like using Chatgpt to get my creative juices flowing, but paragraph by paragraph seems like the power dynamic has shifted.
I'd argue that no one who is aware of the ghostwriters considers the celebrities who used them writers or having written a book.
> writing a book is a daunting endeavor
Absolutely. But it's not one you're going to actually succeed on. Sure, your name will be on the cover, but you won't have written the book. As you say, you'll have edited it.
I've already written two (links are in my bio). At this point I care much less about the accolade than a book on this subject existing.
This is what I struggle with at the moment. All these pairing examples seem too start with "build me something from scratch" but in the real world 99% of my challenges are existing, large, complicated applications.
I'm curious if anyone has seen examples on pairing with GPT for existing apps and how they "compress the context," so to speak, to provide enough background on the project that fits within the context limitation and still provides useful results.
I haven't actually run the code it emitted for me (in Python), but the entire experience was genuinely mind-boggling.
I see why people have been leaning on it. Once we get project- and repo- and institutional environment-level awareness, which I assume will be soon, I can't imagine not using the assistance.
Everything is going to change. I knew that, but now I know it.
As an aside I'm so happy now that GPT-4 exists, the other day I had the VP engineering discuss GPT-4 in a meeting and was considering firing and closing down junior positions.
I'm also mixed on this but not surprised that junior positions will be less in the few years, it makes me wonder if CS grads would pursue software engineering at all if the future if the pipeline from junior to senior is eliminated.
Did anybody ask him/her what VP engineering was for now that GPT4 can do all the non-tech stuff for their senior devs as well as the junior tech stuff?
My question to you is, if your company is probably going to fire more people, but it’s impossible stop. So what advice would you have to CS grads now? Just quit. Go find a job as an electrician. This isn’t worth it?
There is no guarantee that it continues, but assuming that everything stops is ridiculous.
Unless the progress halts, it will be able to do everyone's jobs. You, me, VP or X, everyone.
In fact, GPT-4 is probably already able to do most white collar work better than humans. People somehow feel superior when they test it and realize it can't one-shot their entire project off-the-cuff without reference to documentation or being connected to any tools or having the ability to iterate at all.
I'm yet to see much white collar work it can do better than humans. Maybe it can do bits and pieces "better" (more like faster) than humans. But it requires so much oversight that I wouldn't say it can do it better than humans yet. Unless you have a counter example?
In fact, I just out of curiosity tested what 3.5turbo would say about a NPM module. It hallucinated something that suggests it wasn't in its training data.
So I tried again with GPT4, and GPT4 told me it wasn't a real NPM module (might be new enough to be after the cutoff; really like that it now appears less likely to hallucinate) but added "but assuming it exists and is designed to X, I can provide you with a general guide on how to use it." where X was inferred from the package name. I then cut and pasted the documentation, and asked it how it might implement that package.
It came up with a decent first attempt. I had the real code up in a separate window, so could compare and contrast. I asked it to expand on certain features. Asked it how it might DRY up some code. It did. It ended up with something significantly simpler and more elegant than the real thing (which is why I won't give the name - the dev in question did perfectly OK job and doesn't deserve to be hung out to dry) within about 15 minutes of prompting. In the process it taught me new things about some of the packages the original depended on, and leveraged features of that which the original code didn't take advantage of.
I think a lot of the attempts that ends up with describing it as a failure comes from basically asking it to be a magic senior dev where they don't see the back and forth on requirements etc. that goes on behind the scenes with a human senior dev. When you treat it the way a more senior developer would interact with a somewhat less experienced peer which has encyclopaedic knowledge of all of the available tooling - discussing requirements, asking questions about the chosen solution, doing quick reviews etc. - the results are amazing. It's like pair programming without the annoying, slow typing (I don't like peer programming...).
Even with the low speed of GPT4, the turnaround time is vastly better, and the outcome seems to be more reliable for a lot of aspects.
I'm curious, how does the VP of engineering figure senior developers are created? Out of thin air?
I wonder if we'll see more of this short-term thinking that will eventually create a serious shortage of experienced engineers in the future. I realize no individual executive is under any obligation to ensure the well-being of the industry at large, but if none of them care about it, they are dooming it.
I'm talking about RL in the sense of interacting with an environment during training. Isn't GPT trained by predicting the next word in a large but static piece of text?
Why not just learn to program it yourself? Why are you wasting your time trying to get chatGPT to give you the right thing? Most of the time you spent learning prompt engineering you could have spent on your actual problem.
EDIT: removed the hot take for peace.
But there are other areas where things really shine. Complex SQL is the biggest one I've found. It's often intricate, tedious, and annoying, but usually not too terribly hard to review/read/QA (at least for read queries!).
GPT turning plain language into SQL and knowing about fancy SQL capabilities that I normally have to fully google the syntax for every time I use anyway is a wonderful way of reducing friction.
There's probably no shortage of other tasks like that, though some of the more potentially-destructive ones of them (dealing with Terraform, say) that I find similarly tedious probably deserve a lot more detailed review before trusting the output.
The takeaway at the end was: ChatGPT is more useful for very senior developers. Knowing how to structure your question, what to ask for, what you don't want, actually lead to helpful collaboration with ChatGPT. Where it totally goes off the rails is when you ask broad questions, it doesn't work, then ChatGPT just starts throwing answers against the wall (try this, try that).
I get the derision towards "learning prompt engineering," and generally agree, but in this case I think you could also argue that it's understanding (through experience) how to structure thought and questions in the same way you would going into the problem on your own. I've come around to seeing that interaction with ChatGPT as a useful exercise in and of itself.
You might as well ask why bother learning your editor.
Given the small number of steps this does not look like a waste of time to me at all even if starting from the point of view of being unaware of how best to prompt it, but with some basic understandibg of the approach, most of the time you can do better.
Would be interesting to see these models return a kind of coefficient indicating how “often” a semantic context has appeared in the training data..
For a second there I thought I needed to read more articles per day in order to not miss the release of a new model!