The process goes;
1. Write some code
2. Get AI assistance with some part
3. Check it's right, make changes
4. Write tests
5. Put up for code review
6. CI runs checks
7. Release Candidate
8. Release Testing
There are many, many chances to catch errors. If you don't have the above in place, i'd focus on that first before using AI assistance tools.
Yes, yes they are! HN is now inundated with examples and the situation is only going to get worse. People with zero understanding of code, who take hours to convert a single line from one language to another (and even then don’t care to understand the result) are shipping and selling software.
¹ Last paragraph: https://news.ycombinator.com/item?id=35133929
Crucially, those have context around them. And as you navigate them more and more you develop an intuitive sense for what is trustworthy or not. Perhaps you even find a particularly good blog which you reference first every time you want to learn something from a particular language. Or in the case of Stack Overflow, your trust in an answer is enhanced by the discussion around it: the larger the conversation, the more confident you can be of its merits and limitations.
When you get all your information from the same source and have no idea what it referenced, you lose the all the other cues regarding the validity, veracity, or usefulness of the information.
Often incorrect context, incredibly biased context (no you don't know what you want, here's a complete misdirection), or just plain outdated. So basically the same thing as ChatGPT.
People who blindly copy and paste and ship are always going to do that. Everyone else isn't. It's really that simple.
And as you learn more from different sources, you get better at identifying them.
> or just plain outdated
Which you can plainly see by looking at the date of publication. In the case of Stack Overflow, it is common that popular questions have newer answers which invalidate old ones and replace them at the top.
> People who blindly copy and paste and ship are always going to do that.
Yes, they will. You seem to agree that is bad. So wouldn’t it follow that it is also bad that the pool of people doing that is now increasing at a much larger rate?
Without a solution we're just whining about bad actors existing.
Of course they do.
Ship fast, break things, look cool, cash out, leave someone else to fix the mess you have created, it's the new black.
Even if they're not, I find "scouring code I didn't write for potential errors of any magnitude" to be much harder than "writing code". I admit there's a sweet spot where errors would be easy to spot or where the AI is getting you unstuck, but it's not trustworthy enough at the moment for me to take the risk.
The you test it manually with a few more inputs, then you write some automated tests for it (maybe with GPT's assistance).
Why not? And if I miss them, my editor, compiler or general testing should pick that up.
The second shortcoming is that that I have to switch over to ChatGPT and it's messy to give it my existing code when it's more than just toy code. It would be a lot more effortless if it was integrated like Copilot (if we ignore the fact that this means sending all your code to OpenAI...).
Still, it's great for boilerplate, general algorithms, data translanslation (for small amounts of data). It's a great tool when exploring.
Edit: Hmm, maybe it can add uncertainty markers if I just ask...
--
Building a bridge over a lake using only toothpicks would be extremely challenging due to the limited strength of toothpicks. (70%) However, it is possible to build a toothpick bridge by using a truss structure. A truss structure involves using triangles to distribute the weight of the bridge evenly across the structure. (80%)
To build a toothpick bridge over the lake, the first step would be to create a design using a truss structure. The design should take into consideration the width of the bridge, the width of the lake, and the strength of toothpicks. (90%) The toothpicks should be laid in layers to increase their strength. (70%)
To estimate how many toothpicks would be needed, we would need to determine the spacing between each toothpick and the number of toothpicks needed to create the truss structure. The number of toothpicks required would also depend on the thickness and quality of the toothpicks used. (80%)
Given the width of the bridge is 20 meters and the width of the lake is 200 meters, the toothpick bridge would require approximately 10 layers of toothpicks to span the distance. However, without a detailed design, it is impossible to estimate the exact number of toothpicks needed to build the bridge. (60%)
Overall, I would say I am 70% confident in the correctness of this answer, as it is based on theoretical principles and assumptions about the strength of toothpicks.
--
It's... okay, not great. The blatantly wrong part is marked 60%, which is the lowest certainty it assigned to anything, but that's still really high for how wrong it is.
Bad developers don't even know if the code they write themself is close to correct. AI doesn't make that situation any worse. It actually improves the situation.
But out of curiosity, to give it something harder, I asked ChatGPT w/GPT3.5 to write me an interrupt handler for a simple raster effect for the Commodore 64 in 6502 assembly, and it got tantalisingly close while being oh-so-wrong in relatively basic ways which suggests that it hasn't "grasped" how to handle control flow in a language without clearly delineated units.
GPT4 appears to have gotten it right (it's ~30 years since I've done this myself), though the code it wrote was a weird mix of hacks to save the odd cycle followed by blatant cycle wasting stuff that suggest it's still not seen quite enough "proper" C64 demo code.
I often paste chunks of code into it just to get a detailed line-by-line explanation of exactly what that code is doing. It's really helpful, especially for convoluted code that you're trying to get a good feel for.
In fact I read every piece of code I write, right after I write it, and probably more times after. It's a good practice, because as an human, my first take at a piece of code is often subtly wrong, or correct but missing important edge cases.
Speaking of, I didn't deliberately insert that typo in the above paragraph, but I did notice it when I read this post before submitting it, and would normally have corrected it.
When I've asked for code it has been very wrong, the sweet spot appears to be things that you don't know off hand but can verify easily.