AI programming tools should be added to the Joel Test
blog.waleson.com
blog.waleson.com
So...why should they be included?
I really worry about this "often wrong" part - you only know they are wrong if you already know what you're doing. Otherwise you end up trying to use hallucinated APIs & libraries, or produce code not better than copy & pasting StackOverflow answers (which is what the AI was trained on anyway).
Code has a built-in form of easy fact checking, which makes it one of the most appropriate applications for LLMs. It's much harder to spot a hallucinated fact in a paragraph of prose than it is to spot a hallucinated API method.
The skills you most need to develop in order to take advantage of LLM assisted programming are code reading, code review, manual and automated testing and being really good at thinking of edge-cases that might not be covered.
It turns out these are important skills for being a great developer already - LLMs just force the issue on them a little more.
Or if it hallucinates a method name, it might be in a code path that goes untested. How often are people using these tools to also write comprehensive test suites?
If you don't do that, the chances LLM-generated code will introduce weird bugs are high. But the chances you yourself (or your coworkers) will introduce weird bugs are high as well - QA, code reviews and testing are important no matter what helped write the code.
So I keep wondering if we just save time by introducing more unknown bugs using GPT?
I guess this also has a lot to do with what code is written. I would be much more concerned with a system level C++ library than some JavaScript CRUD.
Gpt-4 often handles errors well. The generated code is easy to review if you understand what you asked for (if it generates tests and examples too-which it can). Etc
What about when it confabulates a module name that's been squatted by a malware distributor?
Just like if someone opens a PR against your project on GitHub you should review the dependencies they are trying to bring in as well.
I'm sure there are good use cases, but I just haven't seen anything life-changing in my experiments with various LLM products. The only thing I use day to day is IntelliCode, which saves a bit of typing here and there.
I often see GPT/Gemini propose code solutions that refer to non-existing libraries or methods inside of these libraries. Perhaps the solution is to use specialized AI for coding is more advised.
And dont use the generic ones. However these hallucinations are still facts. The issue is that humans are also prone to hallucinations; this is perhaps the most challenging aspect to solve in AI. Everything the AI says must be vetted/fact-checked.
Back to coding with AI: i noticed AI, as described by the original author, is perhaps best integrated as an assistant. It works well when given enough context and a simple task that could otherwise take hours to complete but is now in seconds. So think of:
- given this function: …
- write test cases that covers the most common inputs.
LLM’s are still long way to go to replace programmers. But at the pace it’s going, it feels scary sometimes.
If a robot could paint your house, but made three small errors, would you refuse to use it? Or would you just fix the three small errors by painting over them?
There's some kind of John Henry complex going on in this AI discussion.
I've worked a good number of hours with Claude Opus and it has never produced non-compiling code (ChatGPT 4 does that for me), but it can create quite subtle bugs, which is missed by the "just make sure it compiles"-type comments in this thread.
I do think AI fits #9. The fact that current AI tools are not meeting data security requirements are due to the market demands and maturity:
- price needs to be low to attract adopters.
- low price? These service providers will hoard data
- data needs to be collected for training
So i think long term, there will be more premium AI tools that “promise” to not collect your data. Perhaps self-hosted? Self hosting with AI is not attractive, at least not for consumers or small businesses.
People seem not to trust companies which make these promises, which is unfortunate for the industry.
In general I am concerned that LLMs will discourage innovation in programming language design - why write a better Python if GPT can just automate the tedium away?
I believe the knowledge I gained was NDA, so I'm keeping it generic.
I don't think so negatively. My bet: innovation in programming language design will emerge that will make it a lot "less necessary/helpful" to use AIs.
Just one example: Quite some programmers claim that AI take a lot of "tedium" from the programming away. But what if we could create programming languages that mostly get rid of this "tedium" (e.g. by using higher-level abstractions to abstract away the tedious, repetitive tasks)?
Why is AI treated very different than say cloud? Most companies don't have problem with putting all data in Github or AWS or Office 365, but lot of them freaks out if any AI can access the data. I don't think OpenAI/copilot enterprise plan T&C/privacy policy is very different than Github or AWS.
I suspect when any AI model will start using patents databases for training - it will be a watershed moment for what one can do with open data. Old regulations simply would not put up any meaningful fight against volume and quality of model hallucinations, that may become valuable and patentable inventions and improvements according to the same regulations.
Most encryption is either provider managed where they can easily look into data, or the customer also shares key with them(e.g. EC2 instance getting encrypted S3 data with key in EC2). I have never seen someone using AWS treating amazon as a threat.
The answer is (from my job experience): it is not reallt treated differently. There exist treaties with Microsoft/Azure for Windows, Office 365, Azure DevOps, ... But these are considered as inconvenient necessities that the company would be happy if they were not needed, thus I believe these treaties are watched like a hawk.
They think that anything you say to an LLM is instantly added to its "knowledge" of the world - it has a perfect memory and hoovers up even the slightest piece of new information that you expose to it.
That's not how these things work. But it's hard to convince people of that, especially when the training data used by the models is a closely kept secret!
I also spun up an internal chat UI[2] to replace ChatGPT so people can feel comfortable discussing proprietary data with the LLM endpoint.
The only thing that would make it more secure would be running inference engines internally, but I wouldn't have access to as good of models, and I'd need a _lot_ of hardware to match the speeds.
[1] - https://marketplace.visualstudio.com/items?itemName=AndrewBu...
[2] - https://github.com/mckaywrigley/chatbot-ui (legacy branch)
So if company will not allow me to use Copilot, that would be a negative factor from me.
I don’t really use it to do anything but tedious stuff and for searching for documentation that google will refuse to show anymore. and it provides sources so you can verify. It really does feel (to me) like the magic of google search’s majestic era, like 2010-2015. It just tends to give the correct answers at an extraordinarily high rate and can be poked and prodded in the right direction without a lot of work.
I wouldn't object to it doing fuzzy "Google Knows Best" searches for normies, but I really wish verbatim was still an option.
On the other hand, I rarely use Google these days anyway.
What kinds of questions are you asking?
For me, treating someone like a junior assist means teaching the respective person a huge lot of stuff in a very short time, so that he gets useful as fast as possible. :-)
ChatGPT-4, by comparison, seems to have gotten noticeably worse at writing code since it was first released. I have no hard data to show that, but I have a very strong feeling that it has.
Now, they're both supposedly based on the same OpenAI stuff, but I think Microsoft must be adding some "secret sauce" to Copilot. Different training data? Different system prompt?
Unfortunately, it overheats my laptop so I can't actually use it, and I primarily do support, I don't program enough at my job to justify paying for copilot.
If the CPU use was lower, I don't see why I'd ever go without.
If a prospect client or company bans it, it's a hard no from me.
I understand that might be too extreme a red line for some, but for me, life's too short to wait for laggards to catch up with the inevitable.