GitHub's Copilot lies about its documentation. Why would I trust it with my code
shkspr.mobi
shkspr.mobi
Today, I sense a similar mindset among those who oppose AI assistants in software development. They tend to focus on the perceived limitations of these tools, often failing to see the bigger picture of how AI can enhance productivity. Instead, they quickly latch onto any shortcoming to justify their resistance, missing the potential long-term benefits.
Whenever some new technology is being sold, someone always tries to delegitimize any criticism by citing some superficially similar criticism of some past technology that was adopted. Or, fairly frequently, they just reflexively cite Socrates's/Plato's OG critique of writing.
> Today, I sense a similar mindset among those who oppose AI assistants in software development. They tend to focus on the perceived limitations of these tools, often failing to see the bigger picture of how AI can enhance productivity. Instead, they quickly latch onto any shortcoming to justify their resistance, missing the potential long-term benefits.
So your message is: "Who are you going to believe, the sales pitch or your own lying eyes? You should focus on selling yourself on the new thing."
I think you should listen to the criticism, and try to understand and empathize with it instead of trying to dismiss and disregard it. There's something there. Maybe the people who critique it are the careful and thoughtful developers, and the ones who sing its praises are the sloppy ones who aren't careful enough to see the problems or just don't care (e.g. the kind who see their job as closing tickets not writing working software)?
Maybe this thing will get adopted despite its shortcomings and change everything, like NFTs did, but that doesn't mean those shortcomings aren't important.
I think there might be a few things to consider here:
1. Is it without any loss in productivity? Or are they so productive/skilled that they can still outpace most other devs even without the IDE's additional functionality?
2. Is there a loss in productivity, and they just don't realize it? Just because they can still do a good enough job without an IDE doesn't necessarily mean that they couldn't do a better job as experienced IDE users.
This is similar to the arguments people make against using Rust. "I can code C just fine and never have any issues, I don't need Rust"
It's also similar to this Copilot discussion. The people who say "I don't see any benefit from LLMs when coding", could it be that they just... don't have enough experience using LLMs? We've -just- had that come up in HN[0], where everybody in the comments was showing how they got the answer just fine out of their LLM of choice, while the OP was all about how AI couldn't possibly understand that bit of code.
But I guess you're right that we can't trust word of mouth and don't know for sure until/unless we have an agreed measure of productivity and can run with/without or before/after type experiments.
This is a fact that needs more attention, and it's why I prefer to onboard people with setup documents they work through rather than a canned working development setup, even though the former takes "longer" to get people up and running. The long term benefits in forcing people to touch and familiarize themselves with as much as possible far outweigh any short-term time gain.
Give competent developers - i.e. people trained to solve optimization problems in their sleep - a metric to optimize for. Call it "productivity". See what happens.
I can't think of a worse idea than giving copilot to a junior dev... I use chatgpt a lot, but I've got lots of experience. One has to understand how the big picture works very well to not find yourself led down the wrong path.
Most non-technical business stakeholders don't really understand what we do. You don't see shows about awesome devs doing awesome shit, like you do for lawyers, doctors, business people, etc. All they know about us is that we're nerds who work like two hours a day and can't be bothered to go into the office.
They do, however, understand ChatGPT and how it can be used to create whatever you want (see also: Eric Schmidt's missive to Stanford business undergrads that made the front page here some time back).
Given this, they are EXTREMELY incentivized to give/force upon ChatGPT to juniors. If they can give this to juniors and maintain feature velocity with an acceptable amount of quality degradation, then they can eventually move software development to lower-cost countries with weaker worker protections wholesale (and use good old fashioned protectionism to make sure that those countries don't try to steal our slice of the pie).
Quality issues can be fixed by expert consultants who know their shit, but that market won't be big enough for all of us.
You do, but it's even more ridiculous nonsense fantasy than a typical doctor or a lawyer show (e.g. all the hacker girl scenes like https://www.youtube.com/watch?v=hkDD03yeLnU, https://www.youtube.com/watch?v=u8qgehH3kEQ).
But there's a key difference here.
I don't remember an IDE trying to autocomplete something which didn't exist. Similarly, an IDE never pointed me to documentation which it had just invented. If it had, I'd've been one of those people living out my days in Nano.
Why would it be any different with "AI"?
Also, Vim (and Emacs) and notepad are not even in the same class of editor.
IDEs made it easier to navigate a huge codebase with things like autocomplete and one-click refactor. This made it easier to make huge codebases that _require_ these QoL add-ons over time. I think this was a factor that made outsourcing software development (to countries with weak or non-existent worker rights) possible at scale.
In a similar vein, highly paid software engineers are using LLMs to generate more and more of the code they ship.
If highly paid SWEs can write code by shoveling a sentence into Copilot, _literally anyone can._
We seem hellbent on destroying our profession for the sake of "productivity" (i.e. using less of our brainpower to churn our features faster).
But what do I know? I'm just a simpleton that uses vim and writes code the old-fashioned way.
Whether you scale that strategy once your business becomes viable is dependent on whether you give a shit about software quality or not. That's the long-term concern I have with this flavor of AI.
And configuring tools through IDEs is very opaque. How do I add a flag to the compiler command line in this IDE? Who knows. How do I change the CMake config? Maybe I can edit the CMakeLists directly, or maybe that will make the IDE blow up, so I have to use some arcane GUI built atop an already arcane text UI.
Plus, you can usually forget about scripting your workflow. Unless the exact workflow you need is built in to the IDE already, you'll be clicking the same sequence of buttons over and over ad infinitum. An IDE is one big complex tool in place of a handful of simpler composable tools, and the only way to extend an IDEs functionality is to make it more complex instead of adding another simple tool or composing existing tools in a different way.
That said, I would keeo this thing the hell away fro junior engineers and I wouldn't want a junior who has used AI his entire career anywhere near my codebase. AI is gasoline and without someone with experience behind the scenes, they are just going to be 10x faster at knocking out unmaintainable spagetti code.
When you run a saas company, code is only ONE factor among many you are maintaining. The other arguable more important one is the domain specific knowledge that exists in the heads of your talent. Give me the nerd that uses emacs with 5+ years of experience thinking though and solving problems without AI. I'll let him have copilot because he'll be the one writing and maintaining the code. More importantly, he'll have teh experience to know when copilot is leading him astray.
. . . Except it got one so wrong I couldn’t trust it and had to be just a meticulous as I would’ve been in the first place, so that was a net loss of time.
I do not see how it could fit into my workflow other than how I use it today — a google expert I can talk to about anything.
What is your workflow like?
In addition to Copilot, I frequently use ChatGPT to get quick answers or examples, even beyond Laravel's already excellent documentation. It’s also incredibly useful for quickly learning the basics of setting up the infrastructure for my projects, saving me time and effort in the process.
This is the core problem right here.
If the industry wasn't so inefficient then ai wouldn't be so effective.
Runner up is copilot integration in the IDE that I ask to transform data or reorganize lines of code. It's grunt work and having smart automation around it is helpful.
Anything requiring "thinking" generally falls over pretty quickly. For example I've found broader feature code generation has generally been crap or worse and often using APIs that just don't exist. Catch is that LLMs have been sold as thinking machines that will replace people and VCs have thrown all their money into that fire.
P.S. VIM has a growning following and vscode user counts are astronomical, so I'm not sure your point stands.
This is the use case that I struggle with the most.
Let's say that I want to know how to reheat a baguette. If I throw that into ChatGPT, it will give me an answer based on...who actually knows? It could be from the set of responses it found most often in its training. It could be from reddit.
It could be mashing these up into something completely incorrect.
Perplexity and Kagi Assistant are better in that they give you some of the sources they used in generating an answer, but now we're back to a search engine with extra steps.
With a search engine, I can go through each source and know (a) that the answers aren't altered and (b) who wrote the answer and whether I should trust them.
To wit: I once used ChatGPT 4 to find me good restaurants in Houston based on /r/houston subreddit, with sources. (This is the best way to find new restaurants in my opinion.) It gave me five recommendations that (a) were not the best recs and (b) had fake Reddit URLs for sources.
How can anyone trust that?
Also, neither vim nor emacs are simple text editors.
Quite frankly, IDEs at the time were really bad. They were slow and clunky. Intellisense didn't always work. People who relied too much on IDEs didn't understand their project's build process and couldn't debug it. The code templating stuff spit out awful code.
I kept using Emacs for a long time because I was more productive with it. I had a lot of customizations and add-ons that made it a dream to browse the codebase and churn out good code.
I finally switched to VSCode some time after it was open sourced because the plugins for my languages had gotten really good, and emacs development felt like it had slowed down and everything was starting to get bitrotten.
As for IDEs in general, they only propose functions that do exist.
Aside from the fact that it's not all that comparable, it's just flat-out wrong in its description of events.
This is the biggest misconception people have about AI. You shouldn't trust a Copilot by default, but more often than not it will be very handy if you know what you do.
It's not like GH tells you to use Copilot exclusively for writing boilerplate code. What the OP tried to do was a legitimate use case, wasn't it?
Because i don't know what i do but have enough experience to know if the Copilot knows what it's doing.
Or sometimes I use it for languages which I'm not god at, but still can read it. I'm a 10+ year Java programmer. But sometimes I want some small bash / python / powershell script.
For example a powershell script to change power state profile. It spits something out in seconds, I can read it, understand it and ask for changes.
With a large java project I don't use it much. But sometimes, and then I also read the code the check if it does what I want to happen. Sometimes I can just describe a problem in English, copilot will create something that I think is wrong, but it gives me an idea on how to solve it myself.
Then copilot was blocked at work. ¯\_(ツ)_/¯
If I ask it something it doesn't know, it says "Hmm. I don't know that."
So, yeah, it is surprising that this supposedly advanced AI doesn't have the same level of functionality as a smart speaker.
To each there own, I have been using these tools for a while and have a grasp of what is possible and not.
…you have to admit that this is just outrageously shit.
Come on
You’re sending a response from a help bot to a user with a url in it and you didn’t bother to add a check to make sure the links it returns actually exist.
That’s just lazy. I wouldn’t want to use it either.
I don’t.
Or any product where the developers have been in such a rush to pour AI into their products that they’ve failed to integrate it in a way that is useful and meaningful.
It doesn’t seem like such a big ask.
If you’re integrating AI and you’re not doing the most basic output sanitation… you deserve responses like this.
From github? With their resources and experience? It’s a fuxking disgrace.
The burden of proof is on the new generation, not on old timers. And this is coming from a new generation software engineer.
I will some day in the near future migrate my projects off of GitHub. It was nice having a centralized place for open source projects for a while, though; a lot more people saw stuff I was working on than otherwise would if I just had a cgit instance running somewhere. But that's okay: the point of me publishing it wasn't for other people to see, necessarily.
And I'm sure eventually there will be another anyway. I certainly remember how Sourceforge went downhill.
With code it does the same thing and its prediction skills aren't tuned by humans perfectly yet.
As Open Ai notes here are the limitations of ChatGPT
Limitations
ChatGPT sometimes writes plausible-sounding but incorrect or nonsensical answers. Fixing this issue is challenging, as: (1) during RL training, there’s currently no source of truth; (2) training the model to be more cautious causes it to decline questions that it can answer correctly; and (3) supervised training misleads the model because the ideal answer depends on what the model knows(opens in a new window), rather than what the human demonstrator knows.
ChatGPT is sensitive to tweaks to the input phrasing or attempting the same prompt multiple times. For example, given one phrasing of a question, the model can claim to not know the answer, but given a slight rephrase, can answer correctly.
The model is often excessively verbose and overuses certain phrases, such as restating that it’s a language model trained by OpenAI. These issues arise from biases in the training data (trainers prefer longer answers that look more comprehensive) and well-known over-optimization issues.1, 2 Ideally, the model would ask clarifying questions when the user provided an ambiguous query. Instead, our current models usually guess what the user intended.
While we’ve made efforts to make the model refuse inappropriate requests, it will sometimes respond to harmful instructions or exhibit biased behavior. We’re using the Moderation API to warn or block certain types of unsafe content, but we expect it to have some false negatives and positives for now. We’re eager to collect user feedback to aid our ongoing work to improve this system.
Um, no. Not so for me. No copilot at all. Guess we shouldn't trust you either ;)
A key skill is learning when to say "I'm sorry, I don't know that."
AI is currently trained to be overly friendly and helpful - to the point of deception. They often "hallucinate" entire software packages. At which point, they become worse than useless.
And is it wise to judge one tool’s quality based on a totally different tool’s ability in a different domain?
Honestly you’re mostly demonstrating why people aren’t reliable.
If you tolerate broken tools, that's fine, more power to you. I don't. If shit doesn't work it goes in the bin.