GitHub Copilot – Lessons
medium.com
medium.com
As many have witnesses, the Copilot Chat has been just terrible. I give up on my hopes for the current wave of the AI evolution.
What is uniform across all LLM is the lack of nuance. Even when they manage to generate some domain specific characters that do make sense, there’s always lack of detail and nuance to the domain, regardless how you try to trick your prompts to retrieve something from a very specific context. I’m impressed by how much easier it’s become for me to get a digest of longer texts, but at the same time I’m disappointed by the quality of the results. It is very rare that I get what I actually ask for. It’s like talking to a mid level consultant who pretends to know everything but the output is rather questionable, and you just give up and seek to end the meeting.
Beyond the LLMs themselves, don't you think this is related to the data used for training? I mean, this echoes what you find as a developer on the wild, you can find answers to "all" simple things, much less discrete answers to complex stuff, and almost nothing for very complex stuff and/or specific programming domains.
I might instinctively know what to write, hit tab, and miss what copilot actually wrote isn't what I wanted to write. For example, it might add when you meant to subtract from a list but my quick scan didn't catch it. So you originally miss it, bump into a bug when writing tests, and spend more time fixing it than if you had just typed it out to begin with.
Hey, I'd love to hear more about why you think Copilot Chat performs poorly.
I personally had a very similar experience and eventually decided to build a solution for it while in YC (shameless plug but I wrote about the Copilot bugs that annoyed me here: https://docs.double.bot/copilot)
You still have to keep your hands on the wheel and you need driving expertise. But since writing uncommitted code has less disastrous potential consequences, it is much more usable.
I still don't believe the people who claim they are using AI to write or rewrite entire codebases. Maybe for the first version of toy projects. But I've yet to see anyone using AI to automatically write entire features that span an enterprise software codebase.
Aider wrote 58% of the code in the last release, and >40% of the previous few.
The release history page [0] plots this stat for each release over the last 12+ months. The overall trend is pretty cool, especially since Claude 3.5 Sonnet.
It’s not an enterprise code base, but it’s not a toy code base either.
Currently continue uses vector search of chunks which is just a crutch. I am not sure what copilot or Aider does, but the right context is key.
Another way to improve a coding assistant is to move away from simple paradigms to agent based Workflow that can deploy sub agents and break down tasks. All in the back autonomously while you code and then it surfaces suggestions or changes. This will get possible with increasing inference speeds (groq) and better agent frameworks.
Most devs at our company say LLMs are useless for coding and github copilot is a glorified autocomolete costing 20USD. I think the tech and ecosystem will improve a lot over time and they underestimate LLM abilities due to their bias.
I probably type less than half the code I commit, even without AI assistance?
I'm not sure this is true, Copilot Chat uses GPT4o but I've never seen a clarifying statement that the more immediate inline Copilot uses anything more than a 3 variant.
You would think that if that was the case, it would have been heavily sold.
The thing is, that if people using Copilot are assuming 4 but actually seeing the results of 3, it goes a significant way to explaining why they are sometimes underwhelmed or disappointed.
I'm happy to be corrected with a link that specifically clarifies that inline Copilot uses 4+.
Copilot, Cody, Jetbrain's AI thing. All of them are helpful tools for a small set of tasks, like a super-charged code completion tool. But nothing more than it.
I don't think anyone at our shop has high hopes for LLM's in programming beyond efficiency anymore. I'd like to see Github copilot head in a direction where it's capable of auto-updating documentation such as JSdoc when functionality changes, LLM's are already excellent at writing documentation on "good" code, but the real trick is keeping it up-to-date as things change. I know this is also a change-management issue, but in the world where I mainly work, the time to properly maintain things isn't always prioritized by the business at large. Which obviously costs the business down the line, often grievously so, but as long as "IT" has very little pull in many organisations it's also just the state of things. I'd personally love for them to get better at writing and updating tests, but so far we've been far less successful with it than the author has.
As far as efficiency and quality goes our in-house measurements point in two directions. For inexperienced developers quality has dropped with the use of LLM's, which in our house is completely down to how employees (and this is not just developers) tend trust LLM's more than they would trust search results. So much so that a lot of AI usage has basically been banned from the wider organisation by the upper decision makers because quality is important in what we do. Yes I know this is ironic when you look at how they prioritize IT in an organisation where 90% of our employees use a computer 100% of their working hours. Anyway, as far as efficiency goes there are two sides. When used as fancy auto-complete we see an increase in work output across every kind of developer, however, when used as a "sparring partner" we see a significant decrease. We don't have the resources to do a lot of pair-programming and a couple of developers might do direct sparring on computation challenges for 1-2 hours a week. They are free to do so more, and they aren't punished for it as we don't do any sort of hourly registration on work, but 1-2 hours is where it's at on average. Sometimes it'll increase if they are dealing with complex business processes or if we're on-boarding some one new.
> Copilot is very useful to scan existing code for any errors or missed edge cases
Aside from tests I think this is the one part of the article I really haven't seen in our very anecdotal testing. But maybe this is down to us still learning how to adopt it properly or a difference in coding style? Anyway, almost all of our errors aren't with the actual code but rather with a misrepresentation/misunderstanding/unreported-change-in of business logic, and this has been the area where LLM's have been the weakest for us.
In terms of what the article calls skill atrophy, this is probably why I limit my usage to snippets on steroids. I tried GPT4 more directly a few times and while it was all impressive at the start, it’s all surface level and the hallucinations are bad but incredibly subtle.