Re: I Don't Use Copilot
vivekhaldar.com
vivekhaldar.com
When a code suggestion is correct it’s nice but it honestly feels like a very rare occurrence so far.
What often ends up happening is I will get suggested a snippet that is close to, but not exactly what I need. So if I accept the suggestion, I end up doing a fair amount of editing anyways to correct the suggestion.
Even worse, occasionally I will accept a suggestion and only later notice a subtle mistake.
The place I’ve seen the biggest benefit is with writing comments or user facing text copy.
I see people on Twitter and Hacker News talking about how much their productivity is being boosted by Copilot and ChatGPT…
But I’m just left scratching my head because I’m barely certain I’m seeing ANY boost from either of these tools so far. I feel like I’m missing something?
Same, I got burned one time by accepting a suggestion I thought I understood.
I think the general problem is that Copilot shifts you from writing code to reading code, and sometimes reading is harder. You can't really take a "yeah that seems right" attitude because it's just throwing guesses at you and seeing what sticks. The safe way to use it is as a jumping off point for writing your own code.
IMO it’s a nice 10-15% speed boost for web dev. And much more than that for unit tests. I expect you can measure a massive jump in test coverage by engineers that use copilot because it takes work that most people hate and speeds it up significantly.
Most code on github is by definition average, which is usually a few levels below the quality of an expert programmer. So it's not surprising the output of code LLMs is poor to average quality. For any larger block of code it outputs, a non-trivial amount of editing is required if a higher bar of quality is wanted.
I'm a crappy marketing copy writer. A professional writer could do much better. But with chatGPT I too can write hack-quality marketing-copy that roughly conveys the a message and give it others to put in various marketing outlets.
By definition, half the code out there is of below median quality. Whether or not half the code out there is of quality below the arithmetic mean depends on assumptions about the distribution of code quality. I would suspect that much more than half the code out there is "below average", so to speak.
This is what hits me hard, as when I write it myself these things are more clear.
> I see people on Twitter and Hacker News talking about how much their productivity is being boosted by Copilot and ChatGPT…
I use ChatGPT quite a lot, but not really for coding. It does increase my productivity, but only because it better acts as a fuzzy search/query machine than Google. Often I use the two together (which leads to using Bard). I can remember the concept, but not the name, so GPT/Bard tells me the name, and then I can google from there. With this type of searching the hallucinations are not a big issue as they are only minor interruptions and are far outweighed just by the amount of searching I'd need to do to even find the thing I need in the first place. I should mention that I'm a researcher so I'm often looking for new tools, ways things have been done before, and frequently probing for things that are just outside my area of knowledge. But I should also say that I won't trust them to teach me a concept because when I've tested it on things I know, there are serious mistakes. Realistically the accuracy depends on how common/frequent the information is. If it is something that many people write about in a general field, it'll be accurate. If it is a topic, even popular, of a niche subfield, good luck.
More than once has it shoved Python code into my Ruby file. Bash or YML files will often get comments added, but the wrong style for the language (which since I can never freakin' remember the style for every language actually makes me less productive since I have run the code, have it crash, and still go look up the language spec).
It keeps trying to autocomplete large hash manipulations, and it looks right, but subtly wrong in the syntax (missing commas or something) and then it's the same compile, crash, then stare at the code and have to add the same damn change for 50 lines of a hash when I should have just copied and pasted in the first place.
So far: net negative on my productivity. Plus, it sucks on shitty wifi like a plane or most hotels.
Without knowing the specific statistic being reported I’m not sure you can reach that conclusion. Copilots plug-ins can be configured to suggest as you type just like any other auto complete. It may just be that the user has paused and doesn’t need the suggestion being offered.
And it’s not like people use 100% of the suggestions their auto complete tools provide, but that doesn’t make those suggestions garbage.
I haven't tried copilot yet, but chatgpt code generation is probably useful to me 1/3 of time as a stack-overflow/reference-doc replacement. "I need a thing foo to perform bar in context baz" which I then will right myself using the pattern. But the lookup was way faster than googling tons of results.
From what I understand (and I might be totally wrong),
Copilot is useless for dealing with codebases where
you often need to call internal functions
Perhaps this is the next big frontier or opportunity.LLMs have finally gotten to the point where they are sometimes useful. And local project-level code intelligence (intellisense, static analysis, etc) has been pretty good for a while.
The first team to marry these two things is going to really have something great. A true force multiplier.
It certainly doesn't seem like it will be easy, but it also certainly doesn't seem to be impossible.
It's quite excellent even for large projects. It still saves a lot of mundane coding
Oh, but it will learn that you're driving to work and will be able to prioritize that above the airport and the arena? Oh yay, so when I drive to someplace I drive to 5 times a week it can give me directions I don't need, but when I drive somewhere I haven't been before it can only tell me how to go to a bunch of places other people went on the way.
It's practically useless if your knowledge of your codebase and language is better than 0.
It's probabbly useful for languages and crap you don't use a lot or where you'd end up copypasting 90% of it anyway, like crappy webdev. (Good webdev is a different story)
I'd rather use it to explain what the fuck I was thinking when I wrote this shit 6 months ago, or what Nate was thinking when he seemingly fucked it all up (or maybe he fixed it). And maybe, maybe, playing the analogy back in reverse, it will get to where I tell it I'm trying to go somewhere specific and it gives me turn by turn directions, instead of arbitrarily suggesting shit that I drive past. Like I could say I'm trying to extend this to accept triangles and not just squares and it would say, oh use the visitor pattern and here are the places to change it, or something. But right now it would just flip it's shit and start suggesting triangle, polygon, stars, circles or something like that. Do you want to reverse a linked list of triangles? No, clippy. I know better than you how to change this shitty code. Thanks anyway.
Some things I've done this week:
- I was asked to write a white paper on a topic I had previously built a slide presentation and given a talk on. I exported the outline, grabbed a few links to blog posts on the broader topic, and asked ChatGPT (with link-crawling) to write the whitepaper. I didn't like the output at all. So I asked chatgpt for several options on the outline for the whitepaper, and several options for narrative. I think selected an overall narrative arch for the paper, and the outline (with a few manual tweaks) and asked it to rewrite the paper with a revised prompt. It was pretty good and then I asked to change the style and tone. I then asked it to critique itself and suggest several improvements, several of which I then asked it to take into account and revise the paper. Then I asked to several variations of the document for different audiences. I picked out a few different parts I liked from each variant and ran another pass asking for improvements, including suggestions for visual aids. This all only took about 10 minutes, after which I shared this first draft with a colleague for review. That would have taken me a whole afternoon (and likely longer since I procrastinate when starting a blank page)
- I wanted a chrome plugin to use the archive.ph api for gated papers without needing to trust that sketchy extensions might steal my web history. I asked chatgpt to generate the code. Upon review I made a couple tweaks and had a private extension in about 3 minutes.
- I wanted to produce a menu for guests that staying with us this weekend. There were several dietary restrictions and preferences that need to be accounted for. I entered these in requesting a variety options. I selected those options, then asked for recipes and a suggested shopping list. I then asked for wine and beer pairings and got some good generic responses of varieties - but the brand-specific suggestions were bad. Time 2 minutes
- I wanted to some pro/con brainstorms on a new technology decision - I asked it to search for expert blogs and summarize their arguments. I also got it to generate hello-world level examples for each to compare and contrast. I then had it generate code for a more complex use case using the tool for each framework. I was then able to make a much more informed decision and start evaluating deeper a smaller set of options 2 from the dozen+ that existed. Time - 20 minutes
All of these have a pattern of iteratively working with a tool to semi-automate what I want. For now it sort hits a sweet spot of either summarizing research I would normally google and low-key generation of document or code that doesn't need a tone of context outside either what I researched or have on hand and can be self-contained.
Maybe I'm one of the hangers on who refuses to use it, not out of IP concerns but because I'd hate for it to cause my skills to rot or for me to become too dependent on it. I do sometimes have GPT-4 write code for me, but somehow I feel like it's different reading and reviewing its code than to punch a few characters into my IDE and have code magically appear.
For example, I gave Copilot something like this:
const textFile = longfile
const numLines =
I expected it to give me a regular expression that would count the number of new line characters, but instead it split the string on \n and asked for the length of that array. These are decent sized files (300ish lines) and this code is running in a hot loop over thousands of them, but I decided to give it a chance and benchmarked both choices.It turns out that even with all of the extra allocations, the regular expression solution was 4-5x slower than splitting and getting the length. Not something I would have guessed, but now I know.
I'm patching in a loop now (though I feel like a regular for loop is more readable than a loop with indexOf).
In my case, after benchmarking inside my actual app on real data, all three perform equally well (including the String.split version). Dropping the Regex made the difference between this being the heavy part of the process to it being a rounding error, and there isn't much additional value to be gained. With that in mind, I'd rather go for legibility over speed.
I'm going to go with the for loop, though, because the extra memory usage could come back to bite me later.
I'm coming from Kotlin where I expect there to be a standard library function for simple tasks like this, but I clearly need to get back in the habit of writing loops by hand.
I found it to just be… not how I wanted to work. I’m not sure it ever made me faster? Like it typed things for me, so there’s some savings there, but because I wasn’t precisely sure what all of this magic code does before I read thru it and renamed everything, it just felt like I was the copilot, reviewing the others work.
It sure is cool but not how I want to work.
I need to write a cron expression like once in couple years. So, I don't have working memory of it. So, to write it, I need to
1. Google "cron run something periodically".
2. Filter spam such as geeksforgeeks and irrelevant stackoverflows
3. Find either good enough approximation
4. If didn't work, go to cron man and start reading it. Example that I look for is 3rd from below in examples table https://docs.oracle.com/cd/E12058_01/doc/doc.1014/e12030/cro...
OR
I can go to chatgpt/copilot and directly ask what I need
"Tell me a cron expression to run command on each third Monday of a month and explain it"
and it will give me good enough approximation right from the start. I won't need to fight popups that will shame me for using adblockers, I won't need to go through tons of spam. I will get much better headstart, that will save my time and energy.
ChatGPT responds:
> 0 0 0 ? * MON#3 *
Regarding MON#3, it explains: "Day of the week (MON#3): This is the crucial part of the expression. We use MON to specify that we want the command to run on a Monday. Then we append #3 to indicate that we want the third occurrence of Monday. This ensures that the command will run on the third Monday of the month."
It does not mention though, even with the smallest hint, that standard cron does not support this syntax. Or a 7-field syntax.
(By "standard cron" I mean Vixie cron derivatives like the one that ships with Debian)
The correct answer in this case would have been: "This is not possible with standard cron syntax alone. You can either use a different cron implementation that supports MON#3, or schedule for every Monday, but then in your script check if it's the third Monday."
I'd rather use something like https://crontab.guru/ to verify my syntax, thus I have it bookmarked for the 1-2 times a year it comes up.
An LLM is probably the only thing which could tell you this.
> Recent LLMs are fine in terms of smaller things like cron
This is obviously false given the contents of this thread.
Indeed, SO (well, SU) got this right: https://superuser.com/questions/428807/run-a-cron-job-on-the...
And crontab(5) explicitly calls out why this is impossible:
Note: The day of a command's execution can be specified by two fields —
day of month, and day of week. If both fields are restricted (ie,
aren't *), the command will be run when either field matchesNot in my experience. GPT4 works a lot better than previous LLMs. Have you used GPT4 yet? Typically in my experience people who talk about LLMs not being that useful (or "pathological liars") have not used them recently, or at all. So, I mean, not a useful opinion worth considering then.
The key thing to me is it's nice for certain kinds of work. Loosely/offhand, those categories:
- Writing stuff from scratch: Sometimes useful. Particularly if it's a "write it, get it working, forget it" script.
- "Polishing" code with close attention to detail/docs/implementation: This is an area where I'm more likely to immerse myself in understanding the code, and wouldn't be likely to use Copilot, as it would break the immersion.
- Writing accessory scripts: E.g. a sqlite alembic migration script; I really am not a DB guy, so outlined some steps in a comment, read what Copilot wrote, tested it out, called it a day.
- Adding stuff in an existing codebase: In cases where there's hundreds of functions across dozens of files, the idea of copilot knowing the functions and suggesting them for me, without my needing to look them up, sounds great. I'd definitely like to attempt this with work's codebase someday. I think a lot of precursor tools could do much of this; but I just never got in the habit of needing them.
I spend more time with ChatGPT currently, but I'm pretty sure I'm nowhere near being a heavy user there, either. I think 75% of my peers are basically doing nothing with AI tools currently. (It helps that I do side projects a lot more than they do, too...)
I think the big difference is what type of programming people do. LLMs are great at handing you code that is routine, but I find them problematic when I try to do anything even a bit uncommon. I spend more time debugging than it would take to write it myself. I think a lot of this confusion is coming down to treating coding as coding and not looking at tasks. The same way writing a fantasy novel doesn't make you good at writing poetry, or a legal document, or a technical report.
I recently started learning Go and I felt like I hated it, I couldn't make any mistakes or think about the syntax because co-pilot was constantly trying to impress me.
One thing I really hate is it will always try to auto-complete my comments. Sorry robot, I am a divine being and these comment ramblings are mine alone.
You think it's bad now, imagine 5 years from now. This version of Copilot is just the beginning. Will you care about skill rot when everyone else around you knows the skills to effectively use the new generative AI models and is producing 5x?
I'd hate for it to cause my skills to rot
Well, good news: it's nowhere near good or useful enough for that. Currently it's mostly helpful with boilerplate stuff and coding quiz style challenges.It doesn't know anything about which algorithm might be appropriate for a novel challenge, it doesn't know anything about your database or codebase, etc.
My personal opinion so far is that this is going to be a very useful tool. A big leap will occur if/when it can work in tandem with knowledge of your local code+data. However I don't see it ever replacing higher level reasoning or rotting one's skills. Unless the "skill" in question is memorizing boilerplate stuff.
GPT writes larger blocks with more context, on demand. It's quite slow and ignores certain commands though which is irritating.
There are cases though where it came in clutch and implemented a solution much better than I could by leveraging some very hidden numpy function call that I couldn't find even if my life depended on it.
Nope. For what it's worth, I use Clojure(Script) these days, and working with the REPL is my preferred way to interactively "chat" with my program.
I asked family member how they use AI for programming. He said its mostly ChatGPT for soft skills like emails or JIRA tickets, or rubber ducking a design problem. He doesn't use any AI tooling for development. It seems both authors recognize this as well.
I feel you missed something here.
ChatGPT: reading and reviewing its code
Copilot: punch a few characters into my IDE and have code magically appear... and you read and review the magically appeared code, just like you do with ChatGPT?
I think combining or crafting your offboard brain/prompts at this point is par for the course, or maybe even a no brainer. Take your knowledge, and things you don't like or don't pique your interest, give it to the robots. That's what I've been doing so I can focus more on the critical path. I don't trust any of these robots, but not everything they produce is bad, some of it is actually pretty good.
I see this mentality a lot, and I don't understand it, but people seem to take satisfaction from doing things the hardest way. While I can appreciate the feeling of accomplishment, and being able to overcome an obstacle while learning, if there is a solution that takes less time to build and create while providing the same outcome, that tool is going to have a home in my tool chest.
If Copilot/GPT produces stuff with subtle problems, and it always will, then it will not be better code than someone with skills.
More testing? Generated by the same LLM? Sure, but it might be testing the wrong thing.
I think that those with skills should lean on them. Let code monkeys use LLM's; experts should be human, not a fleshy bot feeding another bot and uncritically using the spew.
Sure, but humans could be testing the wrong thing too.
They should, but let's not pretend that experts or any human for that matter are infallible.
I wouldn't refer to people as code monkeys, it's rude. I can see on the site listed in your profile you're having issues aligning your navigation <div> (left or center?). Like you I too make mistakes. In this case ChatGPT recommend using a <container>, while I'm not sure if it's the right path forward, it seems to fix the issue on your page keeping the <nav> left aligned.
My navigation is where I want it to be. It's not a mistake. Nice try.
Humans may be infallible, but if they wrote the code, they'll have an easier time figuring out where their assumptions were wrong because at least there were assumptions. An LLM just regurgitates statistically-massaged word lists.
Your navigation looks like utter trash. I've seen high school graduates do considerably better. It is a mistake. Just own it.
You've taken a stance based on regurgitated talking points. Do you get upset with yourself when you learn something from stackoverflow?
I'd like to think that we should all want to formulate our questions well.
Having someone(something) complete our questions is like having someone interrupt you half-way through a thought and finish your sentence for you.
For the one completing the thought, they think they're being helpful. For the one being interrupted, there's a decent chance it's a <wtf?> moment.
Being able to form a complete thought by yourself is something that's important. Don't throw this away....
I'm sometimes asked to explain a piece of my code.
If I told my boss, "I dunno. ChatGPT wrote that part, and to me it looks like it works OK" I'd probably be fired.
At some level somebody with engineering knowledge is going to have to understand the problem domain, the code, the data, the use cases, ancillary data stores, etc.
So it is more like it is predicting what I already wanted to type, instead of being something like pair programming and giving snippets.
Outside of Leetcode-style stuff it's never going to produce reams of correct code in non-trivial projects with so little supervision that this is a major problem.
A similar thing happened with me using Adobe products at a student discount, only to realize its hefty $60 per month price tag was not something I was willing to spend.
And if the gained productivity is only marginal at best, it’s better to not grow dependent on these technologies.
This puts you closer to institutionalized domestic violence. Give up your existing career and become dependent on [service provider]; you won't ever need to be able to provide for yourself because my behavior will never change for the worse, right? No provider has ever decided to become withholding, vindictive or capricious, right?
Your lighter won't do this to you.
^ I'd say that's a pretty good litmus test of whether you should let the skill rot.
Somewhere along my family tree we didn't pass this skill forward and it was lost.
Personally I see it as a rite of passage for a senior developer to take note of and embrace their own limits.
Recently started doing this in a for me new (but old) language: Python. Copilot is incredibly helpful in learning the ideosyncracies of it and its libraries. And the suggestions do actually keep getting better and better. Been using it for about 3 months now.
At least for Python I deeply recommend it. It accellerates the learning of a new language.
If I’m using an unfamiliar library, I’ll type a comment on what I want to do then start attempting to use the library. Copilot is amazing at getting a close enough solution. I can often skip reading a lot of irrelevant documentation.
Most of my current work is creating new React components. In these tasks, Copilot probably saves me 20 or 30 minutes a day on average. Often more. It's not always right, but often close enough, and I'm not committing my code without testing it anyway.
A simple example. If I’m using a library, like datefns, I can describe what I’m looking to do, copilot will find the correct method for me, but might reference the wrong local variable. This is even more helpful when a library has a bunch of similar APIs (parse, parseUTC, parseISO, etc, etc).
It also really helps with the under documented prep work you might need to do. Some libraries have really poor default configurations and setup documentation. They require non-obvious setup that can be hard to track down in documentation. Copilot does an amazing job of getting you in the ballpark. Even if it’s missing settings, it become incredibly clear how a library intends for me to use it.
> But the exact same argument can be made of every advance that has continously raised the level of abstraction of programming over the decades.
Languages and libraries are written (at least up until now) by humans, who _care_ about writing good working code. Generally, if you use a library, it will do it's job flawlessly, in the bounds of what it's expected to do. You rely on those people writing and maintaining it to perfect the job it was designed to do. Co-pilot is making suggestions, each a entirely independent "guess" at what you're trying to accomplish (apologies I don't know _that_ much about ML, but don't think this is far fetched). Meaning that each suggestion produces code in "untested waters" and wasn't code written specifically to do this job, used over and over by others... It's not like it's part of a project and a bug report going to filed for a bug...
Edit: Moved the top portion below, as it was sort of just repeating what was said:
I think the comparison of sifting through copilot suggestions to sifting through errors in google is that, generally you are looking for the error that fixes your problem. Not 5 different (probably working) solutions, which each may contain different bugs. Meaning that I (or X developer of any level would) _need_ to continue searching google to find a working answer to my problem. But validating a co-pilot answer for any potential flaws is much more error prone.
This makes sense since it wasn't trained on the private codebase but not using existing helpers, naming, abstractions makes the results largely useless since only a minority of output is useful. Like the author mentions what these tools do with the input is a gray area and nobody wants to publicize their code through chatbots. It'd be cool if there was a selfhosted alternative out there, trainable on private codebases.
I've also tried using the bots for answering questions about tech problems I've been having. Answers are mostly off point, have a lot of nonrelated info and most importantly are generated slowly. In the time I get an answer from a chatbot I can scan through a dozen Stack Overflow pages.
As far as I'm concerned we're not there yet but never say never.
I personally love these toys and the only downside is that they are services and not local tools (yet).
As for copilot, it's ok. GPT-4 is much much better when you learn how to use it. Busy work is completely automated. I need to transform strings in one form into structs? Just paste a representative sample of the data and tell it to. Is the data dirty? Just tell it how it's dirty and it will handle those edge cases. Want unit tests to verify the function it wrote works for you? Tell it to write the unit tests, then read them, point out what it got wrong, then check again.
When I was a mouthy teenager, my mom used to call this "It's easy to shit when your butt is full".
Sounds better in Slovenian. It's the same idea as The Man in The Arena – it's easy to have opinions when all your needs are met. If you really believe in what you're saying, get in the trenches and start doing the work.
For example: All the people railing against capitalism while getting large parts of their comp from equity and stocks and such. And making top 5% incomes. If you really believe capitalism is bad, why don't you leave Tech, sell off all your equity/stocks, and go live in a commune somewhere? hmmmm
I find this hard to believe. I know that $10 / month (for anything) is prohibitive for some people in some regions but hard to believe that if you already have a computer, internet connection and a cell phone, that an incremental $10 / month would break the bank.