How I use LLMs as a staff engineer
seangoedecke.com
seangoedecke.com
In my opinion, using LLMs to write code comes as a faustian deal where you learn terrible practices and rely on code quantity, boilerplate, and indeterministic outputs - all hallmarks of poor software craftsmanship. Until ML can actually go end to end on requirements to product and they fire all of us, you can't cut corners on building intuition as a human by forgoing reading and writing code yourself.
I do think that there is a place for LLMs in generating ideas or exploring an untrusted knowledge base of information, but using code generated from an LLM is pure madness unless what you are building is truly going to be thrown away and rewritten from scratch, as is relying on it as a linting, debugging, or source of truth tool.
All the other distilled models and qwen coder and similar are a large step below the above models in terms of most benchmarks. If someone is running a small 20GB model locally, they will not have the same experience as those who run the top of the line models.
The result works ok, nobody cares if the code is good or bad. If it’s bad and there are bugs, doesn’t matter, no humans will look at it anymore - Claude will remix the slop until it works or a new model will rewrite the whole thing from scratch.
Realized during writing this that I should’ve added the extract of requirements in the comment of the index.ts of the package, or maybe a README.CURSOR.md.
My defense is that Karpathy does the same thing, admitted himself in a tweet https://x.com/karpathy/status/1886192184808149383 - I know exactly what he means by this.
At the absolute minimum this should require including a highly detailed function specification in the prompt context and sending the output to a full unit test suite.
Lordy. Is this where software development is going over the next few years?
Personally, I’m on the fence. But having conversations with others, and some requests from execs to implement different AI utils into our processes… making me to be on the safer side of job security, rather than dismiss it and be adamant against it.
I've heard exactly the same stories from my friends in larger tech companies as well. Every all hands there's a push for more AI integration, getting staff to use AI tools and etc., with the big expectation that development will get faster.
(I know it because I'm in charge of maintaining all processes around LLM keys, their usages, Cursor stuff and etc.)
> I'm in charge of maintaining all processes around LLM keys
Does management look to you for insight on which staffers are appropriately committed to leveraging AI?
If we take the premise at face value, then this is a time management question, and that’s a part of pretty much every performance evaluation everywhere. You’re not rewarded for writing some throwaway internal tooling that’s needed ASAP in assembly or with a handcrafted native UI, even if it’s strictly better once done. Instead you bash it out in a day’s worth of Electron shitfuckery and keep the wheels moving, even if it makes you sick.
Hyperbole aside, hopefully the point is clear: better is a business decision as much as a technical one, and if an LLM can (one day) do the 80% of the Pareto distribution, then you’d better be working on the other 20% when management come knocking. If I run a cafe, I need my baristas making coffee when the orders are stacking up, not polishing the machine.
Caveats for critical code, maintenance, technical debt, etc. of course. Good engineers know when to push back, but also, crucially, when it doesn’t serve a purpose to do so.
<<never uses AI>>
I've already seen several rounds of slacks: "why aren't you using <insert LLM coding assistant name>?" off the back of this reporting.
These assistants essentially spy on you working in many cases, if the subscription is coming from your employer and is not a personal account. For one service, I was able to see full logging of all the chats every employee ever had.
Executives, directors, and managerial staff have had their heads up their own asses since the dawn of civilization. Riding the waves of terrible executive decisions is unfortunately part of professional life. Executives like the idea of LLMs because it means they can lay you off; they're not going to care about your opinion on it one way or another.
> Being very anti-LLM code instead of trying to understand how it can improve the speed might be detrimental for your career.
You're making the assumption that LLMs can improve your speed. That's the very assumption being questioned by GP. Heaps of low-quality code do not improve development speed.
It becomes some sort of muscle memory, where I can predict whether using LLM would be faster or slower. Or where it's more likely to give bad suggestions or not. Basically treating it as googling skills.
Yeah, that intuition is so important. You have to use the models a whole bunch to develop it, but eventually you get a sort of sixth sense where you can predict if an LLM is going to be useful or harmful on a problem and be right about it 9/10 times.
My frustration is that intuition isn't something I can teach! I'd love to be able to explain why I can tell that problem X is a good fit and problem Y isn't, but often the answer is pretty much just "vibes" based on past experience.
1) Being up to date with the latest capabilities - this one has slowed down a bit, and my biggest self-learning experience was in August/September, and most of the intuition still works. However, although I had time to do so, it's hard to ask my team to drop 5-6 free weekends of their lives to get up to speed
2) Transition period where not everyone is on the same page about LLM - this I think is much harder, because the expectations from the executives are much different than on the ground developers using LLMs.
A lot of people could benefit on alignment of expectations, but once again, it's hard to explain what is possible and not possible if your statements will become nullified a month later with a new AI model/product/feature.
Not sure I agree with that. I would say there are classes of problems where LLMs will generally help and a brief training course (1 week, say) would vastly improve the average (non-LLM-trained) engineer's ability to use it productively.
In real, traditional, deterministic systems where you explicitly design a feature, even that has difficulty being coherent over time as usage grows. Think of tab stops on a typewriting evolving from an improvised template, to metal tabs installed above the keyboard, to someone cutting and pasting incorrectly and reflowing a 200 page document to 212 pages accidentally because of tab characters...
If you create a system with these models that writes the code to process a bunch of documents in some way or so some kind of herculean automation you haven't improved the situation when it comes to clarity or simplicity, even if the task at hand finishes sooner for you in this moment.
Every token generated has an equal potential to spiral out into new complexities and whack a mole issues that tie you to assumptions about the system design while providing this veneer that you have control over the intersections of these issues, but as this situation grows you create an ever bigger problem space.
And I definitely hear you say, this is the point where you use sort of full stack interoception holistic intuition about how to persuade the system towards a higher order concept of the system and expand your ideas about how the problem could be solved and let the model guide you ... And that is precisely the mysticism I object to because it isn't actually a kind productiveness, but a struggle, a constant guessing, and any insight from this can be taken away, changed accidentally, censored, or packaged as a front run against your control.
Additionally the nature of not having separate in band and out of band streams of data means that even with agents and reasoning and all of the avenues of exploration and improving performance will still not escape the fundamental question of ... What is the total information contained in the entire probabilistic space. If you try to do out of band control in some way like the latest thing I just read where they have a separate censoring layer, you just either wind up having to use another LLM layer there which still contains all of these issues, or you use some kind of non transformer method like Bayesian filtering or something and you get all of the issues outlined in the seminal spam.txt document...
So, given all of this, I think it is really neat the kinds of feats you demonstrate, but I object that these issues can be boiled down to "putting a serious amount of effort into learning how to best apply them" because I just don't think that's a coherent view of the total problem, and not actually something that is achievable like learning in other subjects like math or something. I know it isn't answerable but for me a guiding question remains why do I have to work at all, is the model too small to know what I want or mean without any effort? The pushback against prompt engineering and the rise of agentic stuff and reasoning all seems to essentially be saying that, but it too has hit diminishing returns.
My experience has been that it's slightly improved code completion and helped with prototyping.
Probably, but when the time comes for layoffs the ones that will be the first to go are those that are hiding under a rock, claiming that there is no value to those LLM’s even as they’re being replaced.
The need for real coding skills however, won't.
First, what LLMs/GenAI do is automated code generation, plain and simple. We've had code generation for a very long time; heck, even compiling is automated generation of code.
What is new with LLM code generation is non-deterministic, unlike traditional code generation tools; like a box of chocolates, you never know what you're going to get.
So, as long as you have tools and mechanisms to make that non-determinism irrelevant, using LLMs to write code is not a problem at all. In fact, guess what? Hand-coding is also non-deterministic, so we already have plenty of those in place: automated tests, code reviews etc.
Don’t get me wrong—I’ve seen productivity gains both in LLMs explaining code/ideation and in actual implementation, and I use them regularly in my workflow now. I quite like it. But these people are itching to eliminate the cost of maintaining a dev team, and it shows in the level of wishful thinking they display. They write a snake game one day using ChatGPT, and the next, they’re telling you that you might be too slow—despite a string of record-breaking quarters driven by successful product iterations.
I really don’t want to be a naysayer here, but it’s pretty demoralizing when these are the same people who decide your compensation and overall employment status.
And this is the promise of AI, to eliminate jobs. If CEOs invest heavily in this, they won't back down because no one wants to be wrong.
I understand some people try to claim AI might make net more jobs (someday), but I just don't think that is what CEOs are going for.
They might not have to. If the results are bad enough then their companies might straight-up fail. I'd be willing to bet that at least one company has already failed due to betting too heavily on LLMs.
That isn't to say that LLMs have no uses. But just that CEOs willing something to work isn't sufficient to make it work.
Pre-LLM, that's how Boeing destroyed itself. By creating value for the shareholders.
If your experienced employees, who are giving an honest try to all these tools, are telling you it’s not a silver bullet, maybe you should hold your horses a little and try to take advantage of reality—which is actually better—rather than forcing some pipe dream down your bottom line’s throat while negating any productivity gains by demotivating them with your bullshit or misdirecting their efforts into finding a problem for a given solution.
Maybe I’m being overly cynical, but assuming this isn’t a race to the bottom and people will get rich being super productive ai-enhanced code monsters, to me, looks like a conceited white collar version of the hustle porn guys that think if they simultaneously work the right combo of gig apps at the right time of day in the right spots then they can work their way up to being wealthy entrepreneurs. Good luck.
What seems far more likely to me is that computer scientists will be doing math research and wrangling LLMS, a vanishingly small number of dedicated software engineers work on most practical textual coding tasks with engineering methodologies, and low or no code tooling with the aid of LLMs gets good enough to make custom software something made by mostly less-technical people with domain knowledge, like spreadsheet scripting.
A lot of people in the LLM booster crowd think LLMs will replace specialists with generalists. I think that’s utterly ridiculous. LLMs easily have the shallow/broad knowledge generalists require, but struggle with the accuracy and trustworthiness for specialized work. They are much more likely to replace the generalists currently supporting people with domain-specific expertise too deep to trust to LLMs. The problem here is that most developers aren’t really specialists. They work across the spectrum of disciplines and domains but know how to use a very complex toolkit. The more accessible those tools are to other people, the more the skill dissolves into the expected professional skill set.
If LLMs make average devs 10x more productive, Jevon's Paradox[1] suggests we'll just make 10x more software rather than have 10x fewer devs. You can now implement that feature only one customer cares about, or test 10x more prototypes before building your product. And if you instead decide to decimate your engineering team, watch out because your competitors might not.
I think that's the sort of spot where better tools might be appropriate. I know what I want to do, but it's a mess to do it. I suspect that will be better at facilitating growth instead of stunting it.
The entire reason they hire us is to let them know if what they think makes sense. No one is ideologically opposed to AI generated code. It comes with lots of negatives and caveats that make relying on it costly in ways we can easily show to any executives, directors, etc. who care about the technical feasibility of their feelings.
Unfortunately, that hasn't been my experience. But I agree with you comment generally.
Edit: fix typo in last sentence
And very often, if the LLM produces a poopoo, asking it to fix it again works just well enough.
I've yet to encounter any LLM from chatGPT to cursor, that doesn't choke and start to repeat itself and say it changed code when it didn't, or get stuck changing something back and forth repeatedly inside of 10-20 minutes. Like just a handful of exchanges and it's worthless. Are people who make this workflow effective summarizing and creating a fresh prompt every 5 minutes or something?
I estimate a sizable portion of my successful LLM coding sessions included at least a few resets of this nature.
Learning takes time.
This is the most important thing in my opinion. This is why I switched to showing tokens in my chat app.
https://beta.gitsense.com/?chat=b8c4b221-55e5-4ed6-860e-12f0...
I treat tokens like the tachometer for a car's engine. The higher you go, the more gas you will consume, and the greater the chance you will blow up your engine. Different LLMs will have different redlines and the more tokens you have, the more costly every conversation will become and the greater the chance it will just start spitting gibberish.
So far, my redline for all models is 25,000 tokens, but I really do not want to go above 20,000. If I hit 16,000 tokens, I will start to think about summarizing the conversation and starting a new one based on the summary.
The initial token count is also important in my opinion. If you are trying to solve a complex problem that is not well known by the LLM and if you are only starting with 1000 or less tokens, you will almost certainly not get a good answer. I personally think 7,000 to 16,000 is the sweet spot. For most problems, I won't have the LLM generate any code until I reach about 7,000 since it means it has enough files in context to properly take a shot at producing code.
> Are people who make this workflow effective summarizing and creating a fresh prompt every 5 minutes or something?
I work on one small problem at a time, only following up if I need an update or change on the same block of code (or something very relevant). Most conversations are fewer than five prompt/response pairs, usually one-three. If the LLM gets something wrong, I edit my prompt to explain what I want better, or to tell it not to take a specific approach, rather than correcting it in a reply. It gets a little messy otherwise, and the AI starts to trip up on its own past mistakes.
If I move on to a different (sub)task, I start a new conversation. I have a brief overview of my project in the README or some other file and include that in the prompt for more context, along with a tree view of the repository and the file I want edited.
I am not a software engineer and I often need things explained, which I tell the LLM in a custom system prompt. I also include a few additional instructions that suit my workflow, like asking it to tell me if it needs another file or documentation, if it doesn't know something, etc.
* If you ask it to solve a problem and nothing more, chances are the code isn't the best as it will default to the most common solutions in the training data.
* If you ask it to refactor some code idiomatically, it will apply most common idiomatic concepts found in the training data.
* If you ask it to do both at the same time you're more likely to get higher quality but incorrect code.
It's better to get a working solution first, then ask it to improve that solution, rinse/repeat in smallish chunks of 50-100 loc at a time. This is kinda why reasoning models are of some benefit, as they allow a certain amount of reflection to tie together disparate portions of the training data into more cohesive, higher quality responses.
That does not match my experience at all. You obviously have to use your brain to review it, but for a lot of problems LLMs produce close to perfect code in record time. It depends a lot on your prompting skills though.
Are you sure? I've definitely had cases where an inexperienced pair programmer made my code worse.
*I don’t know to what extent it’s worthwhile discussing whether you could call these the same model vs. entirely different, for any two products in the same family. Outside of simply quantising the same model and nothing else. Maybe you could include distillations of a base model too?
I built my chat app around this idea and to save money. When it comes to coding, I feel Sonnet 3.5 is still the best but I don't start with it. I tend to use cheaper models in the beginning since it usually takes a few iterations to get to a certain point and I don't want to waste tokens in the process. When I've reached a certain state or if it is clear that the LLM is not helping, I will bring in Sonnet to review things.
Here is an example of how the conversation between models will work.
https://beta.gitsense.com/?chat=bbd69cb2-ffc9-41a3-9bdb-095c...
The reason why this works for my application is, I have a system prompt that includes the following lines:
# Critical Context Information
Your name is {{gs-chat-llm-model}} and the current date and time is {{gs-chat-datetime}}.
When I make an API call, I will replace the template strings with the model and date. I also made sure to include instructions in the first user message to let the model know it needs to sign off on each message. So with the system prompt and message signature, you can say "what do you think of <LLM's> response".
You still need to state your assertions with precision and keep a model of the code in your head.
Its possible to be be precise at an higher level of abstraction as long as your prompts are consistent with a coherent model of the code.
This is a fantastic quote and I will use this. I describe the future of coding as natural language coding (or maybe syntax agnostic coding). This does not mean that the llm is a magic machine that understands all my business logic. It means what you've described - I can describe my function flow in abstracted english rather than requiring adherence to a syntax
Its like having to forever be the most miserable detective in the world; no mystery, only clues. A method that never existed, three different types that express the same thing, the cheeky smile of your coworker who says he can turn the whole backend into using an ORM in a day because he has Cursor, the manager who signs off on this, the deranged PR the next day. This continual sense that less and less people even know whats going on anymore...
"Can you make sure we support both Mongo and postgres?"
"Can you put this React component inside this Angular app?"
"Can you setup the kubernetes with docker compose?"
If youre getting junior devs just pooping out code and sending to review thats really bad and should be a pip-able offense in my opinion.
If you're truly saving time by having an LLM write boiler plate code, is there maybe an opportunity to abstract things away so that higher-level concepts, or more expressive code could be used instead?
5 lines of code written with just the core language and standard library are often much easier to read and digest than a new abstraction or call to some library.
And it’s just an unfortunate fact of life that many of the common programming languages are not terribly ergonomic; it’s not uncommon for even basic operations to require a few lines of boilerplate. That isn’t always bad as languages are balancing many different goals (expressiveness, performance, simplicity and so on).
In a way LLMs are ushering in a kind of boilerplate renaissance IMO. When you can have an LLM refactor a massive amount of boilerplate in one fell swoop it starts to not matter much if you repeat yourself - actually, really logically dense code would probably be harder for LLMs to understand and modify (not dissimilar from us…) so it’s even more of a liability now than in the past. I would almost always rather have simple, easy-to-understand code than something elegant and compact and “expressive” - and our tools increasingly favor this too.
Also I really don’t give a shit about how to best center a div nor do I want to memorize a million different markup tags and their 25 years of baggage. I don’t find that kind of knowledge gratifying because it’s more trivia than anything insightful. I’m glad that with LLMs I can minimize the time I spend thinking about those things.
Other things are complicated to abstract for the boilerplate they avoid. The kind of thing that avoids 100 lines of code but causes errors that take 20 minutes to understand because of heavy use of reflection/inferred types in generics/etc. The older I get, the more I think "clever" reflection is more of a sin than boring boilerplate.
> I don’t do this a lot, but sometimes when I’m really stuck on a bug, I’ll attach the entire file or files to Copilot chat, paste the error message, and just ask “can you help?”
The "reasoning" models are MUCH better than this. I've had genuinely fantastic results with this kind of thing against o1 and Gemini Thinking and the new o3-mini - I paste in the whole codebase (usually via my https://github.com/simonw/files-to-prompt tool) and describe the bug or just paste in the error message and the model frequently finds the source, sometimes following the path through several modules to get there.
Here's a slightly order example: https://gist.github.com/simonw/03776d9f80534aa8e5348580dc6a8... - finding a bug in some Django middleware
I've had the experience of seeing some junior dev posting error messages into ChatGPT, applying the suggestions of ChatGPT, and posting the next error message into ChatGPT again. They ended up applying fixes for 3 different kinds of bugs that didn't exist in the code base.
---
Another cause, I think, is that they didn't try to understand any of those (not the solutions, and not the problems that those solutions are supposed to fix). If they did, they would have figured out that the solutions were mismatches to what they were witnessing.
There's a big difference between using LLM as a tool, and treating it like an oracle.
I just had a case where I was adding stuff to two projects, both open at the same time.
I added new fields to the backend project, then I swapped to the front-end side and the LLM autocomplete gave me 100% exactly what I wanted to add there.
And similar super-accurate autocompletes happen every day for me.
I really don't understand people who complain about "AI slop", what kind of projects are they writing?
One of my favourite things is to ask it if it thinks there are any bugs - this helps a lot with validating any logic that I might be exploring. I recently ported some code to a different environment with slightly different interfaces and it wasn't working - I asked o1 to carefully go over each implementation in detail why it might be producing a different output. It thought for 2 whole minutes and gave me a report of possible causes - the third of which was entirely correct and had to do with how my environment was coercing pandas data types.
There have been 10 or so wow moments over the past few years where I've been shocked by the capabilities of genai and that one made the list.
LLMs can absolutely bust out some corporate docs super crazy fast too... probably a reasonable thing to re-evaluate the value though
I had a working and complete version of Apple MapKit JS rendering a map for an address (along with the server side token generation), and last night I told it I wanted to switch to Google Maps for "reasons".
It nailed it on the first try, and even gave me quick steps for creating the API keys in Google Dev Console (which is always _super_ fun to navigate).
As Simon has said elsewhere in these comments, it's all about the context you give it (a working example in a slightly different paradigm really couldn't be any better).
I have a project with a bunch of tests already, then I pick a test file and write `public Task Test` and wait a few seconds, in most cases it writes down a pretty sane basis for a test - and in a few cases it figured out an edge case I missed.
Ah, now it makes sense.
Traditional auto complete can finish the statement you started typing, LLMs often suggest whole lines before I even type anything, and even sometimes whole functions.
And static types can assist the LLM too. It's not like it's an either or choice
"Almost all the completions I accept are complete boilerplate (filling out function arguments or types, for instance). It’s rare that I let Copilot produce business logic for me"
My experience is similar, except I get my IDE to complete these for me instead of an LLM.
Hard for an IDE auto complete to do this.
I find Copilot is great if you add a small comment describing the logic or function. Taking 10s to write a one line sentence in English can save 5-10 mins writing your code from scratch. Subjectively it feels much faster to QA and review code already written.
Having good typing and DTOs helps too.
…needn’t say more.
Copilot was utter garbage when I switched to cursor+claude, it was like some alien tech upgrade at first.
The "dumb" autogenerated stuff is incredible. It's like going from bad autocomplete to Intellisense all over again.
The world of python tooling (at least as used by my former coworkers) put my expectations in the toilet.
Then I reflected, how very true it was. In fact, as of writing this there are 138 comments and I started simply scrolling through what was shown to assess the negative/neutral/positive bias based upon a highly subjective personal assessment: 2/3 were negative and so I decided to stop.
As a profession, it seems many of us have become accustomed to dealing in absolutes when reality is subjective. Judging LLMs prematurely with a level of perfectionism not even cast upon fellow humans.. or at least, if cast upon humans I'd be glad not to be their colleagues.
Honestly right now - I would use this as a litmus test in hiring and the majority would fail based upon their closed-mindedness and ability to understand how to effectively utilise tools at their disposal. It won't exist as a signal for much longer, sadly!
We need to trust machines more than humans because machines can't get responsibility. That code that you pushed and broke prd - you can't point at the machine.
It is also predictability/growth in a sense. I can assess certain people and know what they will probably get wrong and develop the person and adjust it. If that person uses LLMs it disguises that exposure of skill and leads to a very hard signal to read as a senior dev, hampering their growth.
I did see a paper recently on the impact of AI/LLMs and danger to critical thinking skills - it's a very real issue and I'm having to actively counter this seemingly natural tendency many have.
With respect to signals, mine was around the attitude in general. I'd much rather work with someone who goes "Yes, but.." than one who is outright dismissive.
Increasing awareness of the importance of context will be a topic for a long time to come!
See this is what I don't get about the AI Evangelists. Every time I use the technology I am astounded at the amount of incorrect information and straight up fantasy it invents. When someone tells me that they just don't see it, I have to wonder what is motivating them to lie. There is simply no way you're using the same technology as me with such wildly different results.
Most of these people who aren't salesmen aren't lying.
They just cannot tell when the LLM is making up code. Which is very very sad.
That or they could literally be replaced by a script that copy/pastes from stack-overflow. My friend did that a lot and it definitely helped features ship but doesn't make maintainable code.
Prompting styles are incredibly different between different people. It's very possible that they are using the same technology that you are with wildly different results.
I think learning to use LLMs to their maximum effectiveness takes months (maybe even years) of effort. How much time have you spent with them so far?
I don’t know what technology you are using but I know that I am getting very different results based on my own prompt qualities.
I also do not really consider hallucinations to be much of an issue for programming. It comes up so rarely and it’s caught by the type checker almost immediately. If there are hallucinations it’s often very minor things like imagining a flag that doesn’t exist.
"is this idiomatic C?"
"not just “how does X work”, but follow-up questions like “how does X relate to Y”. Even more usefully, you can ask “is this right” questions"
"I’ll attach the entire file or files to Copilot chat, paste the error message, and just ask “can you help?”"
It’s hard to impossible to discuss these things without a concrete problem at hand. Most of the prompt is the context provided. I can only talk from my experience which is, that how you write the prompt matters.
If you share what you are trying to accomplish I could probably provide some more appropriate insights.
- ai-assisted-programming: https://simonwillison.net/tags/ai-assisted-programming/
- prompt-engineering: https://simonwillison.net/tags/prompt-engineering/
And my series on how I use LLMs: https://simonwillison.net/series/using-llms/
And they aren’t perfect, but they sure can save a lot of time once you know how to use them and understand what they are and aren’t good at.
This is how I use AI at work for maintaining Python projects, a language in which I am not at all really versed. Sometimes I might add “this is how I would do it in …, how would I do this in Python?”
I find this extremely helpful and productive, especially as I have to pull the code onto a server to test it.
Asking a LLM to translate between languages works really well most of the time. It's also a great way to learn which libraries are the standard solution for a language. It really accelerated my learning process.
Sure, there is the occasional too literal translation or hallucination, but I found this useful enough.
When I stray out of this (e.g. I started doing a lot of IoT, ML and Robotics projects, where I can't always use TypeScript). I think one key thing that LLMs have helped me is that I can ask why something is X without having to worry about sounding stupid or annoying.
So I think it has enabled me at least a way to get out of the TypeScript zone more worry free without losing productivity. And I do think I learn a lot, although I'm relating a lot of it on my JS/TS heavy experience.
To me the ability to ask stupid questions without fear of judgment or accidentally offending someone - it's just amazing.
I used to overthink a lot before LLMs, but they have helped me with that aspect, I think a lot.
I sometimes think that no one except LLMs would have the patience for me if I didn't filter my thoughts always.
I have brains and can verify if it's correct or not.
I do thoroughly review of the the LLM answers, and hardly every directly copy paste answer, so I feel this way I still learn the language.
One thing that is not mentioned -- code review. It is not great at it, often pointing out trivial or non issues. But if it finds 1 area for improvement out of 10 bullet points, that's still worth it -- most human code reviewers don't notice all the issues in the code anyway.
I need to avoid LLM use to ensure my coding ability stays up to par.
--
I work on Graphite Reviewer (https://graphite.dev/features/reviewer). I'm also partly dyslexic. I lean massively on Grammarly (using it to write this comment) and type-safe compiled languages. When I engineered at Airbnb, I caused multiple site outages due to typos in my ruby code that I didn't see and wasn't able to execute before prod.
The ability for LLMs to proofread code is a godsend. We've tuned Graphite Reviewer to shut up about subjective stylistic comments and focus on real bugs, mistakes, and typos. Fascinatingly, it catches a minor mistake in ~1/5 PRs in prod at real companies (we've run it on a few million PRs now). Those issues it catches result in a pre-merge code change 75% of the time, about equal to what a human comment does.
AIs aren't perfect, but Im thrilled that they work as fancy code spell-checkers :)
CoPilot is used for simple boilerplate code, and also for the autocomplete. It's often a starting point for unit tests (but a thorough review is needed - you can't just accept it, I've seen it misinterpret code). I started experimenting with RA.Aid (https://github.com/ai-christianson/RA.Aid) after seeing a post on it here today. The multi-step actions are very promising. I'm about to try files-to-prompt (https://github.com/simonw/files-to-prompt) mentioned elsewhere in the thread.
For now, LLMs are a level-up in tooling but not a replacement for developers (at least yet)
1. Try to write some code
2. Wonder why my IDE is providing irrelevant, confusing and obnoxious suggestions
3. Realize the AI completion plugin somehow turned itself back on
4. Turn it off
5. Do my job better than everyone that didn't do step 4
The question I keep asking myself is, "Should we be making tools that auto-write code for us, or should we be using this training data to suss out the missing tools we have where everyone writes the same code 10 times in their careers?"
Such an unnecessary flex.
Case in point, this person has around around 7 years of professional experience at just two companies, Zendesk and GitHub. I don't mean this as a personal dig in any way (truly) but this simply isn't what we used to mean by a "Staff" level software engineer.
This person is early-mid career, which we used to just call "Software engineer" then "Senior Software Engineer" and now (often enough) "Staff Software Engineer"
Not because of lack of skill, but I don't care. I could ask for a fancier one and most likely get it, but why?
It also helps if you realize staff+ is just a way to financially reward people who don’t want to be managers so you end up with these unholy engineer/architect/project manager/product manager hybrids that have to influence without authority.
I like this definition (which comes with a whole book): https://staffeng.com/
At the end, I ask it to give me a quiz on everything we talked about and any other insights I might have missed. Instead of typing out the answers, I just use Apple Dictation to transcribe my answers directly.
It's only recently that I thought to take the conversation I just had, and have it write a blog post of the insights and ah-ha moments I had, and have it write a blog post. It takes a fair bit of curation to get it to do that, however. I can't just say, "write me a blog post on all we talked about". I have to first get it to write an outline with the key insights. And then based on the outline, write each section. And then I'll use chatgpt's canvas to guide and fine-tune each section.
However, at no point do I have to specifically write the actual text. I mostly do curation.
I feel ok about doing this, and don't consider it AI slop, because I clearly mark at the top that I didn't write a word of it, and it's the result of a curated conversation with 4o. In addition, I think if most people do this as a result of their own Socratic methods with an AI, it'd build up enough training data for next generation of AI to do a better job of writing pedagogical explanations, posts, and quizzes to get people learning topics that are just out of reach, but there hadn't been too many people able to bridge the gap.
The two I had it write are: Effects as Protocols and Contexts as Agents: https://interjectedfuture.com/effects-as-protocols-and-conte...
How free monads and functors represent syntax for algebraic effects: https://interjectedfuture.com/how-the-free-monad-and-functor...
This is key - if it's marked clearly as AI-generated or assisted, it's not slop. I think this is an important part of AI ethics that most people can agree with.
Coding assistant LLMs have changed how I work in a couple of ways:
1) They make it a lot easier to context switch between e.g. writing kernel code one day and a Pandas notebook the next, because you're no longer handicapped by slightly forgetting the idiosyncrasies of every single language. It's like having smart code search and documentation search built into the autocomplete.
2) They can do simple transformations of existing code really well, like generating a match expression from an enum. They can extrapolate the rest from 2-3 examples of something repetitive, like converting from Rust types into corresponding Arrow types.
I don't find the other use cases the author brings up realistic. The AI is terrible at code review and I have never seen it spot a logic error I missed. Asking the AI to explain how e.g. Unity works might feel nice, but the answers are at least 40% total bullshit and I think it's easier to just read the documentation.
I still get a lot of use out of Copilot. The speed boost and removal of friction lets me work on more stacks and, consequently, lead a much bigger span of related projects. Instead of explaining how to do something to a junior engineer, I can often just do it myself.
I don't understand how fresh grads can get use out of these things, though. Tools like Copilot need a lot of hand-holding. You can get them to follow simple instructions over a moderate amount of existing code, which works most of the time, or ask them to do something you don't exactly know how to do without looking it up, and then it's a crapshoot.
The main reason I get a lot of mileage out of Copilot is exactly because I have been doing this job for two decades and understand what's happening. People who are starting in the industry today, IMO, should be very judicious with how they use these tools, lest they end up with only a superficial knowledge of computing. Every project is a chance to learn, and by going all trial-and-error with a chatbot you're robbing yourself of that. (Not to mention the resulting code is almost certainly half-broken.)
Just last night I did a quick test on Cursor (first time trying it). Opened up my IRC bot project and asked it to "add relevant handlers for IRC messages".
It immediately recognised the pattern I had used before and added CTCP VERSION, KICK, INVITE and 433 (nickname already in use). It didn't try to add everything under the sun and just added those. Took me 20 seconds.
That's bad because it makes "not training your juniors" the default path for senior people.
I can assign the task to one of my junior engineers and they will take several days of back and forth with me to work out the details--that's annoying but it's how you train the next generation.
Or I can ask the LLM and it will spit back something from its innards that got indexed from Github or StackOverflow. And for a "junior engineer" task it will probably be correct with the occasional hallucination--just like my junior engineers. And all I have to do for the LLM is click a couple of keys.
With all the talk of o1-pro as a superb staff engineer-level architect, it took me awhile to re-parse this headline to understand what the author, apparently a staff engineer, meant
I stick to a "no copy & paste" rule and that includes autocomplete. Interactions are a conversation but I write all my code myself.
I would be so bored if my job consisted of writing prompts all day long.
LLMs make it quicker for me to:
- Decipher obscure error messages
- Knock out a quick exploratory prototype of a new idea, both backend and frontend code
- Write boiler plate code against commonly used libraries
- Debug things: feeding a gnarly bug plus my codebase into Gemini (for long context) or o3-mini can save me a TON of frustration
- Research potential options for libraries that might help with a problem
- Refactor - they're so good at refactoring now
- Write tests. Sometimes I'll have the LLM sketch out a bunch of tests that cover branches I may have not bothered to cover otherwise.
I enjoy working like this a whole lot more than I enjoyed working without them, and I enjoyed programming a lot prior to LLMs.
I enjoy writing code, but I enjoy seeing getting a feature out even more. In fact, I don't quite enjoy the part of writing basic logic or tweaking CSS which an intern can easily do.
I don't think anybody is writing prompts all day long. If you don't actually know how to write code, maybe. But at this point, a professional software engineer still works with a code base with a hands-on approach most of the time, and even heavy LLM users still spend a lot of time hand writing code.
- imprecise semantic search
- simple auto-completion (1-5 tokens)
- copying patterns with substitutions
- inserting commonly-used templates