the level of teaching involved would always mean the overall velocity of work slowed down.
some people say you can throw them the drudge work but i find that if you're doing coding right (e.g. you dont let your code base degenerate into a mess of boilerplate), there is barely any drudge work to do.
Me, too. But that doesn't mean I'm a great developer, just a shitty manager.
I suspect that will change sooner or later. Models will be cultivated over time the way we cultivate full-time employees now, with an acquired awareness of what they're building, new skills picked up in the process, and insight into how the larger system works.
There's a long-running instance of the model on the provider that's allocated to my organisation? Or are you thinking more of a server-side memory system, similar to the (currently very fallible) ones like Honcho and Mem0?
What happens when my org stops paying that provider? Do I get to take the now senior agent with me to the next provider? Does the provider have to delete it (and all that learning is lost forever)? Does that now become a free agent that can be hired by the next organisation like an employee (one that probably doesn't know how to keep industry secrets to itself)?
Karpathy's notion of a self-maintained wiki may suggest a direction for work in that area. As he points out ( https://gist.github.com/karpathy/442a6bf555914893e9891c11519... ) this is basically the latest of several attempts at implementing Vannevar Bush's "Memex" concept. I think it will find application in some form at organization-wide scope.
If the trend towards centralization holds, then yes, who gets to own and maintain the state in a stateful AI model is indeed going to be a big question. I hope it doesn't hold.
I just cant pretend that I'll get extra productivity while I'm training them.
Certain professions lend themselves well to a apprentice/master framework because the apprentice doing the drudge work that requires less skill frees up time. This can increase overall productivity while the apprentice is trained. Development isnt really like that though.
It is an investment, but it pays back for itself is my point.
I let the AI make decisions all the time. I often approve them, and I sometimes revert them. Most of the time they’re really good decisions based on my initial intent, but followed by analysis I didn’t make but agree with.
There's clearly some level where you want a human making decisions for even the most vibey of project, because without some kind of a spec about what you're trying to build and what features you want you'd get nonsense.
But like... maybe don't stress the details too much.
Yes, clearly. There was a meme out there, "just make something cool idk".
Statements like "Don't let AI make decisions" are made because of the loss of control we experience as mechanical parts of our work (such as writing to files) gets automated.
> A computer can never be held accountable
> Therefore a computer must never make a management decision > A computer can never be held accountable
> Therefore a computer must never make a management decision
It makes sense. But most decisions, by quantity, are not management decisions.The difference is profound, and takes more than a couple of days to get your head around the implications. I'd summarise it as: "if you give a computer the same input it always produces the same output, but if you give a model the same input it always produces different output". Add to that the output is often wrong and it can't reliably follow instructions, and the difference is so great it breaks most of your intuitions.
The reward working with this piece of unreliable jelly is it can be far smarter than you (think the difference between a man with a shovel and a 20 ton excavator - they can literally find bugs in minutes that would take a human hours or days), and they know far more than you.
The engineering challenge is to make this near random machine produce a reliable product. It isn't easy.
The hype you see around them is it's trivially easy to get it to produce a feature rich but very unreliable product, as Anthropic demonstrates with their vibe coded claude-cli. I refuse to use it now. Among its other charms, it triggers a BSOD on windows: https://github.com/anthropics/claude-code/issues/30137 (Granted, it's just another Windows bug: https://learn.microsoft.com/en-ca/answers/questions/5814272/..., but if you are shipping to Windows you should be working around such bugs.)
I think this solidified an idea for me. A tool being smarter than me but inconsistent, is useless.
I can work with people who are smarter than me, because I can trust them, and I can trust them to own up or be held accountable for screw ups.
For a calculator, I can only hold myself accountable. However I cannot hold myself accountable for not knowing something I dont know.
I thought I could let this fly through keeper, but the idea keeps bugging me.
If the programmer does not vet and understand what the LLM has done, we call it vibe coding. If he takes personal responsibility for every line of code, we call it software engineering.
Lots of people use LLMs to vibe code (I do), and lots of people use them as a software engineering aid (I do that too).
The choice depends on the quality of code you need to produce. If an engineer is designing an aircraft wing, he would choose the highest grade of materials with known parameters. If he is building a kite for his son, an old newspaper would do in a pinch. Using 7075 Aluminium wouldn't just be overkill, it would mean his son doesn't get to see the fun his dad has building a flying machine in an afternoon.
A good engineer understands the properties of every tool at his disposal, and knows when to deploy each.
They're great at some stuff and terrible at other stuff in ways that are very hard to predict.
I'm figuring out new and better ways to use them in a daily basis, and I've been an almost daily user for nearly three years.
Let's say models can exactly and correctly write any code you ask of them.
- How do you break down a project into a sequence of requests to models?
- How can you most effectively parallelize the work - models will never be instant, so there will always be benefits in working out how best to use several agents at once
- Now that the models can handle the details of Lean, and Swift-UI, and Oracle stored procedures, and thousands of other technologies that you never got around to learning in the past... what can you do with those and how do you pick which projects to go after?
- How do you collaborate with other engineers and designers and product people in a world where you can churn out the right code reliably in a few minutes?
The models we have today are already effective enough to change the shape of our work as software engineers. As the models continue to improve figuring out and adapting to whatever that new shape is becomes even more complicated.
You need to be able to estimate how long an approach is likely to take in order to correctly prioritize your work.
Coding agents throw a wrench in this because some of the stuff that used to take a long time doesn't take a long time any more. I'm finding my 25+ years of experience in estimating software has been completely thrown off - it's taking a whole lot of work for me to start building up those intuitions about what's quick and what takes a long time.
With coding agents you need to think very carefully about how you design the agentic loop such that the agent has the right tools and information available to it to compete the goal.
I've been writing a lot more about that here: https://simonwillison.net/guides/agentic-engineering-pattern...
If one has been reading a wide variety of books/papers/articles/whatever their whole life, and one has been mindful of how to communicate with the "written word" as it were, it takes about 3 hours to be wildly effective with this technology. I think it took longer to learn google-fu than it did to learn how to use this technology effectively.
Disagree. It takes a lot of experimenting to find the right balance between sufficient guardrails and insane halluciations. And it'll be different depending on work domains.
I'm still refactoring AI workflows every week after more than a year or so and still working on it. Will probably be a perpetually ongoing effort as models change.
If you spend a year walking in circles, someone can easily close the gap with one step. Especially if models and harnesses are supposedly getting more powerful all the time.
I kind of agree in general that it is a learned skill, but considering how unclear people generally are when they communicate, I'm guessing it'll take longer than a weekend to be able to catch up, especially catch up to people who've been working on precise and careful communication and language for years already in a professional environment.
That said - what we have learned in the last year could be compressed quite a lot - there are a lot steps we could skip, and 'learn by failure' that need not be repeated.
It takes a while to get the subtleties of it, it's among the most highly nuanced things we've ever encountered.