I feel bad for people who haven't yet experienced how useful these models are for programming.
Some also just prefer manually entering everything. Those people I will never understand.
I feel bad for people who haven't yet experienced how useful these models are for programming.
Some also just prefer manually entering everything. Those people I will never understand.
For reference, I program systems code in C/C++ in a large, proprietary codebase.
My experiences with OpenAI(a year ago or more), and more recently, Cursor, Grok-v3 and Deepseek-r1, were all failures. The later two started out OK and got worse over time.
What I haven't done is asked "AI" to whip up a more standard application. I have some ideas(an ncurses frontend to p4 written in python similar to tig, for instance), but haven't gotten around to it.
I want this stuff to work, but so far it hasn't. Now I don't think "programming" a computer in english is a very good idea anyway, but I want a competent AI assistant to pair program with. To the degree that people are getting results, to me it seems they are leveraging very high-level APIs/libraries of code which are not written by AI and solving well-solved, "common" problems(simple games, simple web or phone apps). Sort of like how people gloss over the heavy lifting done by language itself when they praise the results from LLMs in other fields.
I know it eventually will work. I just don't know when. I also get annoyed by the hype of folks who think they can become software engineers because they can talk to an LLM. Most of my job isn't programming. Most of my job is thinking about what the solution should be, talking to other people like me in meetings, understanding what customers really want beyond what they are saying, and tracking what I'm doing in various forms(which is something I really do want AI to help me with).
Vibe coding is aptly named because it's sort of the VB6 of the modern era. Holy cow! I wrote a Windows GUI App!!!. It's letting non-programmers and semi-programmers(the "I write glue code in Python to munge data and API ins/outs" crowd) create usable things. Cool! So did spreadsheets. So did Hypercard. Andrej tweeting that he made a phone app was kinda cool but also kinda sad. If this is what the hundreds of billions spent on AI(and my bank account thanks you for that) delivers then the bubble is going to pop soon.
Usually that's because of context: LLMs are not very good at understanding a very large amount of context, but if you don't give LLMs enough context, they can't magically figure it out on their own. This relegates AI to only really being useful for pretty self-contained examples where the amount of code is small, and you can provide all the context it needs to do its job in a relatively small amount of text (few thousand words or lines of code at most).
That's why I think LLMs are only useful right now in real-world software development for things like one-off functions, new prototypes, writing small scripts, or automating lots of manual changes you have to do. For example, I love using o3-mini-high to take existing tests that I have and modifying them to make a new test case. Often this involves lots of tiny changes that are annoying to write, and o3-mini-high can make those changes pretty reliably. You just give it a TODO list of changes, and it goes ahead and does it. But I'm not asking these models how to implement a new feature in our codebase.
I think this is why a lot of software developers have a bad view of AI. It's just not very good at the core software development work right now, but it's good enough at prototypes to make people freak out about how software development is going to be replaced.
That's not to mention that often when people first try to use LLMs for coding, they don't give the LLMs enough context or instructions to do well. Sometimes I will spend 2-3 minutes writing a prompt, but I often see other people putting the bare minimum effort into it, and then being surprised when it doesn't work very well.
And in terms of time saved, if I am just changing string constants, it’s not going to help much. But if I’m restructuring the test to verify things in a different way, then it is helpful. For example, recently I was writing tests for the JSON output of a program, using jq. In this case, it’s pretty easy to describe the tests I want to make in English, but translating that to jq commands is annoying and a bit tedious. But o3-mini-high can do it for me from the English very well.
Annoying to do myself, but easy to describe, is the sweet spot. It is definitely not universally useful, but when it is useful it can save me 5 minutes of tedium here or there, which is quite helpful. I think for a lot of this, you just have to learn over time what works and what doesn't.
Maybe one of my problems is that I tend to jump into writing simple code or tests without actually having the end point clearly in mind. Often that works out pretty well. When it doesn’t, I’ll take a step back and think things through. But when I’m in the midst of it, it feels like being interrupted almost, to go figure out how to say what I want in English.
Will definitely keep toying with it to see where I can find some utility.
The areas that I've found LLMs work well for are usually small simple tasks I have to do where I would end up Googling something or looking at docs anyway. LLMs have just replaced many of these types of tasks for me. But I continue to learn new areas where they work well, or exceptions where they fail. And new models make it a moving target too.
Good luck with it!
Maybe that's why I don't like them. I'm always in a flow state, or reading docs and/or a lot of code to understand something. By the time I'm typing, I already know what exactly to write, and thanks to my vim-fu (and emacs-fu), getting it done is a breeze. Then comes the edit-compile-run, or edit-test cycle, and by then it's mostly tweaks.
I get why someone would generate boilerplate, but most of the time, I don't want the complete version from the get go. Because later changes are more costly, especially if I'm not fully sure of the design. So I want something minimal that's working, then go work on things that are dependent, then get back when I'm sure of what the interface should be. I like working iteratively which then means small edits (unless refactoring). Not battling with a big dump of code for a whole day to get it working.
If I've got a clear idea of what I want to write, there's no way I'm touching an LLM. I'm just going to smash out the code for exactly what I need. However, often I don't get that luxury as I'll need to learn different file system APIs, different sets of commands, new jargon, different standard libraries for the new languages, new technologies, etc...
Mostly I use it for stupid templates stuff for which it isn’t bad. It’s not the second coming but it definitely speeds you up
None of this is particularly unique to software engineering. So if someone can already do this and add the missing component with some future LLM why shouldn’t they think they can become a software engineer?
Did you catch the sarcasm there?
Are you a manager by any chance? The non-coding parts of my job largely require domain experience. How does an LLM provide you with that?
That's okay.
It's not my responsibility to convince or convert them.
I prefer to just let them be and not engage.
It's like showing someone from 1980 a modern smart phone and them saying, yeah but it can't read my mind.
This leads me to believe that the issue is not that llm skeptics refuse to see, but that you are simply unaware of what is possible without them--because that sort of fuzzy search was SOTA for information retrieval and commonplace about 15 years ago (it was one of the early accomplishments of the "big data/data science" era) long before LLMs and deepnets were the new hotness.
This is the problem I have with the current crop of AI tools: what works isn't new and what's new isn't good.
> It's so self-evident, that I don't know how to take the request for examples seriously
Do you see why people are hesitant to believe people with outrageous claims and no examples
They "hallucinate", they "know", they "think".
They're just the result of matrix calculus on which your own pattern recognition capacities fool you into thinking there is intelligence there. There isn't. They don't hallucinate, their output is wrong.
The worst example I've seen of anthropomorphism was the blog from a searcher working on adverse prompting. The tool spewing "help me" words made them think they were hurting a living organism https://www.lesswrong.com/posts/MnYnCFgT3hF6LJPwn/why-white-...
Speaking with AI proponents feels like speaking with cryptocurrencies proponents: the more you learn about how things work, the more you understand they don't and just live in lalaland.
Maybe hype is overly beneficial to them but if you promise me 1500 and I get 1100 then I will underwhelmed.
And especially around LLM marketing hype is fairly extreme.
When businessmen sell me "artificial intelligence", I come prepared for lots of fuckery.
Stitching together well-known web technologies and protocols in well-known patterns, probably a good success rate.
Solving issues in legacy codebases using proprietary technologies and protocols, and non-standard patterns. Probably not such a good success rate.
As the parent says, while far from perfect, they're an incredible aid in so many areas. When used well, they help you produce not just faster but also better results. The only trick really is that you need to treat it as a (very knowledgeable but overconfident) collaborator rather than an oracle.
I say "intern" in the sense that its error-prone and kind of inexperienced, but also generally useful. I can ask it to automatically create a lot of the bootstrapping or tedious code that I always dread writing so that I can focus on the fun stuff, which is often the stuff that's pawned off onto interns and junior-level engineers. I think for the most part, when you treat it like that, it lives up to and sometimes even surpasses expectations.
I mean, I can't speak for everyone, but whenever I begin a new project, a large percentage of the first ~3 hours is simply copying and pasting and editing from documentation, either an API I have to call or some bootstrapping code from a framework or just some cruft to make built-in libraries work how you want. I hate doing all that, it actively makes me not want to start a new project. Being able to get ChatGPT to give me stuff that I need to actually get started on my project has made coding a lot more fun for me again. At this point, you can take my LLM from my cold dead hands.
I do think it will keep getting better, but I'm also at a point where even if it never improves I will still keep using it.