I trusted an LLM, now I'm on day 4 of an afternoon project
nemo.foo
nemo.foo
I think the key to being successful here is to realize that you're still at the wheel as an engineer. The llm is there to rapidly synthesize the universe of information.
You still need to 1) have solid fundamentals in order to have an intuition against that synthesis, and 2) be experienced enough to translate that synthesis into actionable outcomes.
If youre lacking in either, youre at the same whims of copypasta that have always existed.
That's complicated, but I wouldn't say the resulting software is complex. You gave an LLM a repetitive, translation-based job, and you got good results back. I can also believe that an LLM could write up a dopey SAAS in half the time it would take a human to do the same.
But having the right parameters only takes you so far. Once you click generate, you are trusting that the model has some familiarity with your problem and can guide you without needing assistance. Most people I've seen rely entirely on linting and runtime errors to debug AI code, not "solid fundamentals" that can fact-check a problem they needed ChatGPT to solve first place. And the "experience" required to iterate and deploy AI-generated code basically boils down to your copy-and-paste skills. I like my UNIX knowledge, but it's not a big enough gate to keep out ChatGPT Andy and his cohort of enthusiastic morons.
We're going to see thousands of AI-assisted success stories come out of this. But we already had those "pennies on the dollar" success stories from hiring underpaid workers out of India and Pakistan. AI will not solve the unsolved problems of our industry and in many ways it will exacerbate the preexisting issues.
That was about 15 years ago, I no longer have the same enthusiasm you do.
Some of us are capable of wanting for things better than a coin-operated REST API. The kind of imagination used to put people on the moon, that now helps today's business leaders imagine more profitable ways to sell anime pornography on iPhone. (Don't worry, AI will disrupt that industry too.)
Basically if they are good at utilizing junior developers and interns or apprentices they probably will do well with an LLM assistant.
Since AI will capitulate and give you whatever you want.
You also have to learn how to ask without suggesting because it will take whatever you give it and agree.
I love LLMs. I agree with OP them expanding my hobby capacity as well. But I am constantly saying (in effect) “you sure…?” and tend to have a pretty good bs meter.
I’m still working to get my partner to that stage. They’re a little too happy to accept an answer without pushback or skepticism.
I think being ‘eager to accept an answer’ is the default mode of most people anyway. These tools are likely enabling faster disinformation consumption for the unaware.
Getting the LLM to pull out well-known names of concepts is for me the skill you can't get anywhere else. You can describe a way to complete a task and ask for what it's called and you'll be heading down arxiv links right away. Like yes the algorithm to find the closest in edit distance and length needle string in a haystack is called Needleman–Wunsch, of course Claude, everyone knows that.
Junior devs can get plenty of value out of them too, if they have discipline in how they use them - as a learning tool, not as a replacement for thinking about projects.
Senior devs can get SO much more power from these things, because they can lean on many years of experience to help them evaluate if the tool is producing useful results - and to help them prompt it in the most effective way possible.
A junior engineer might not have the conceptual knowledge or vocabulary to say things like "write tests for this using pytest, include a fixture that starts the development server once before running all of the tests against it".
* I put commercial there as a qualifier because there's some thought that in the future, very specifically-trained smaller models (open source) on particular technologies and corpuses (opt-in) might yield useful results without many of the ethical minefields we are currently dealing with.
On the other hand if you understand what needs to be done, and how to direct the work the productivity boost can be massive.
Claude 3.5 sonnet and O1 are awesome at code generation even with relatively complex tasks and they have a long enough context and attention windows that the code they produce even on relatively large projects can be consistent.
I also found a useful method of using LLMs to “summarize” code in an instructive manner which can be used for future prompts. For example summarizing a large base class that may be reused in multiple other classes can be more effective than having to overload a large part of your context window with a bunch do code.
If a topic is well represented in those places, then you will get your answer quicker and it can be to some extent shaped to your use case.
If the topic is not well represented there, then you will get circular nonsense.
You can say “obviously, that’s the training data”, and that’s true, and I do find it obvious personally, but the reaction to LLMs as some kind of second coming does not align with this reality.
So you're using it like Wikipedia? I find when learning something new (and non coding related) YouTube is infinitely better than an LLM. But then I prefer visual demonstration to tutorial or verbal explanation.
If I have a question about a first order language or framework feature or pattern, it works great. If I have a question about a second order problem, like an interaction between language or framework features, or a logical inconsistency in feature behavior, then it usually has no idea what’s going on, unless it turns out to be a really common problem such as something that would come up when working through a tutorial.
For code completion, I’ve just turned it off. It saves time on boilerplate typing for sure, but the actual content pieces are so consistently wrong that on balance I find it distracting.
Maybe I have a weird programming style that doesn’t mesh well with the broader code training corpus, not sure. Or maybe a lot of people spend more time in the part of problem-space that intersects with tutorial-space? I am not very junior these days.
That being said I definitely do use LLMs to engage with tutorial type content. For that it is useful. And outside of software it is quite a bit better for interfacing with Wikipedia type content. Except for the part where it lies to your face. But it will get better! Extrapolating never hurt anyone.
It's pretty good at writing screens in broad strokes. You will have to fill in some details.
The exact details of correctly threading data through; or prop drilling vs alternatives; the rules around wrapping screens to use them in React Navigation? It's terrible at them.
Did you puzzle about this sentence specifically? Imagine you don't know jack about flying helicopters, then Tank uploads the Helicopter Pilot Program (TM) directly to your brain; it would feel like magic.
Conversely, if you know a lot about helicopters, just not enough to fly a B-212, and the program includes instructions like "Press (Y) and Left Stick to hover", you'd know it's confabulating real world piloting with videogames.
That's the same with LLMs, you need to know a lot of the field to ask the right questions, and recognize/correct slop or confabulation, otherwise they seem much more powerful and smart than they really are.
> I think the key to being successful here is to realize that you're still at the wheel as an engineer. The llm is there to rapidly synthesize the universe of information.
Bingo. OP is like someone who is complaining about the tools, when they should be working on their talent. I have a LOT of hobbies (circuits, woodworking, surfing, playing live music, cycling, photography) and there will always be people who buy the best gear and complain that the gear sucks. (NOTE: I"m not implying claude is "the best gear", but it's a big big help.)
I think the only problem with LLMs is synthesis of new knowledge is severely limited. They are great at explaining things others have explained, but suck hard at inventing new things. At least that's my experience with Claude: it's terrible as a "greenfield" dev.
My college professor was also willing to say "I don't know, ask me next class"
I just asked Claude a question I am pretty sure was not in its training data.
https://www.quora.com/How-many-Humans-can-we-fit-on-the-Moon
Edit: couldn't resist, and dammit!!
Response: Ah, I see what you're doing! Since the Moon has no atmosphere, there’s technically no air to create any kind of airspeed velocity. So, the answer is... zero miles per hour. Unless, of course, you're asking about the speed of the horse itself! In that case, we’d just have to know how fast the astronaut can gallop without any atmosphere to slow them down.
But really, it’s all about the fun of imagining a moon-riding astronaut, isn’t it?
This is a perfect example of where not knowing the “domain” leads you astray. As far as I know “newborn width” is not something typically measured, so Claude is pulling something out of thin air.
Indeed you are showing that something not in the training data leads to failure.
Edit: it also doesn't account for the fact the moon is more or less a sphere, and not a flat plane.
This is a key differentiator that I see in humans over LLMs: knowing ones limits.
Granted, they are not (can't be) as rigorous as the tests your professor took, but new models are run through test suites before being released, too.
That being said, I saw my college professors making up things, too (mind you, they were all graduated from very good schools). One example I remember was our argument with a professor who argued that there is a theoretical limit for the coefficient of friction, and it is 1. That can potentially be categorised as a hallucination as it was completely made up and didn't make sense. Maybe it was in his training data (i.e. his own professors).
I agree with the "I don't know" part, though. This is something that LLMs are notoriously bad.
Also, you method for attesting your professors accuracy is inherently flawed. That little piece of paper on their wall doesn't correlate with how accurate they are; it doesn't mean zero, but it isn't foolproof. Hate to break it to you, but even heroes are fallible.
I don't use Claude, so maybe there's a huge gap in reliability between it and ChatGPT 4o. But with that disclaimer out of the way, I'm always fairly confused when people report experiences like these—IME, LLMs fall over miserably at even very simple pure math questions. Grammatical breakdowns of sentences (for a major language like Japanese) are also very hit-or-miss. I could see an LLM taking the place of, like, an undergrad TA, but even then only for very well-trod material in its training data.
(Or maybe I've just had better experiences with professors, making my standard for this comparison abnormally high :-P )
EDIT: Also, I figure this sort of thing must be highly dependent on which field you're trying to learn. But that decreases the utility of LLMs a lot for me, because it means I have to have enough existing experience in whatever I'm trying to learn about so that I can first probe whether I'm in safe territory or not.
For as rich a culture the Japanese have, there's only about 1XX million speakers and the size of the text corpus really matters here, the couple billion of English speakers are also highly motivated to choose English over anything else because Lingua Franca has homefield advantage
To use LLM's efectively you have to work with knowledge of their weaknesses, Math is a good example, you'll get better results from Wolphram Alpha even for the simple things, which is expected
Broad reasoning and explanations tend to be better than overly specific topics, the more common a language, the better the response If a topic has a billion tutorials online, an LLM has a really high chance of figuring out first try
Be smart with the context you provide, the more you actively constrain an LLM, the more likely it is to work with you I have friends that just use it to feed class notes to generate questions and probe it for blindspots until they're satisfied, the improvements on their grade s make it seem like a good approach, but they know that just feeding responses to the LLM isn't trustworthy, so they do and then they also check by themselves, the extra time valuable by itself, if just to improve familiarity with the subject
They are language models, not calculators or logic languages like Prolog or proof languages like Coq. If you go in with that understanding, it makes a lot more sense as to their capabilities. I would understand the parent poster to mean that they are able to ask and rapidly synthesize information from what the LLM tells them, as a first start rather than necessarily being 100% correct on everything.
YMMV. /shrugs/
But, once you've had AI help you solve some gnarly problems, it is hard not to be amazed.
And this is coming from a gal who thinks the idea of self-driving cars is the biggest waste of resources ever.
Sorry, maybe I should've been clearer in my response—I specifically disagree with the "college professor" comparison. That is to say, in the areas I've tried using them for, LLM's can't even help me solve simple problems, let alone gnarly ones. Which is why hearing about experiences like yours leaves me genuinely confused.
I do get your point about people disagreeing with modern AI for "political" reasons, but I think it's inaccurate to lump everyone into that bucket. I, for one, am not trying to make any broader political statements or anything—I just genuinely can't see how LLMs are as practically useful as other people claim, outside of specific use cases.
I think that its because often the libraries you use are niche or have a a few similar versions, the LLM really commonly hallucinated solutions and would continually suggest that library X did have that capability. I think because often in hardware projects you often hit a point where you can't do something or you need to modify a library, but the LLM tries to be "helpful" and it makes up a solution.
And then you’ll paste in the error, and they’ll just say “ok I see the problem” and output the exact same broken code lol.
I’m guessing the problem is lack of training data. Most TS codebases are mostly just JS with a few types and zod schemas. All of the neat generic stuff happens in libraries or a few utilities
The number of projects I’ve done where my notes are the difference between hours of relearning The Way or instant success. Google doesn’t work as some niche issue is blocking the path.
ESP32, Arduino, Home Assistant And various media server things.
Public Arduino, RPi, Pico communities are basically peak cargo cult, with the blind leading the blind through things they don't understand. The noise is vastly louder than the signal.
There's a basically giant chasm between expereinced or professional embedded developers that mostly have no need to ever touch those things or visit their forums, and the confused hobbyists on those forums randomly slapping together code until something sorta works while trying to share their discoveries.
Presumably, those communities and their internal knowledge will mature eventually, but it's taking a long long time and it's still an absolute mess.
If you're genuinely interested in embedded development and IoT stuff, and are willing to put in the time to learn, put those platforms away and challenge yourself to at least learn how to directly work with production-track SoC'a from Nordic or ESP or whatever. And buy some books or take some courses instead of relying on forums or LLM's. You'll find yourself rewarded for the effort.
It won't because the RPi are all undocumented, closed-source toys.
It would be an interesting experiment to see which chips an LLM is better at helping out with: RPi's with its hallucinatory ecosystem or something like the BeagleY-AI which has thousands of pages of actual TI documentation for its chips.
It would be really nice if the LLMs could cover for this and circumvent where RPi's keep getting used because they were dumped under cost to bootstrap a network effect.
I'm not sure they will. There's a kind of evaporative cooling effect where once you get to a certain level of understanding you switch around your tools enough that there's not much point interacting with the community anymore.
Eventually I read some actual documentation and realised it was just spouting very plausible sounding nonsense - and confident at it!
The same thing happened a year or so ago when I tried to get a much older ChatGPT to help me with with USB protocol problems in some microcontroller code. It just hallucinated APIs and protocol features that didn’t actually exist. I really expected more by now - but I now suspect it’ll just never be good at niche tasks (and these two things are not particularly niche compared to some).
For the best of both worlds make the LLM first 'read' the documentation, and then ask for help. Make a huge difference in the quality and relevance of the answers you get.
What is your main project ? Do you LLM that? (
I wager you're not a rust expert and should maybe reconsider using rust in your main project.
FWIW asking LLM whether you should use rust ~ asking it about the meaning of life. Important questions that need answers, but not right away! (A week or 2 tops)
If you need to synthesize the universe of information with LLM.. that is not the universe you want to live or play in
That's a nice way of putting it.
The very moment when you try to go off the beaten path and do something unconventional or stuff that most people won't have written a lot about, it gets more tricky. Just consider how many people will know how to configure some middleware in a Node.js project... vs most things related to hardware or low level work. Or even working with complex legacy codebases that have bits of code with obscure ways of interacting and more levels of abstraction that can be reasonably put in context.
Then again, if an LLM gets confused, then a person might as well. So, personally I try to write code that'd be understandable by juniors and LLMs alike.
It was so wrong that I wonder what version of the C standard it was even hallucinating.
counter point:
https://github.com/ggerganov/llama.cpp/pull/11453
> This PR provides a big jump in speed for WASM by leveraging SIMD instructions for qX_K_q8_K and qX_0_q8_0 dot product functions.
> Surprisingly, 99% of the code in this PR is written by DeekSeek-R1. The only thing I do is to develop tests and write prompts (with some trials and errors)
It is able to figure out some things that I know do not have much training data at all.
It is looking at the manual and figuring things out. "That doesn't make sense. Wait, that can't be right. I must have the formula wrong."
I just seen that in the chain of thought.
Sure, LLMs can improve but they're ultimately still bound by the constraints of the type of data they're trained on and don't actually build world models through a combination of high bandwidth exploratory training (like humans) and repeated causal inference.
That's an interesting thought. I think there are ways to automate this, and some IDEs / tools track this already. I've seen posts by both Google and Amz providing percentages of "accepted" completions in their codebases, and that's probably something they track across codebases automatically.
Also on topic, here's aider's "self written code" statistics: https://aider.chat/HISTORY.html
But yeah I agree that "written by" doesn't necessarily imply "autonomously", and for the moment it's likely heavily curated by a human. And that's still ok, IMO.
r = (rgba >> 24) & 0xff;
...and then pause, it's pretty good at guessing: g = (rgba >> 16) & 0xff;
b = (rgba >> 8) & 0xff;
a = rgba & 0xff;
... for the next few lines. I don't really ask it to do more heavy lifting than that sort of thing. Certainly nothing like "Write this full app for me with these requirements [...]"I have fought the "lowest cognitive load" code-style fight forever at my current gig, and I just keep losing to the "watch this!" fancytowne code that mids love to merge in. They DO outnumber me, so... fair deuce I suppose.
There is value in code being readable by Juniors and LLMs -- hell, this Senior doesn't want to spend time figuring out your decorators, needless abstractions and syntax masturbation. I just want to fix a bug and get on with my day.
Since it's just a placeholder I often ask for a funny twist but it's rarely ever anything like it.
But especially the ability to just see some of the stuff it produces, and now to see its thought process, is incredibly useful to me already. I do have autism and possibly ADD though.
I hope for a rennaisance of somewhat more rigorous programming languages: you can typecheck the LLM suggestions to see if they're any good. Also you can feed the type errors back to the LLM.
This is a good take that tracks with my (heavy) usage of LLMs for coding. Leveraging productive-but-often-misguided junior devs is a skill every dev should actively cultivate!
Which is fine, if it's a conscious choice for yourself.
Which I completely agree, I use LLMs for the cases where I do know what I'm trying to do, I just can't remember some exact detail that would require reading documentation. It's much quicker to leverage a LLM rather than going on a wild goose chase of the piece of information I know exists.
Also it's a pretty good tool to scaffold the boring stuff, asking a LLM "generate test code for X asserting A, B, and C" and editing it to be a proper test frees up mental space for more important stuff.
I wouldn't trust a LLM to generate any kind of business logic-heavy code, instead I use it as a quite smart template/scaffold generator.
I still have the skills to search the web if the magic piano disappears.
Don't know why you are trying to come up with a situation that doesn't exist, what's your point exactly against this quite narrow use-case?
Feels like this is only worthwhile because the junior dev learns from the experience; an investment that yields benefits all around, in the broad sense. Nobody wants a junior around that refuses to learn in perpetuity, serving only as a drag on productivity and eventually your sanity.
There's still incredible accumulated value here, but it's at the other end. The more times you successfully use an LLM to produce working code, the more you learn about how to use them - what they're good at, what they're bad at, how to effectively prompt them.
https://www.youtube.com/live/outcGtbnMuQ?si=oTMA02ns_BJDRS4c...
Advances since then have indeed been remarkable.
Six months ago, I tried building this app with ChatGPT and got nowhere fast.
Building it with Claude required a gluing together a few things that I didn't know much about: JavaScript audio processing, drawing on a JavaScript canvas, an algorithm for bilinear interpolation.
I don't write JavaScript often. But I know how to program and I understand what I'm looking at. The project came together easily and the creative momentum of it felt great to me. The most amazing moment was when I reported a bug—I told Claude that the audio was stuttering whenever I moved the controls—and it figured out that we needed to use an AudioWorklet thread instead of trying to play the audio directly from the React component. I had never even heard of AudioWorklet. Claude refactored my code to use the AudioWorklet, and the stutter disappeared.
I wouldn't have built this without Claude, because I didn't need it to exist that badly. Claude reduced the creative inertia just enough for me to get it done.
That's the next step for me in learning AI... playing with different integrated editor tools.
I do see the LLMs ingesting more and more documentation and content and they are improving at giving me right answers. Almost two years ago I don't believe they had every python package indexed and now they appear to have at least the documentation or source code of it.
So it's handy to get a quick list of "all packages which do X", but it's worse then useless to have it speculate as to which one to use or why, because of the hallucination problem.
The edge between being actually more productive or just “pretend productive” using large language models is something that we all haven’t completely figured out yet.
I use an NPM script to automate concatenating the spec + source files + prompt, which I then copy/paste to o1. So far this has been working somewhat reliably for the early stages of a project but has diminishing returns.
Aider also has a copy/paste mode to use web ui interfaces/subscriptions instead apis.
I definitely use and update my CONVENTIONS.md files and started adding a second specification file for new projects. This + architect + "can your suggestion be improved, or is there a better way?" has gotten me pretty far.
It’s incredibly worrying that it needs to be explained again and again that LLMs are different from people, do not behave like people, and should not be compared to people or interacted like people, because they are not people.
I suppose the wide range of negative and positive experiences people seem to have working with LLMs is related to the wide range of expectations people have for their interactions in general.
I can also tell when it’s stuck in some kind of context swamp and won’t be any more help, because it will just keep making the same stupid mistakes over and over and generally forgetting past instructions.
At that point I take the last working code and paste it into a new chat.
I have a custom prompt that instructs gpt4o to get aggressive about attacking anything I say (and, importantly, anything it says).
Here's my result for the same question:
https://chatgpt.com/share/67984aa9-1608-8012-be93-a77728ab8e...
I've been doing that since way before LLMs were a thing.
Like many others are saying, you need to be in the drivers seat and in control. The LLM is not going to fully complete your objectives for you, but it will speed you up when provided with enough context, especially on mundane boilerplate tasks.
I think the key to LLMs being useful is knowing how to prompt with enough context to get a useful output, and knowing what context is not important so the output doesn’t lead you in the wrong direction.
USB-CDC is cooler than that, you can make the Pico identify as more than just one device. E.g. https://github.com/Noltari/pico-uart-bridge identifies as two devices (so you get /dev/ttyACM0 and /dev/ttyACM1). So you could have logs on one and image transfers on another. I don't think you're limited to just two, but I haven't looked into it too far.
You can of course also use other USB protocols. For example you could have the Pico present itself as a mass-storage device or a USB camera, etc. You're just limited by the relatively slow speed of USB1.1. (Though the Pico doesn't exactly have a lot of memory so even USB1.1 will saturate all your RAM in less than 1 second)
Made me wanna join in your garage and help out with the project :)
Cursor & Claude got the boilerplate set up, which was half the mental barrier. Then they acted as thought partners as I tried out various implementations. In the end, I came up with the algorithm to make the thing performant, and now I'm hand-coding all the shader code—but they helped me think through what needed to be done.
My take is: LLMs are best at helping you code at the edge of your capabilities, where you still have enough knowledge to know when they're going wrong. But they'll help you push that edge forward.
I asked claude to write the initial version. It came up with a complicated class based solution. I spent more than 30 minutes getting a good abstract to come out. I was copy pasting typescript errors and applying fixes it suggested without thinking much.
In the end, I gave up and wrote what I wanted myself in 5 minutes.
0] https://github.com/cloudycotton/browser-operator/blob/main/s...
It has been my experience 1 code clown can poison a project with dozens of reasonably talented engineers active. i.e. clowns often go through the project smearing bad kludges over acceptable standards to appear like their commit frequency means something.
This is why most developers secretly dream of being plumbers. Good luck, =3
If .com, email me (it's in my profile) and I can see if there is a reason your account is getting so heavily captcha'd.
A bit on a tangent, but has there been any discussion of how junior devs in the future are ever going to get past that stage and become senior dev calibre if companies can replace the junior devs with AIs? Or is the thinking we'll be fine until all the current senior devs die off and by then AI will be able to replace them too so we won't need anyone?
1. CS/Eng degree 2. ??? 3. Senior dev!
Companies are not saving money by paying for AI tools if they continue to hire the same number of people. The only way it makes financial sense, and for the enormous amounts of money being invested into AI to reap profits, is if companies are able to reduce the cost of labor. First, they only need 75% of the junior devs they have now, then 50%, then 25%.
It won't happen all at once, and as tasks done by current juniors are incrementally taken over by AI, the expected entry skillset will evolve in line with those changes. There will always be junior people in the field, but their expected knowledgebase and tasks will evolve, and even if 100% of the work currently done by juniors is eventually AI-ified, there will still be juniors, they just will be doing completely different things, and going through a completely different learning process to get there.
> Companies are not saving money by paying for AI tools if they continue to hire the same number of people.
Companies which have a fixed lump of tech work (in practice, none, actually) will save money because they will hire fewer total workers because output per worker will increase, but they will still have people who are newer and more experienced within that set, because the
More realistic companies that either make money with tech work or that apply internal effort to tech as long as it has net positive utility may actually end up spending more on tech, because each dollar spent gives more results. This still saves money (or makes more money), but the savings (where it is about savings, and not revenue) will be in the areas tech is applied to, not tech itself.
I don't see AI replacing that. AI is a tool with the instant Q&A intelligence of a junior dev but it's not actually doing the job of a junior dev. That's a subtle distinction.
Most demeaning and depressingly toxic thing I've read today...
I’ve come to a similar conclusion - for now at least it’s best applied at a fairly granular level. Make me a red brick wall there rather than „hey architect make me a house“.
I do think OP tried a bit too much new stuff in one go though. USB plus zig is quite a bit more ambitious than the traditional hello world in a new lang
Imma gonna have to work this into a convo some day. Just to see the “wait, what??” expressions on people’s faces.
I love it!
* Get AI to write tests
* Use copy/paste. No IDE
* Use python (not because it's better than zig)Learning how to build or create with a new kind of word processor is a skill unto itself.
We get it. They’re not superintelligent at everything yet. They couldn’t infer what you must’ve really meant in your heart from your initial unskillful prompt. They couldn’t foresee every possible bug and edge case from the first moment of conceptualizing the design, a flaw which I’m sure you don’t have.
The thing that pushes me over the line into ranting territory is that computer programmers, of all people, should know that computers do what you tell them to.
You've been here since 2016 and this is the kind of posting that finally gets to you? How in the world have you avoided all the shitposts in the last decade? What is your secret?
I’m currently using them to port a client-side API SDK into multiple languages. This would be a pain in the ass time consuming task but is a breeze with LLMs because the exact behavior I want is clearly defined and relatively deterministic, and it’s also straightforward to test that I’m getting what I intend. The LLM thus gets done in 3 days what would take me 3 weeks (or more) to do by hand.
If the complaint is that it can’t do X, where X is something that would clearly require full AGI and likely true superintelligence — in this case expecting instantaneous, correct code that solves novel problems on the first try - then I have to insist that people are actually expecting Claude to be a Culture Ship Mind, implicitly. They just don’t realize that what they’re asking for his hard, which is itself a psychologically interesting fact, I suppose.
I will argue the opposite of that forever. They're very evidently useful, if you take the time to learn how to apply them.
are you claiming LLMs function like computer program instructions? like they clearly don't operate like that at all.
I think LLMs have uncovered what we have always known in this industry: that people are, by default, bad at communicating their intent clearly and unambiguously.
If you express your intent to an LLM with sufficient clarity and disambiguation, it will rarely screw up. Often, we don’t have time to do this, and instead we aim for the sweet spot of sufficient but not exhaustive clarity. This can be fine if you are experienced with that particular LLM and you have a good feel for where its sweet spot actually is. If you miss that target, though, the LLM will not correctly infer your intended subtext. This is one of the things that requires experience. In fact, even the “same” LLM will change in its behavior and capabilities as it undergoes fine tuning. Sometimes it will even get worse at certain things.
All of this is to say, of course, you’re right that it’s not a compiler. But I think people fail in their application of LLMs for much the same reason that novice coders fail to get compilers to guess what they intended.
If those are your only two reference points, yes they're closer to the former.
But the biggest problem is how much "pixie that does something you neither wanted nor asked for" gets mixed in. And I think a lot of the complaints you're saying are about lack of mind reading are actually about that problem instead.
Right. The problem isn't that the tool isn't perfect, it's that you get a lot of excitable people with incentives pretending that it is or will soon be perfect (while simultaneously scaring non-technical people into thinking they'll be replaced with a chat bot soon).
There are certainly luddite types who are outright rejecting these tools, but if you have hands-on, daily experience, you can see the forest for the trees. You quickly realize that all of the "omg this thing is sentient" or "we can't let what we've got into the world, it's too dangerous" fodder like the Google panic memo are just covert marketing.
"I learned that I need to stay firmly in the driver’s seat when tackling new tech."
Er, that's pretty much what a pilot is supposed to do! You can't (as yet) just give an AI free reign over your codebase and expect to come back later that day to discover a fully finished implementation. Maybe unless your prompt was "Make a snake game in Python". A pilot would be supervising their co-pilot at all times.
Comparing AIs to junior devs is getting tiresome. AIs like Claude and newer versions of ChatGPT have incredible knowledge bases. Yes, they do slip up, especially with esoteric matters where there are few authoritative (or several conflicting) sources, but the breadth of knowledge in and of itself is very valuable. As an anecdote, neither Claude nor ChatGPT were able to accurately answer a question I had about file operation flags yesterday, but when I said to ChatGPT that its answer wasn't correct, it apologised and said the Raymond Chen article it had sourced wasn't super clear about the particular combination I'd asked about. That's like having your own research assistant, not a headstrong overconfident junior dev. Yes, they make mistakes, but at least now they'll admit to them. This is a long way from a year or two ago.
In conclusion: don't use an AI as one of your primary sources of information for technology you're new to, especially if you're not double-checking its answers like a good pilot.
LLMs are jittery apprentices. They'll hallucinate measurements, over-sand perfectly good code, or spin you in circles for hours. I’ve been there back in the GPT-4 days especially, nothing stings like realising you wasted a day debugging AI’s creative solution to a problem you could've solved in 20 minutes.
When you treat AI like a toolbelt, not a replacement for your own brain? Magic. It’s killer at grunt work like; explaining regex, scaffolding boilerplate, or untangling JWT auth spaghetti. You still gotta hold the blueprint. AI ain't some magic wand: it’s a nail gun. Point it wrong, and you’ll spend four days prying out mistakes.
Sucks it cost you time, but hey, now you know to never let the tool work you. It's hopefully a lesson OP learns once and doesn't let it sour their experience with AI, because when utilised properly, you can really get things done, even if it's just the tedious/boring stuff or things you'd spend time Google bashing, reading docs or finding on StackOverflow.
For anything remotely complex, this is dead on. I use various models daily to help with coding, and more often than not, I have to just DIY it or start brand new chats (because the original context got overwhelmed and started hallucinating).
This is why it's incredibly frustrating to see VCs and AI founders straight-up gaslighting people about what this stuff can (or will) do. They're trying to push this as a "work killer," but really, it's going to be some version of the opposite: a mess creator that necessitates human intervention.
Where we're at is amazing, but we've got a loooong way to go before we can be on hover crafts sipping sodas Wall-E style.
No design. Hardware & software. 2 different platforms. A new language. Zig. Unrealistic time expectations.
A senior SWE would've still tanked this, just in different ways.
Personally, I'd still consider it a valuable experiment, because the lessons learned are really valuable ones. Enjoy round 2 :)
Experienced folks aren't surprised by this. LLMs are fast for boilerplate, research, and exploring ideas, but they're not autonomous coders. The key is you staying in charge: detailed prompts, critical code review, iterative refinement. Going back to web interfaces and manual pasting because editor integration felt "too easy" is a massive overcorrection. It's like ditching cars for walking after one fender bender.
Ultimately, this wasn't an AI failure, it was an inexperienced user expecting too much, too fast. The "lessons learned" are valid, but not AI-specific. For those who use LLMs effectively, they're force multipliers, not replacements. Don't blame the tool for user error. Learn to drive it properly.
Learning to properly prompt an LLM to get a net gain in value is a skill in it of itself.
Microsoft’s offering is literally called “copilot”. That is exactly what they’re marketing it as.
A junior dev faking competence while plagiarizing like crazy.
The plagiarizing part is why the junior dev from hell might not get fired: laundering open source copyrights can have beancounter alignment.