We already put natural language between us and the bytes. Hence why most keywords and variable names (a hard part of computer science) are in simple English and it is considered a net positive.
We already put natural language between us and the bytes. Hence why most keywords and variable names (a hard part of computer science) are in simple English and it is considered a net positive.
Since then, I spent a week trying to get cursor to work, and after dealing dealing with all the bugs, and restarting the composer each time with a new prompt, was able to get what I would consider a quality output for a moderately complex app (a parimutuel betting market).
The issue isn’t that LLMs are terrible, it’s the software like cursor is buggy and poorly written.
It should know that I don’t want to use code from an old version of the library I am using because the new library I am using is already in my projects dependencies.
It should let me set up preferences for different programming languages. And preferences for all programming languages.
So when I give it a prompt, it looks at the dependencies and language rules I already have set up, adds those to the prompt and produces the quality output I’m seeing now without me having to manually specify all those things.
Short version: LLMs rule the software is just shitty.
I was able to write a plugin for ComfyUI (a 60k loc python/js codebase) in 2 hours thanks to semantic search. It's not an exercise I'm versed in.
It wasn't that different from the kind of internal monologue I'd have held in my head had I done it on my own, including misguided confidence that gets crushed 5 minutes later as you read other parts of the code that show you had the wrong understanding of how it actually works.
In this context, LLMs can be very useful because a ground truth already exists to compare their replies against.
That sounds like a similar finding to what I had (comparing to copilot in my own case).
My point, which my post maybe didn’t make so well, was a huge amount of the prompt should’ve been written for me in order to get to an acceptable result sooner.
Similar experience trying to use GenAIScript, btw, and peering inside the box the code and product is pretty well incomprehensible.
Yes, the AI is writing the buggy parts that upset me but my point was creating a good quality prompt would’ve taken a lot less time if Cursor had had some reasonable defaults.
There is the story that von Neumann flew off the handle the first time he saw an assembler.
>>How dare you waste compute cycles on this frivolity? Just use machine code like everyone else.
That was in the 1940s when labour was very cheap and compute was insanely expensive. We’re talking hundreds to thousands of programmers’ salaries for the cost of one computer.
No, AI is a shitty tool that has yet to prove its utility. Autocomplete works by analyzing the official API and interface, it's completely different than AI which hallucinates meaning between words and also stuff that it was fed before it met you.
> variable names (a hard part of computer science)
Naming is for software engineering, not CS. One more confusion by people who want to sell us AI at all cost.
You can (and should) give the AI access to your existing codebase and any relevant documentation to use as context if you want good results. If you give the AI zero context for the problem it is trying to solve, of course it will struggle. If you give it all the necessary context, it will do much better.
I've found that just uploading the documentation of the API or library you are working with before asking the AI questions about it makes a huge difference in the quality of its output.
But modern AI tools are far beyond "auto complete". (I actually turn off those in-line completions, I feel they ruin flowstate). The tools now are fully prompted, with multi-file editing, with full codebase context, with web/search and doc integration, and for "on the rails" development are producing high quality code for "easier" tasks.
These modern models and tools can solve nearly every single leet code problem faster than you. They can do every single Advent of Code problem likely 10X-100X faster than you can.
In my professional, high standards, very legal and contract driven web app world, AI tools are still very useful for doing "on the rails" development. Is it architecting entire systems? No of course not (yet). Is it emulating existing patterns and extending them for new functionality 10X faster than a Jr or Mid? Yes it is. Is it writing nearly perfect automated tests based on examples? Yes it is. It is scaffolding new ideas and putting down a great starting point? Yep. And it's even able to iterate on featurework pretty well, and much faster than Jr/Mid.
The kind of work I'd give to a Jr/Mid and expect to take 2-3 days before they need serious feedback up and down the change, these AI are doing in about 30 seconds, maybe 90 seconds if you need to iterate a few times on the prompt.
I get that "AI" is a buzzword that is pumping valuations and making business people see $$$.
But coding assistants are not that. For many programmers, they are quickly becoming valuable tools that do in fact speed up development.
That's expected, since all the leetcode problems have ready-to-use solutions on the internet.
(In fact, the reason they ask leetcode questions isn't to test your IQ, it's to know if you've read the obvious and available literature.)
>That's expected, since all the leetcode problems have ready-to-use solutions on the internet.
1) If the implication is "The model knows the answer and regurgitates it like lyrics to a song" then I would push back. Put a leet code problem into deepseek r1 chain-of-reasoning model and watch it spend 2 minutes spitting out 5000 words thinking through every single facet of the problem and genuinely solving it at a level that is higher than 95% of programmers.
And point 2)
If you do believe it's fundamentally about how much the model has been trained on, then it has seen your CRUD app and has already seen 10,000 times the feature or system you're about to write -- so it should be a foregone conclusion that it can also do all of that development work too. Only the higher order architecting and proprietary domains should be challenging for it, as there would be far less examples to train on (scarcity) or the model doesn't understand a complex solution (architecting systems at scale is something it can't do).
(I also point out how well these models did for Advent of Code 2024, when there were zero examples in the training data for it).
This one is funny because for something like leet code, nearly everyone just reads the best answers, learns them and learns how to regurgitate them in an interview environment.
It's not "thinking", it's regurgitating an internet search and padding it out with markov-chain style text autocompletion.
The 5000 words of padding do not actually provide any value, it's verbal white noise to fill space.
> ...then it has seen your CRUD app and has already seen 10,000 times the feature or system you're about to write
Well, yes. Lots of pointless waste in software engineering. Fortunately I don't write CRUD apps and AI does nothing at all for me in a professional context.
Rather: it tests whether you are sufficiently docile and devoted to be willing to cram lots of leetcode exercise books that have no relevance for the programming concepts that the job will involve, just for a lottery ticket for a somewhat well-paid position.
I know that there is so much more to programming and related topics that is sooo much deeper (in particular if non-trivial mathematics becomes involved) than these leetcode-style brainteasers. So I strongly prefer to read about such deeply intellectually inspiring topics related to programming instead of jumping through the idiotic hoops that other people want me to.
Indeed, I thus fail the test for docility and devotedness, but I honestly can't take organisations seriously that demand such jumping through hoops.
Claude in 2025, especially with Project feature is far better. It can complete CRUD project on its' own, and all I have to do is to fix glaring issues and design API before.
Which might be not impressive to someone, but it is good at that. And few years ago, it would not be possible.
Then the project management tooling does a lot more, like automatically reverse engineer existing databases and so on.
That's exactly what happens, and why I think the whole hype is a joke. I have tried all the models and tools though, it's always an annoying mess.
> tools can solve nearly every single leet code problem faster than you
That would be useful if I was paid to "leet code" or solve Christmas games. This is not a good rebutal though but it made me smile.
> The kind of work I'd give to a Jr/Mid
Good, but I don't want to know what happens in 20 years when there are no more juniors to feed the AI and work on becoming seniors. I will be retired by then and I'll enjoy writing my own open-source stuff.
>Naming is for software engineering, not CS.
I figured they were referencing the “two hard problems of computer science”, those two being naming things, cache invalidation and off by one errors.
Everybody knows the hardest problems in software engineering are assembling promo packets and building consensus on number of spaces per indent.
As I'm not a native English speaker, I disagree. I learned programming long before I got decent in English, and even today I just consider the English keywords in programming languages to be some "abstract mathematical concept" that by mere coincidence is named after some real, existing English word. Even today, being somewhat decent in English, I stil think this way when I see program code.
I actually would insist that this is a much more useful way to think about good programming, since this way you have no difficulties to ask yourself all the time whether it would make sense to replace some "English-named" concept by something more useful, but which has no analogue in the English language (or any other natural language).
"Natural language" is about far more than individual words.
Not to mention that our existing programming languages have a deterministic output given the same code and the same compiler.
LLMs do not.
Thus, LLM prompts are an entirely different class of tool than a programming language.
This should be obvious to anyone who has written code, but alas.