The end result? Non-mainstream languages don't get much easier to get into because average Joe isn't already proficient in them to catch AI's bugs.
People often forget the bitter lesson of machine learning which plagues transformer models as well.
The end result? Non-mainstream languages don't get much easier to get into because average Joe isn't already proficient in them to catch AI's bugs.
People often forget the bitter lesson of machine learning which plagues transformer models as well.
Just found 3 race conditions in 100 lines of code. From the UTF-8 emojis in the comments I'm really certain it was AI generated. The "locking" was just abandoning the work if another thread had started something, the "locking" mechanism also had toctou issues, the "locking" also didn't actually lock concurrent access to the resource that actually needed it.
This is one of the "here be demons" type signatures of LLM code generation age, along with comments like
// define the payload struct payload {};
I've copied thousands of lines of complex code into an LLM asking it to find complex problems like race conditions and it has found them (and other unsolicited bugs) that nobody was able to find themselves.
Oh it just pattern matched against the general concept of race conditions to find them in complex code it's never seen before / it's just autocomplete, what's the big deal? At that level, humans are glorified pattern matchers too and the distinction is meaningless.
The counter point is how LLMs can't find a missing line in a poem when they are given the original.
PAC learning is basically existential quantification...has the same limits too.
But being a tool to find a needle is not the same as finding all or even reliability finding a specific needle.
Being being a general programming agent requires much more than just finding a needle.
True, but describing a limitation of the tech can't be used to make the sort of large dismissals we see people make wrt LLMs.
The human brain has all sorts of limitations like horrible memory (super confident about wrong details) and catastrophic susceptibility to logical fallacies.
Have you not had this issue with LLMs? Because I have. Even with the latest models.
I think someone upthread was making an attempt at
> describing a limitation of the tech
but you keep swatting them down. I didn’t see their comments as a wholesale dismissal of AI. They just said they aren’t great at sufficiently complex tasks. That’s my experience as well. You’re just disagreeing on what “sufficiently” and “complex” mean, exactly.
I'm still convinced that this is true. The more advances we make in "AI" the more i expect we'll discover that we're not as creative and unique as we think we are.
Humans generally have a very high opinion of themselves and their supposedly unique creative skills. They are not eager to have this illusion punctured.
Even our memories are mutable. We will with full confidence recite memories or facts we've learned just moments ago which are entirely fictional. Normal, healthy adults.
[1]https://ai.vixra.org/pdf/2506.0065v1.pdf
The paper is satire, but it's a pretty funny read.
“Pattern matching” is thought of as linear but LLMs are doing something more complex, it should be appreciated as such.
How did you evaluate this? Would be interested in seeing results.
I am specifically interested in the amount of false issues found by the LLM, and examples of those.
There are false positives, and they mostly come from the LLM missing relevant context like a detail about the priors or database schema. The iterative nature of an LLM convo means you can add context as needed and ratchet into real bugs.
But the false positives involve the exact same cycle you do when you're looking for bugs yourself. You look at the haystack and you have suspicions about where the needles might be, and you verify.
You do or you don't.
Recently we've seen many "security researchers" doing exactly this with LLM:s [1]
1: https://www.theregister.com/2025/05/07/curl_ai_bug_reports/
Not suggesting you are doing any of that, just curious what's going on and how you are finding it useful.
> But the false positives involve the exact same cycle you do when you're looking for bugs yourself.
In my 35 years of programming I never went just "looking for bugs".
I have a bug and I track it down. That's it.
Sounds like your experience is similar to using deterministic static code analyzers but more expensive, time consuming, ambiguous and hallucinating up non-issues.
And that you didn't get a report to save and share.
So is it saving you any time or money yet?
It's a hard, time consuming, and meandering process to do this kind of work on a system, and it's what you might have to pay expensive consultants to do for you, but it's also how you beat an expensive bug to the punchline.
An LLM helps me run all sorts of considerations on a system that I didn't think of myself, but that process is no different than what it looks like when I verify the system myself. I have all sorts of suspicions that turn into dead ends because I can't know what problems a complex system is already hardened against.
What exactly stops two in-flight transfers from double-spending? What about when X? And when Y? And what if Z? I have these sorts of thoughts all day.
I can sense a little vinegar at the end of your comment. Presumably something here annoys you?
Thanks for your responses.
Really sorry about the vinegar, not intentional. I may have such personality disorder idk. Being blunt, not very great communication skills.
My vice is when someone writes a comment where I have a different opinion than them, and their comment makes me think of my own thoughts on the subject.
But since I'm responding to them, I feel like it's some sort of debate/argument even though in reality I'm just adding my two cents.
That's not really true for things that are changing a lot. I got a terrible experience last time I've tried to use Zig, for example. The code it generated was an amalgamation between two or three different versions.
And I've even got this same style of problem in golang where sometimes the LLM generates a for loop in the "old style" (pre go 1.22).
In the end LLMs are a great tool if you know what needs to be done, otherwise it will trip you up.
Things that most teams don’t do or half-ass
I use AI assistance to generate code and review code, but I haven't had success trying to use it to update a substantial existing code base in Scala. I have tried using both Claude Code and Cursor and the handful of times I tried there were so many oversights and mistakes that resolving the mess was more effort than doing it manually. I'll probably try again after the next big model releases.
Current frontier models have been the least useful to me when I'm asking them to review performance-critical code inside Scala. These are bits of my code base that I have written to use a lot of mutable variables, mutable data structures, and imperative logic to avoid allocations and other performance impediments. I only use this style for functions highlighted by profiling. I can ask for a code review and even include the reasoning for the unusual style in the docstring, but Claude and Gemini have still tried to nudge me back to much slower standard Scala idioms.
im building one that lets you write and enforce your own rules so you don't get the typical slop
email in profile if you'd like to try it - i can send you a link
Strong typing drastically reduces hallucinations and wtf bugs that slip through code review.
So it’ll probably be the strongly typed languages that receive the proportionally greatest boost in popularity from LLM-assisted coding.
goimports makes everything look the same, the compiler is a nitpicky asshole that won’t let the program even compile if there is an unused variable etc.
That is a really big advantage in the AI era. LLMs are pretty bad at identifying what is and what isn't relevant in the context.
For developers this decision is pretty annoying, but it makes sense if you are using LLMs.
When you have a standard build process that runs go vet, go test, golanci-lint, goimports and compiles the code you can order the LLM to do that as the last step every time.
This way at the very least the shit it produces is well-formed and passes the tests :)
Otherwise they tend to just leave stuff hanging like "this erroring test is unrelated to the current task, let's just not run it" - ffs you just broke it, it passed perfectly before you started messing with the codebase =)
If they do want "popularity" then the counter-counter-point is that it should be easier to get than ever. Just have one proficient person write a lot of idiomatic, relatively isolatable code, and then have an AI generate terabytes upon terabytes of public domain licensed variations and combinations on that code. If you make programming in the small a breeze, people will flock to your language, and then they can discover how to program in the large with it on their own time.
I'm optimistic that most new programming languages will only need a few "real" programmers to write a small amount of example code for the AI training to get started.
CSV is not a complex format.
Why do you reach this conclusion from toying with CSV?
And why do you trust a LLM for economic planning?
When the code is done, it not like the LLM can secretly go flip columns at random
Even for small projects the optimisation criteria is different if the human's role in the equation shifts from authoring to primarily a reviewing based one.
Alternatively every new release will have to come with an MCP for its documentation and any other aspects that might make it easier for an LLM to talk about it and use it accurately.
Also, I had created a custom Node.js/JavaScript BaaS platform with custom Web Components and wanted to build apps with it, I gave it the documentation as attachment and surprisingly, it was able to modify an existing app to add entire new features. This app had multiple pages and Claude just knew where to make the changes. I was building a kind of marketplace app. One time it implemented the review/rating feature in the wrong place and I told it "This rating feature is meant for buyers to review sellers, not for sellers to review buyers" and it fixed it exactly right.
I think my second experience (plain JavaScript) was much more impressive and was essentially frictionless. I can't remember it making a single major mistake. I think only once it forgot to add the listener to handle the click event to highlight when a star icon was clicked but it fixed it perfectly when I mentioned this. With TypeScript, it sometimes got confused; I had to help it a lot more because I was trying to mock some functions; the fact that the TypeScript source code is separate from the build code created some confusion and it was struggling to grep the codebase at times. Though I guess the code was also more complicated and spread out over more files. My JavaScript web components are intended to be low-code so it's much more succinct.
Zig changes a lot. So LLMs reference outdated data, or no data at all, and resort to making a lot of 50% confidence guesses.
With Rust OTOH Claude feels like a great teacher.
I also use coding agents with Elixir daily without issues.
some HDLs should fit the bill: VHDL, Verilog or SystemC
[1] https://docs.amd.com/r/en-US/ug901-vivado-synthesis/Flip-Flo...
Rust code has the property that if it compiles, it usually works. True there are still runtime errors that can occur in Rust, but they're less likely going to be due to LLM hallucinations, which would be caught at compile time.
Can you elaborate a bit here? In my experience, most code I come into contact with isn't nearly defensive enough. Is AI generated code more defensive then the median?
OpenAI uses Rust for their service development as do a lot of other big companies.
It's a lot like Python/Flask, or even a bit like Go. It's incredibly easy to author [1] and deploy, and it runs super fast with no GC spikes or tuning. Super predictable five nines.
Desktop apps sing when written in Rust. A lot of AI powered desktop apps are being written in Rust now.
If you're going to reach for Go or Java (gRPC or Jetty or something) or Python/Flask, Rust is a super viable alternative. It takes the same amount of time to author, and will likely be far more defect free since the language encourages writing in a less error prone way and checks for all kinds of errors. Google did a study on this [2,3].
[1] 99.9% of the time you never hit the borrow checker/lifetimes when writing server code as it's linear request scoped logic. You get amazing error handling syntax and ergonomics and automatic cleanup of everything. You also have powerful threading and async tools if you need your service to do work on the side, and those check for correctness.
[2] "When we've rewritten systems from Go into Rust, we've found that it takes about the same size team about the same amount of time to build it," said Bergstrom. "That is, there's no loss in productivity when moving from Go to Rust. And the interesting thing is we do see some benefits from it. So we see reduced memory usage in the services that we've moved from Go ... and we see a decreased defect rate over time in those services that have been rewritten in Rust – so increasing correctness." https://www.theregister.com/2024/03/31/rust_google_c/
In which world is Rust fantastic for writing desktop applications? Where are the mature Rust UI frameworks?
> Desktop apps sing when written in Rust.
What does this even mean?
> A lot of AI powered desktop apps are being written in Rust now.
For example? And what do you mean by "AI powered desktop apps"?
Rust has come a remarkably long way [1] !
It's better than any language that isn't C/C++. We have bindings to all of the major UI toolkits, plus a lot of native toolkits.
You can also use Electron/Tauri for Javascript, or Dioxus for something more performant. Egui is also really nice for dev tools.
> What does this even mean?
Rewrite It In Rust metrics tend to tell good stories.
Developer blogs (links escaping me right now) show positive performance gains for lots of consumer desktop software written in whole or in parts using Rust. Discord, Slack, lots of other apps are starting to replace under-performing components in Rust.
> For example? And what do you mean by "AI powered desktop apps"?
Stuff like Zed [2] and the open source Photoshop-killer I'm working on (user-guided volumetric rendering, real time instructive Figma). The creator or egui works on Rerun [3], which is industrial / spatial visualization. Etc, etc.
[2] https://zed.dev/
The only mature cross-platform UI frameworks I see in this list are written in C and C++ :) And of all the various Rust bindings for Qt, only a single one ("ritual") offers bindings to the Widgets API. Not only are these bindings unsafe, the project has also been abandoned. https://github.com/KDAB/cxx-qt/?tab=readme-ov-file#compariso.... Of course, this is not suprising, given that the API surface of Qt6 is huge and difficult (or impossible?) to safely wrap in Rust.
(The reality is that Rust's ownership model makes it rather awkward to use for traditional UI frameworks with object trees.)
Tauri looks neat, but the actual UI is a webview, so it's not everyone's cup of tea. (Certainly not mine.)
As you said, egui is cool for simple tools, but not suitable for complex desktop applications.
Yes, Rust has gotten quite a few desktop UI frameworks, but I don't see any industry standards yet. Don't forgot that Qt has been in development for 30 years now! Rust has a lot to catch up.
Obviously, this can also happen with Rust, but certainly less so than with C or C++, I'll give you that. But how is Rust the only way when there's also Java, Go, C#, Swift, Kotlin, etc.?
But I'm not noticing that anymore, at least with Elixir. The gap has closed; Claude 4 and Gemini 2.5 both write it excellently.
Otoh, if you wanted to create an entirely new programming language in 2025, you might be shit outta luck.
This just made me really sad. That effectively means that we'll plateau indefinitely as a civilization (not just on programming languages, but anything where the LLM can create an artificial Lindy effect).
Or, do what you want, and don't worry what language people who have no interest in programming use.