Then there's chain-of-thought being positioned as the next big step forwards, which works by throwing more inferencing at the problem, so that cost can't be amortized over time like training can...
Then there's chain-of-thought being positioned as the next big step forwards, which works by throwing more inferencing at the problem, so that cost can't be amortized over time like training can...
It was an entire web app, with search filters, tree based drag and drop GUIs, the backend api server, database migrations, auth and everything else.
Not once did he need to ask me a question. When I asked him "how long did this take" and expected him to say "a few weeks" (it would have taken me - a far more experienced engineer - 2 months minimum).
His answer was "a few days".
What I'm not saying is "AGI is close" but I've seen tangible evidence (only in the last 2 months), that my 20 year software engineering career is about to change and massively for the upside. Everyone is going to be so much more productive using these tools is how I see this.
Step outside of building basic web/CRUD apps and its accuracy drops off substantially.
Also almost every library it uses is old and insecure.
I have also been developing for 20+ years.
And have heard the exact same thing about IDEs, Search Engines, Stack Overflow, Github etc.
But in my experience at least how fast I code has never been the limiting factor in my project's success. So LLMs are nice and all but isn't going to change the industry all that much.
If nothing really changes in 3-5 years, then I'd call it a flop. But the writing is on the wall that "scale = smarts", and what we have today still looks like a foundational stage for LLM's.
> If nothing really changes in 3-5 years, then I'd call it a flop
Transformers have been used for what 6 years now? Will you in 6 years say "I'll decide if they don't change the world in another 6 years?"
Working with AI-generated code to add new features feels like working with Dreamweaver-generated code, which was also unpleasant. It's not written the same way a human would write it, isn't written with ease of modification in mind, etc.
I've tried using LLMs for some libraries I'm working on, and they failed miserably. Trying to make an LLM implement a trait with a generic type in Rust is a game of luck with very poor chances.
I'm sure LLMs can massively speed up tasks like front-end JavaScript development, simple Python scripts, or writing SQL queries (which have been written a million times before).
But for anything even mildly complex, LLMs are still not suited.
Succeeding on the most common tasks (which isn't exactly what you said) is identical to "they're useful".
But it's also utterly failed to handle mundane tasks, like porting legacy code from one language and ecosystem to another, which is frankly surprising to me because I'd have assumed it would be perfectly suited for that task.
But the worst LLMs? One of my personal tests is "write Tetris as a web app", and the worst local LLM I've tried, started bad and then half way through switched to "write a toy ML project in python".
It’s a very useful tool, not magic.
front-end JS can easily also become very complex
I think a better metric is how close you are to reinventing a wheel for the thousands time. Because that is what LLMs are good at: Helping you write code which nearly the same way has already been written thousands of times.
But that is also something you find in backend code, too.
But that is also something where we as a industry kinda failed to produce good tooling. And worse if you are in the industry it's kinda hard to spot without very carefully taking a hounded (mental) steps back from what you are used to and what biases you might have.
We should not have an entire industry of 10,000,000 devs reinventing the JS/React/Spring/FastCGi wheel. Im sure those humans can contribute in much better ways to society and progress.
I'd have said the opposite. I think LLMs facilitate disposable code. It might use the same paradigms and patterns, but my bet is that most LLM written code is written specifically for the app under development. Are there LLM written libraries that are eating the world?
The code itself is not clean and reusable across implementations, but you don't even need that clean packaged library. You just have an LLM regenerate the same code for every project you need it in.
The LLM itself, combined with your prompts, is effectively the reusable code.
Now, this generates a lot of slop, so we also need better AI tools to help humans interpret the code, and better tools to autotest the code to make sure it's working.
I've definitely replaced instances where I'd reach for a utility library, instead just generating the code with AI.
I think we also have an opportunity to merge the old and the new. We can have AI that can find and integrate existing packages, or it could generate code, and after it's tested enough, help extract and package it up as a battle tested library.
E.g. let's say I'm working on a production thing and features/bugfixes accumulate and some file in the codebase starts to resemble spaghetti. The LLM can help me unfuck that way faster and get to a state of very clean code, across many files at once.
Could potentially mean just a change in time allocation/priority. As it's easier and faster to locate and potentially resolve issues later, it is less important for code to be consistent and perfectly documented.
Not fool proof and who knows how that could evolve, but just an alternative view. One of these big names in the industry said we'll have AGI when it speaks it's own language. :P.
LLMs are definitely suited for tasks of varying complexity, but like any tool, their effectiveness depends on knowing when and how to use them.
SQL is a good target language because the translation from ideas (or written description) is more or less linear, the SQL engine uses entirely different techniques to turn that query into a set of relational operators which can be rewritten for efficiency and compiled or interpreted. The LLM and the SQL engine make a good team.
1. Aasked ChatGPT to write a simple echo server in C but with this twist: use io_uring rather than the classic sendmsg/recvmsg. The code it spat out wouldn't compile, let alone work. It was wrong on many points. It was clearly pieces of who-knows-what cut and pasted together. However after having banged my head on the docs for a while I could clearly determine from which sources the code io_uring code segments were coming. The code barely made any sense and it was completely incorrect both syntactically and semantically.
2. Asked another LLM to write an AWS IAM policy according to some specifications. It hallucinated and used predicates that do not exist at all. I mean, I could have done it myself if I just could have made predicates up.
> But for anything even mildly complex, LLMs are still not suited.
Agreed, and I'm not sure we are any close to them being.
This is probably why there’s such a divide when you try to talk about software dev online. One camp believes that it boils down to duct taping as many ready made components together all in pursuit of impact and business value. Another wants to really understand all the moving parts to ensure it doesn’t fall apart.
Personally, I just use the hell out of Django for that. And since tools like that are already ridiculously productive, I don't see much upside from coding assistants. But by and large, so many of our tools are so surprisingly _bad_ at this, that I expect the LLM hype to have a lasting impact here. Even _if_ the solutions aren't actually LLMs, but just better tools, since we reconfigured how long something _should_ take.
Bad tools often falls in three categories. Too simple, too complex, or unsuitable. For the last two, you'd better switch but there's the human element of sunken costs.
> Current LLMs fail if what you're coding is not the most common of tasks. And a simple web app is about as basic as it gets.
These two complexity estimates don’t seem to line up.
I've used LLMs to generate quite a lot of Rust code. It can definitely run into issues sometimes. But it's not really about complexity determining whether it will succeed or not. It's the stability of features or lack thereof and the number of examples in the training dataset.
What I meant by complexity is not "a task that's difficult for a human to solve" but rather "a task for which the output can't be 90% copied from the training data".
Since frontend development, small scripts and SQL queries tend to be very repetitive, LLMs are useful in these environments.
As other comments in this thread suggested: If you're reinventing the wheel (but this time the wheel is yellow instead of blue), the LLM can help you get there much faster.
But if you're working with something which hasn't been done many times before, LLMs start struggling. A lot.
This doesn't mean LLMs aren't useful. (And I never suggested that.) The most common tasks are, per definition, the most common tasks. Therefore LLMs can help in many areas, and are helpful to a lot of people.
But LLMs are very specialized in that regard, and once you work on a task that doesn't fit this specialization, their usefulness drops, down to being useless.
And for me that is the best case scenario, it takes away the part we have to code / solve already solved problems again and again so we can focus more on the other parts of software engineering beyond writing code.
Personally I much prefer Chatgpt. I give it specific small problems to resolve and some context. At most 100 lines of code. If it gets more the quality goes to shit. In fact copilot feels like chatgpt that was given too much context.
All of my experiences with LLMs have been that for anything that isn't a braindead-simple for loop is just unworkable garbage that takes more effort to fix than if you just wrote it from scratch to begin with. And then you're immediately met with "You're using it wrong!", "You're using the wrong model!", "You're prompting it wrong!" and my favorite, "Well, it boosts my productivity a ton!".
I sat down with the "AI Guru" as he calls himself at work to see how he works with it and... He doesn't. He'll ask it something, write an insanely comprehensive prompt, and it spits out... Generic trash that looks the same as the output I ask of it when I provide it 2 sentences total, and it doesn't even work properly. But he still stands by it, even though I'm actively watching him just dump everything he just wrote up for the AI and start implementing things himself. I don't know what to call this phenomenon, but it's shocking to me.
Even something that should be in its wheelhouse like producing simple test cases, it often just isn't able to do it to a satisfactory level. I've tried every one of these shitty things available in the market because my employer pays for it (I would never in my life spend money on this crap), and it just never works. I feel like I'm going crazy reading all the hype, but I'm slowly starting to suspect that most of it is just covert shilling by vested persons.
In theory I'm learning from the LLM during this process (much like a real code review). In practice, it's very rare that it teaches me something, it's just more careful than I am. I don't think I'm ever going to be less slap-dash, unfortunately, so it's a useful adjunct for me.
I use it to write test systems for physical products. We used to contract the work out or just pay someone to manually do the tests. So far it has worked exceptionally well for this.
I think the core issue of the "do LLMs actually suck" is people place different (and often moving) goalposts for whether or not it sucks.
It works well as a smart documentation search where you can ask follow-up questions or when you know what the output should look like if you see it but can't type it directly from the memory.
For code assistants (aka copilot / cursor), it works if you don't care about the code at all and ok with any solution if it's barely working (I'm ok with such code for my emacs configuration).
When they say treat it like an intern, I'm so confused. An intern is there to grow and hopefully replace you as you get promoted or leave. The tasks you assign to him are purposely kept simple for him to learn the craft. The monotonous ones should be done by the computer.
[0]: https://gist.github.com/simonw/97e29b86540fcc627da4984daf5b7...
After 20 years of being held accountable for the quality of my code in production, I cannot help but feel a bit gaslit that decision-makers are so elated with these tools despite their flaws that they threaten to take away jobs.
Lots of people are terrible at going from 0 to 1 in any project. Me included. LLMs helped me a lot solving this issue. It is so much easier to iterate over something.
And that’s fine if the dev realizes what’s going on but when they attribute their own quirks to AI magic, that’s a problem.
It's almost as if the horde of former kleptocurrency bros have found a promising new seam of fool's gold to mine
I did it in little pieces and started over with fresh context each time the LLM started to get off in the weeds. I'm very happy with the result. The code is clean and well commented, the tests are comprehensive and the app looks nice and performs well.
I could have done all this manually too but it would have taken longer and I probably would have skimped out on some tests and gave up and hacked a few things in out of expedience.
Did the LLM get things wrong on occasion? Yes. Make up api methods that don't exist? Yes. Skip over obvious standard straightforward and simple solutions in favor of some rat's nest convoluted way to achieve the same goal? Yes.
But that is why I'm here. It's a different style of programming (and one that I don't enjoy nearly as much as pounding the keyboard). It's more high level thinking and code review involved and less worrying about implementation detail.
It might not work as well in domains which training data doesn't exist in. Also certainly if someone expects to come in with no knowledge and just paste code without understanding, reading and pushing back, they will have a non working mess pretty shortly. But overall these tools dramatically increase productivity in some domains is my opinion.
If you aren’t on board then it looks impressive but flawed and not even close to living up to the hype.
I've been impressed with the ability to generate "throw away" code for testing out an idea or rapidly prototyping something.
I have a good mental map of the projects I work on because I wrote them myself. When new business problems emerge, I can picture how to solve them using the different components of those applications. If I hadn't actually written the application myself, that expertise would not exist.
Your colleague may have a working application, but I seriously doubt he understands it in the way that is usually needed for maintaining it long term. I am not trying to be pessimistic, but I _really_ worry about these tools crippling an entire generation of programmers.
That sounds quite useful. Does Cursor feed your entire project code (traversing all folders and files) into the context?
My constant suspicion is that most results people are so impressed with were just never validated.
> Analogously, I imagine future software dev to consist mostly of writing specs in natural language.
https://www.commitstrip.com/en/2016/08/25/a-very-comprehensi...?
I like this take. I feel like a significant portion of building out a web app (to give an example) is boilerplate. One benefit of (e.g., younger) developers using AI to mock out web apps might be to figure out how to get past that boilerplate to something more concise and productive, which is not necessarily an easy thing to get right.
In other words, perhaps the new AI tools will facilitate an understanding of what can safely be generalized from 30 years of actual code.
I’d argue web frameworks don’t even help a lot in this regard still. They pile on more concepts to the leaky abstractions of the web. They’re written by people that love the web, and this is a problem because they’re reluctant to hide any of the details just in case you need to get to them.
Coworker argued that webdev fundamentally opposes abstraction, which I think is correct. It certainly explains the mountains of code involved.
It does seem inevitable that some large change will happen to our profession in the years to come. I find it challenging to predict exactly how things will play out.
Isn’t that the point? Degrade the user long enough that the competing user is on-par or below the competence of the tool so that you now have an indispensable product and justification of its cost and existence.
P.S. This is what I understood from a lot of AI saints in news who are too busy parroting productivity gains without citing other consequences, such as loss of understanding of the task or expertise to fact-check.
I'm not a mathematician, hell i did general maths at school. Currently I've been talking through scripting a method to mix dsd audio files natively without converting to tradional pcm. I'm about to use gpt to craft these scripts. There is no way I could have done this myself without years of learning. Now all I have to do is wait half a day so I can use my free gpt o credits to code it for me (I'm broke af so can't afford subs). The productivity gains are insane. I'd pay for this in a heartbeat if I could afford it.
In niche situations it's not helpful at all in writing code that works (or even close). It is helpful as a quick lookup for docs for libs or functions you don't use much, or for gotchas that you might otherwise search StackOverflow for answers to.
It's good for quick-and-dirty code that I need for one-off scripts, testing, and stuff like that which won't make it into production.
Yeah, if you want tic-tac-toe or snake, you can simply ask ChatGPT and it will spit out something reasonable.
But this is not much better than a search engine/framework to be honest.
Asking it to be "creative" or to tweak existing code however ...
The sad part is beginners using the boilerplate code won't get any practice building apps and will completely fail at the complex parts of an app OR try to use AI to build it and it will be terrible code.
I'm probably one of those "heavy users", though I've only been using it for a month to see how well it does. Here's my review:
Large completions (10-15 lines): It will generally spit out near-working code for any codemonkey-level framework-user frontend code, but for anything more it'll be at best amusing and a waste of time.
Small completions (complete current line): Usually nails it and saves me a few keystrokes.
The downside is that it competes for my attention/screen space against good old auto-completion, which costs me productivity every time it fucks up. Having to go back and fix identifiers in which it messed up the capitalization/had typos, where basic auto-complete wouldn't have failed is also annoying.
I'd pay about about $40 right now because at least it has some entertainment value, being technologically interesting.
If what I give it is too open ended, doesn't have enough info, etc, I'll still get a low quality output. Though I find I can steer it by asking it to ask clarifying questions. Asking it to build unit tests can help a lot too in bolstering, a few iterations getting the unit tests created and passing can really push the quality up.
2) Absolutely. Thats like one hour of an engineer salary for a whole month.
Isn't each new model bigger and heavier and thus requries more compute to train?
Until the competition outcompetes you with their new model and you have to train a new superior one, because you have no moat. Which happens what, around every month or two?
> the hardware is getting much better and prices will fall drastically once there is a bit of a saturation of the market or another company starts putting out hardware that can compete with NVIDIA
Where is the hardware that can compete with NVIDIA going to come from? And if they don't have competition, which they don't, why would they bring down prices?
Huge margins lead to a lot of competition trying to catch up, which is what makes market economies so successful.
Eventually one of you runs out of money, but your customers keep getting better models until then; and if the loser in this race releases the weights on a suitable gratis license then your businesses can both lose.
But that still leaves your customers with access to a model that's much cheaper to run than it was to create.
Google however does not sell these, you can only lease time on them via GCP.
Raising the max intelligence of the models tends to raise the intelligence of all the models via distillation.
10% more productive. What does that mean? If you mean lines of code, then it's an incredibly poor metric. They write more code, faster. Then what? What are the long-term consequences? Is it ultimately a wash, or even a detriment?
https://stackoverflow.blog/2024/03/22/is-ai-making-your-code...
Personally I am having a lot of fun, as an iOS developer, creating web games. No market in that, not really, but it's fun and I wouldn't have time to update my CSS and JS knowledge that was last up-to-date in 1998.
However, for a more typical software engineer, where every project is different, you have full lifecycle responsibility from design through coding, occasional production support, future enhancements, refactorings, updates for 3rd party library/OD updates, etc/etc, then how much of your time is actually spent pure coding (non-stop typing) ?! Probably closer to 10-25%, and certainly no-where near 100%. The potential overall time saving from a tool that saves, let's say, 10-25% of your code typing is going to be 1-5%, which is probably far less than gets wasted in meetings, chatting with your work buddies, or watching bullshit corporate training videos. IOW the savings is really just inconsequential noise.
In many companies the work load is cyclic from one major project to the next, with intense periods of development interspersed with quieter periods in-between. Your productivity here certainly isn't limited by how fast you can type.
If you pay Silicon Valley salaries this seems like a no-brainer. There are bigger time wasters elsewhere, but this is an easy win with minimal resistance or required culture change
Short term pricing inefficiency is not relevant to long term impact.
If there is a common shift to using additional runtime compute to improve quality of output, such as OpenAI's GPT-o1, then FLOPs required goes up massively (OpenAI has said it takes exponential increase in FLOPS/cost to generate linear gains in quality).
So, while costs will of course decrease, those $20-30K NVIDEA chips are going to be kept burring, and are not going to pay for themselves ...
This may end up like the shift to cloud computing that sounds good in theory (save the cost of running your own data center), but where corporate America balks when the bill comes in. It may well be that the endgame for corporate AI is to run free tools from the likes of Meta (or open source) in their own datacenter, or maybe even locally on "AI PCs".
Drop the S, I think. There’s no time dimension.
And FLOP is a generalized capability meaning you can do any operation. Hardware optimizations for ML can deliver the same 100B computations faster and cheaper by not being completely generalized. Same way ray tracing acceleration works: it does not use the same amount of compute as ray tracing in general CPU’s.
Still, even with modern accelerators it's lot of computation, and is what drives the price per token of larger models vs smaller ones.
Llama, Mixtral, Stable diffusion and Flux are a lot of fun and free to run locally, you should try them out.
Let me use CG rendering as an example. Back in the day only the big companies could afford to do photoreal 3D rendering because only they had access to the compute and even then it would take days to render a frame.
Eventually people could do these renders at home with consumer hardware but it still took forever to render.
Now we can render photoreal with path tracing at near realtime speeds.
If you could go back twenty years and show CG artists the Unreal Engine 5 and show them it’s all realtime they would lose their minds.
I see the same for A.I., now it’s only the big companies that can do it, then we will be able to do it at home but it will be slow and finally we will be able to train it at home for quick and cheap.
Ok models already run locally; that aside, as the hosted ones are kinda similar quality to interns (though varying by field), the answer is "what you'd pay an intern". Could easily be £1500/month, depending on domain.
When GPT4 was launched last year, the API cost was about $36/M blended tokens, but you can now get GPT4o tokens for about $4.4/M tokens, Gemini 1.5 Pro for $2.2/M or DeepSeek-V2 (as 21B A/236B W model that matches GPT4 on coding) for as low as $0.28/M tokens (over 100X cheaper for the same quality output over the course of about 1.5 years).
The just released Qwen2.5-Coder-7B-Instruct (Apache 2.0 licensed) also basically matches/beats GPT4 on coding benchmarks and quantized can not only can run at a decent speed on just about any consumer gaming GPU, but on most new CPUs/NPUs as well. This is about a 250X smaller model than GPT4.
There are now a huge array of open weight (and open source) models that are very capable and that can be run locally/on the edge.
For ChatGPT in its current state, probably $1K/month.
100,000 ? 500,000 ?
- ok it works, but it won't be useful.
- ok it's useful, but it won't scale.
- ok it scales, but it won't make any money.
- ok it makes money, but it's not going to last.
etc etc
After all, you could have used the exact same response in defense of web3 tech. That doesn't mean LLMs are fated to be like web3, but similarly the outcome that the current expenditure can be recouped is far from a certainty just because there are doubters.
If Copilot came for free and Azure cost a tiny bit more, nobody would even blink.