Now if LLMs are just effective as your experience says, they are indeed extremely useful and you absolutely should see if they can help you.
It’s only when you attempt to build a product — and it could be one person writing one Python script — that uses LLMs in an automated way with minimal human input that you really get insights into LLMs’ strengths and their limitations. You realize it could be useful, but you have to sometimes baby it a lot.
How many people get to step two? That’s a select few. Most people are stuck in the dreamy phase of trying out interactive LLMs.
This is a re-occurring issue with all new technology. Heck it happens with new software frameworks.
So the strengths and weaknesses quickly can become outdated as the strengths grow and weaknesses diminish.
When the first batch of LLMs people tried in 2023 had a lot of weaknesses. At the end of 2024, we can see increases in performance in speed and the complexity of output. People are creating frameworks on top of the LLMs that further increase their value. We went from thousands of tokens in context to millions of tokens pretty fast.
I can see myself dividing problems up into 4 groups:
1. LLMs currently solve the problem
2. It doesn't solve it now, but we are within a couple iteration of next generation models or frameworks to be able to solve it
3. LLMs are still years off from being able to solve this effectively so wait and implement it when it can.
4. LLMs will never solve this.
I think a lot of people building products are in group 2 right now.It seems that the effort required to make an LLM work robustly within a single context (spreadsheet, worddoc, email, whatever) is so gargantuan (honestly) that the returns or even the initial manpower wouldn't be there. So any new feels more or less like bloat, and if not fully useless, then at least a bit anxiety inducing in that you have no clue how much you can rely on it.
The main benefit of LLMs was already abundantly clear: literally just chat with it in day to day work when you can. Ask it questions about accounting, other domains it knows, etc. That's like up to 10-20% performance increase on tasks if you align OK.
Still, they were in search of a unicorn, and it was really tiring to be asked regularly how AI could help my workflows. They were not even spending a real budget on discovering "groundbreaking" use cases, meanwhile hounding us to shove a RAG-bot into every product they owned.
The only thing that made sense was that it was a marketing strategy to promote visibility, but they would not acknowledge that or tell us that directly (but still--it was not their business strategy to get NEW customers).
In my industry the main benefit (so far) is taking all of our human-legible unstructured data and translating it into computer-legible structured data. Loving it.
“It is valuable because you can talk to it” is the same idea that drove a tidal wave of sales for Furby in 1998
Be very careful here if you're using it for anything important! LLMs are quite good at answering questions about accounting in ways which are superficially convincing-looking, yet also complete nonsense. "But the magic robot told me it was okay" will not fly in a tax audit, say.
It might answer questions in a useful way, but you have to make sure you understand the answers and that they match accounting standards or tax rules (and one danger, at least in some places, is that they are different and you might apply the wrong one).
I don’t trust these at all for anything apart from code which I can at least read/rewrite.
It’s quite nice for unit tests I guess. And weird k8s manifests you only write now again like batch/v1 CronJob or whatever.
I’m not panicking about my job just yet..
How can you trust a tool that's right 95% of the time? In the end I wrote a script which handled edge cases explicitly. That took a little bit longer, but the output is deterministic. It took less time than manually cross referencing the output and input would have.
I tried asking GPT to write the conversion script instead, but the script it generated just didn't deal with the edge cases. After a few rounds of increasingly specific directions which didn't seem to be helping, I gave up.
I've been using copilot for development work. It has some magic moments, and it can be great for boilerplate. But then it introduces subtle bugs which are really hard to catch in review, or suggests completely incorrect function signatures and I wonder if it's adding very much at all.
The biggest problem with these tools is that they turn a fun problem solving exercise into an incredibly tedious reviewing exercise. I'd much rather do it myself and understand it fully than have to review the unreliable output of an LLM. I find it much simpler to be correct than to find flaws in other peoples work.
Am I missing something?
Of course many executives don't deal with the obscure stuff day to day (e.g. regular stuff people actually get paid to deal with) and think that LLMs can turn anyone into a superhero overnight :) The amount of times we were told that we were "putting our heads in the sand" regarding the advantages of AI was very annoying.
In reality, the decisions are made by individual executives and managers with their own interests in mind who are being asked by the people they report to what they’re doing in AI. And this goes all the way to the top where the CEOs are basically required to tell shareholders how they’ve implemented AI and how it’s helping them a bunch.
One of the nice things about being in AI right now is that your customers will advertise it and lie about how useful it has been.
(This comment is mainly facetiousness borne out of frustration, but the point stands)
"Mr. President - we must not allow an LLM gap!"
This has a very late-90's vibe and is quite entertaining to watch.
Wallstreet. These companies are largely based on increasing the share price. That's what shareholders want, that's what the board wants, and that's what the C-suite, compensated mostly in stock, wants. Nothing really matters except increasing the stock price, and right now the AI bubble is so damn superheated that if you do not announce something AI, you might be punished in the stock market as investors pull their money and put it somewhere that is willing to bullshit in order to collect some of that sweet hype money.
Because the vast majority of the stock market is literally a "greater fool" system. The point is not to make good companies that will enrich you with strong dividends forever, but rather to pump and dump. The entire point of the public market (for public companies) is that you can bullshit, pump up your stock price, and then dump those pumped up stocks onto a fool because you know the company isn't nearly as valuable as the stock says it is.
Private investors are so goddamned wealthy nowadays that it makes zero sense to engage with the public stock market, and the regulation and demands that it requires of you, unless you are trying to dump on bagholders.
Your average trader is so goddamned stupid you can literally run a company with no fundamentals and no plan and have a $13 billion market cap. Your entire business can be "sell shares repeatedly to these same losers who are blatantly illiterate, can't do basic math, and are outright gambling addicts" and retire on the free money that these people THANK YOU for taking. You can openly enrich yourself off these idiots, and they will vote for it willingly!
The point of being on the public market is to attempt to take advantage of these people and take their money.
For startups that basically implement a frontend to chatgpt or similar… well they have no chance of ever being profitable but investors might not know that.
Most LLM use at the corporate level is happening through Office 365 where Microsoft has put a Copilot button on everything including Word, Outlook and PowerPoint. Execs didn't necessarily ask for it, it protrudes very conspicuously in the UI now.
Do you want to be trying to raise money against them with nothing to tell potential investors about your company's usage of AI?
If you just wait and all of your competitors have a workforce enabled and equipped with the best tools, you are at a disadvantage. It's being the company that put off computerization or automation when everyone else went wild.
And FWIW, I cannot imagine a programmer, in 2024, that remains dismissive of LLMs or the related tooling. While they are grossly oversold by the LinkedIn crowd, and are still a far ways from so-called "prompt engineers" replacing devs, they're a massive accelerator and "second set of eyes", especially if you're doing varied, novel work.
Hi.
> While they are grossly oversold by the LinkedIn crowd,
That is true.
> and are still a far ways from so-called "prompt engineers" replacing devs,
Also true.
> they're a massive accelerator and "second set of eyes", especially if you're doing varied, novel work.
That is not even remotely my experience.
Like I can envision a programmer who would get benefits from it, but bluntly put, the code I work on day to day is far, far too interesting to be handled by copilot, simply because there aren't nearly enough stackoverflow pages about it to be scraped. Honestly if you found yourself able to automate most of your job with copilot, if anything, you have my sincerest condolences. I can't imagine how utterly bored you are in your day-to-day.
IF copilot could get to a place where it could understand, comprehend, and weigh-in on code, that would be incredibly useful. But that's not what Copilot is because that's not what transformers are. They are fancy word probability calculators. And don't get me wrong, that has uses, but it is nothing I'd be comfortable calling a second set of eyes for for anything, save for maybe writing.
While it's certainly possible that this divide is based on how 'hard' the problems people are using them on, my current theory is that some people use them like the proverbial rubber duck - in other words, a way to explore the code, and generate some stuff to work on, while thinking through the problem.
Personally, I have not yet tried it, so I'm curious which side of the discussion I'll fall ...
But so are much older programmers who have seen it all, including the obsoletion of many of their skills, and who are not so dependent on continuing to use them as they could retire anyway.
It's more the middle (programmer) age senior programmers who are less likely to see any use.
I've seen the same pattern with artists' interest in generative AI.
But it's complicated because it IS also dependent on what you're doing. So it's hard to know if something is being dismissed correctly due to domain/expertise, or prematurely due to not putting the work in and figuring out what these tools mean.
I suspect many who find them useless and decry them were sold an exaggerated utility and then were disappointed when they tried to generate libraries or even functions, then feeling deceived when there are errors or flaws, etc.
No, I suspect the large majority (and this has been backed by surveys) of people that are dismissive of them are more senior and have been working in highly specific problem domains for a long time where this is rarely/never a good "general" answer for a problem, and have spent an inordinate amount of time debugging LLM-generated or LLM-Advised code by their peers that contains nefarious and subtle errors that look correct at a glance. I personally can tell you that for what I work on, in my domain, these tools have been a net time suck and not a gain, and I pretty much only use them to ask questions about documentation, which it often gets incorrect anyway (again, in subtle ways that are probably hard for someone who isn't very senior to detect).
Hope that helps.
My take is: if the project is doing something that has been asked a thousand times on stackoverflow and has hundreds of pages in the tutorial content mills, the LLM will tell you something reasonably meaningful about it.
I'd hazard a guess that most people overenthusiastic about those tools are gluing together javascript libs.
This is not necessarily a bad thing, we even asked a LLM today at work to generate some code for a library that we didn't know how to use but seems fairly popular, and the output looked like it would make sense. (Can't tell you how it ended up because I wasn't the one implementing the thing.)
However, we also spent 2 hours in a group debugging session because we're working on a completely custom codebase that isn't documented anywhere on geeksforgeeks, stackoverflow or anywhere else public. I highly doubt that even a local LLM would be able to help, and no way this code is leaving the premises.
There are many billions of lines of high-quality, commented code online, covering just about everything. Millions of projects. All of Linux. All of Android. All of PGSQL and SQLite and MySQL and Apache and Git and OpenSSL and countless encryption libraries and countless data tools, video and audio manipulation, and...
Every single project is absolutely dominated by things that have been done many, many thousands of times. The vast bulk of your projects have zero novelty. They're mixing the same ingredients in different ways. I would think any experienced developer would realize this.
>I'd hazard a guess that most people overenthusiastic about those tools are gluing together javascript libs.
At this point it's comedy how often this "oh I understand that the noobs get value from this, but not us Advanced Programmers". It's absurdist and honestly at this point I just shake my head. My day is filled with C++, Python, Rust, Go, the absolute cutting edge of AI research, and I find these tools absolutely invaluable now. They are a massive accelerator. Zero JavaScript libs or "LOL WEB DEV" programming in my life.
How about a full equivalent of Qt that is proprietary and has absolutely nothing public in it? How is a LLM going to help with that? There is no public info anywhere.
> the absolute cutting edge of AI research
No offense but there are billions of public pages about "AI" research since it's the new gold rush. Of course LLMs have material about all your libs.
Billions? For many of the things I am working on there are zero public pages outside of research papers. I said nothing about working with libs. Again, I'm not asking an AI "here's my project now finish it", I'm working with AIs for the countless little programming challenges and needs. Things that mirror things done in many, many other projects, most having nothing to do with my domain.
As an aside, starting that with "no offense" as an attempt to make it insulting is...weird.
I feel like this discussion is taking place ten years ago. The weird reference to StackOverflow is particularly funny.
And like, even for those who write a lot of boring code, like... cool man. I don't judge people for that. We need all code written and all code is not exciting, novel, or interesting and there's nothing wrong with doing it. Someone's gotta.
I'm just saying that the further up the proverbial complexity chain I went, the less able Copilot was. And once I was quite in the weeds, it seemed utterly perplexed and frankly, not worth the time in asking.
No one takes it as a judgment, and no one is offended. It's just a truth that when people make such claims, they're often exaggerating the uniqueness or novelty of what they're doing.
You described your work in another comment, and what you described is the most bog standard programming in the field. It's always the case.
Maybe you are imagining a case where the entire codebase is generated by a single prompt?
When last I tried anyway, Copilot was, frankly, useless.
Rather, I'd be using something like the Zed editor with its AI Assistant integration and Claude Sonnet 3.5 as the model, where I first provide it context in the chat window (relevant files, pages, database schema, documents it should reference and know) and possibly discuss the problem with it briefly, and only then (with all of that as context in the prompt) do I ask it to author/edit a piece of code (via the inline assist feature, which "sees" the current chat).
But it generally is the most useful for "I know exactly what I want to write or change, but it'll take me 30 minutes to do so, while with the LLM I can do the same in 5 minutes". They're also quite good at "tell me edge-cases I might have not considered in this code" - even if 80% of the suggestions it'll list are likely irrelevant, it'll often come up with something you might've not thought about.
There's definitely problems they're worse than useless at, though.
Where more complex reasoning is warranted, OpenAI o1 series of models can be quite decent, but it's hit or miss, and with the above prompt sizes you're looking at 1-2$ per query.
I was being dismissively rhetorical. I can actually imagine them because we see them on HN constantly, with a Luddism that somehow actually becomes a rather hilarious attempt at superiority. A "well if you actually get use out of them, I guess you're just not at my superior level of Unique Projects and Unique Code". It's honestly just embarrassing at this point.
>but bluntly put, the code I work on day to day is far, far too interesting to be handled by copilot
Hilarious. It's a cliche at this point.
Let me take it further and turn this on its head: The people who usually think LLMs aren't valuable to coders generally work on the most boring, copy/paste prattle. They stick to the same tiny niche all day every day. They're basically implementing the same thing again and again. It's so rote, and they're so profoundly unchallenged, that tooling has no value.
>The people who usually think LLMs aren't valuable to coders generally work on the most boring, copy/paste prattle.
Here you are doing the same thing, aren't you?
Instead of calling people names, the biggest tell of a weak argument, why don't you explain the type of work you do and how using an LLM is faster than if you coded it yourself and and/or also faster than any current way of doing the same thing.
I'm assuming you are a senior+ level coder.
>Instead of calling people names
Who called anyone a name? Luddism? Yes, many HN participants are reacting to AI in a completely common rejection of change / challenge, and it recurs constantly.
>how using an LLM is faster than if you coded it yourself
I am coding it myself. Similar to the other guy who talks about putting an LLM in "charge" of his precious, super-novel code, you're setting up a strawman where using an LLM implies some particular scenario that you envision. In reality I spend my day asking questions, getting broad strokes, getting code commented, asking for API or resources, etc.
Sorry, perhaps I misinterpreted it.
>In reality I spend my day asking questions, getting broad strokes, getting code commented, asking for API or resources, etc.
Can you give me some concrete examples? I'd like to use it, but I'm currently of the mind:
1. If it's boring code, I can write it faster than asking LLM to do it and fixing its issues.
2. If it's not boring code, like say a rules engine or something, I'm not sure the LLM will give me a good result based on the domain.
I mainly stick to back end work, automation, building WebAPI's and DSS engines for the medical field.Maybe I'm under and over thinking it at the same time. FWIW, I typically stick to a single main language, but where I usually work, the companies dictate a GP language for all our stuff: C# in my example. I do a small amount of Python for LLM training, but I'm just starting out with Python. I can see it being useful saying, "convert this C# to Python," but honestly, I'd rather just learn the Python.
You should read up on what Luddism and Luddists were actually about. They didn't think the machines were evil or satanic, which is the common cultural read. They assumed (correctly) that the managerial class would take full advantage of increased productivity of lower-quality goods to flood the market with cheap shit that would put competitors out of business, and let them fire 4/5 of their workforces while doing so. And considering the state of the textile industry today, I think that was a pretty solid set of projections.
Luddites didn't oppose automation on the basis that machines are scary. They were the people who worked the machines that already existed at the time, after all. They opposed them on the basis that the greedy bastards who owned everything would be the only ones actually benefiting from automation, everyone else would get one kind of shaft or another, which again: is exactly what happened.
This, actually is incredibly analogous to my opinions about LLM. It's an interesting tech that has applications but is already being situated to be the sole domain of massive hyperscalers and subject to every ounce of enshittification that follows every tech that goes that way, while putting creatives, and yes some coders, out of a job.
So yes, it was name calling, but also I don't object to the association. In this case, I'm a Luddite. I am suspicious of the motivations and the beneficiaries of automation being forced into my industry and I'm not going to be quiet about it.
What also happened is that everyone can buy clothes incredibly cheaply. Which seems like a widespread benefit.
I think it's just about all industries these days.
Yes, so many quote meanings have been malformed over the years, such as "a rolling stone gathers no moss," is considered good, while originally bad. "Blood is thicker than water," "Money is the root of all evil," etc.
The Luddites were right.
They were primarily opposed to automation because it devalued the work they did and the skills they held. That is the core essence of Luddism. They thought if they destroyed the machines, automation could be stopped. There were some post-facto justifications like product quality, but if that was true they'd have no problem out-competing the machines.
Yes, it is Luddism that drives a lot of the AI sentiment seen on HN, and it is not only utterly futile and basically people convincing themselves and each other while the world moves on. There is no "name calling", and that particular blend of pearl clutching is absurd.
I don't think I'm calling myself a super special snowflake here. These models are just ... bad at sales forecasting.
LLMs aren't entirely useless for me. I'll use ChatGPT to generate code to make plots. That's helpful.
If someone says "Well my software dev project is too unique and novel and they are therefore of no value to me, but I understand it works for those simple folk with their simple needs", there is an overwhelming probability they are...misinformed.
What do you mean by that? In my experience they mostly help when you are doing mainstream, well-trodden work for which a company’s/project’s/domain’s inside knowledge isn’t needed. Maybe you mean the latter by “novel”?
Varied and novel to the programmer. If you're doing the same thing day in and day out it probably isn't much use. If you're like many programmers and you're jumping between libraries and languages and platforms and APIs and domains and spheres, it's immensely helpful.
>doing mainstream, well-trodden work for which a company’s/project’s/domain’s inside knowledge isn’t needed
A core disconnect in this discussion is that many people seem to be arguing from the position of "I tried to generate whole solutions with these tools and it failed, so it's useless".
Everything is well-trodden. One of the commentators in here, who declares themselves a unique snowflake where these tools are useless, does "sockets with encryption and messaging", which is some of the most well-trodden ground in this domain. Everything is glue. Everything is a lot of relatively simple things strung together. And for all of those, portions are helped and accelerated with these tools.
Everyone is working in unique domains and spaces with weird rules and restrictions and needs. The whole point of the special snowflake comment is specifically that people aren't remotely unique in being a special snowflake.
And the Luddism argument isn't "name-calling", it's an observation of an absolute truth on here.
Everything is extremely well-trodden at the code level. You're taking some inputs and generating some outputs. You're calling some functions. You're manipulating some strings or lists or sets. You're sorting some things or filtering things. You're building a client to a given API. You're doing some messaging and some encryption.
I guarantee that your code isn't remotely as unique or novel as you think it is. Of course on here everyone is 240lbs, 6'2" and benches 350, and their magically novel, super unique code is just too secret. So everyone's special.
If you had to explain the domain and context to these tools, you might be using them wrong.
Can you imagine even in principle a piece of evidence that would convince you otherwise?
And yes, in the end all code distills down to shit that is very similar to the countless billions of lines of code on the internet. Everything -- every single project that anyone on this site is working on -- is a bunch of glued together shit where 90% (more like 99% Probably 99.9%) of it is in common with code seen in countless other projects. Utterly regardless of domain or business or project specific uniqueness. Someone would have to be profoundly incompetent to not realize this.
Again, I seem to arguing with people who think that using a tool means feeding it their whole projects and having it rewrite it in Rust or something. In reality a programmer could yield immense value from such tools having a) given it zero lines of their code, b) used zero lines of the code it generated. This isn't a difficult concept, but I'm going to continue getting insane responses by all of the unique people who work on the amazingly unique situations where instead of sorting from A-Z, they sort from Z-A!
> In reality a programmer could yield immense value from such tools having a) given it zero lines of their code, b) used zero lines of the code it generated.
We simply have different experiences of life! I think with the advent of Claude 3.5 Sonnet the LLMs have just about edged out ahead in terms of time saved vs time wasted, for me, but before Sonnet I'm fairly confident they were moderately net negative for me.
Can you give some concrete examples of where they've helped you this dramatically? With links to chat logs? I still don't understand how people are finding them so useful, and I keep asking people and they keep not providing chat logs so I can see what they're doing.
If someone can't use these tools for security or propriety reasons, that's an obvious hard restriction. But saying "oh my code is just too unique, my skills too advanced" is self-deluding nonsense.
>I think with the advent of Claude 3.5 Sonnet the LLMs
We are talking about now. The present. I'm stating the value of these tools now. Their state in the past is irrelevant.
>Can you give some concrete examples
Sorry, NDAs, secret code, secret projects, et al
> Their state in the past is irrelevant.
Apparently also their state in the present? I said "just about edged out ahead", whereas you say "immense value".
I find a number of models fantastically useful in the present. I'm not sure why your one day evaluation of them is relevant to how I utilize them.
It was inevitable that someone was going to do the "but I insist upon writing in some fringe language" bit (when the "my project and domain is too unique" bit faltered), but many models actually have excellent F# abilities. Eh.
> many models actually have excellent F# abilities
Name three? As I said, I badly want this stuff to work!
And of course Github Copilot autocomplete works well just as it does for most other languages.
Even the best programmers have very narrow skills relative to the whole field.
Yeah, I've used 20+ languages and hundreds of technologies and a variety of different types of product and can pick things up quickly. But it's still a drop in the bucket of technologies you can use and types of problem you can solve.
Programmer skills are deep but narrow. LLM skills are shallow but wide. It's an excellent complement for any programmer working outside their deep+narrow expertise.
I've confirmed this asked chatgpt: 9.11 > 9.9 true or false?
True because .11 is greater than .9
Write a speech announcing a momentous scientific discovery - the solution to the long standing question of (48294-1444)*0.3258
LLMs should never do math. They shouldn't count letters or sort lists or play chess or checkers. Basically all of the easy gotcha stuff that people use to point out errors are things that they shouldn't do.
And you pointed out something they do now which is creating and run a python script. That really is a pretty solid, sustainable heuristic and is actually a pretty great approach. They need to apply that on their backend too so it works across all modes, but the solution was never just an LLM.
Similarly, if you ask an LLM a chess question -- e.g. the best move -- I'd expect it to consult a chess engine like Stockfish.
But these aren't "gotcha questions", these are just some of the basic interactions that people will want to have with intelligent assistants. Literally just two days ago I was doing some things with the compound interest formula - I asked Claude to solve for a particular variable of the formula, then plug in some numbers to calculate the results (it was able to do it). Could I have used Mathematica or something like that? Yes of course. But supposedly the whole purpose of a general purpose AI is that I can use it to do just about anything that I need to do. Likewise there have been multiple occasions where I've needed ChatGPT or Claude to work with tables or lists of data where I needed the results to be sorted.
So early on the value was completely in language. But you're absolutely correct that for these tools to really be useful they need to be better than that, and slowly we're getting there. If you're asking a math question as a component of your question, firstly delegate that to an appropriate math engine while performing a series of CoT steps. And so forth.
But it's rapidly making progress. CoT models coupled with actual domain-specific logic engines (math, chemistry, physics, chess, and so on) will be when the promise is actually met by the reality.
So Indeed 9.11 is chronologically higher than 9.8 and chronology is an extremely common use case.
However a grade F will be given by many.
ChatGPT 4o gets both of these cases correct for me.
> [11,9,1,3].sort()
[ 1, 11, 3, 9 ] const list = [3, 11]
list.sort()
console.log(list[0] < list[1])
logs `false`. JavaScript doesn't "think" anything but its native sort function doesn't do what many people expect it would when called on a list of pure numbers.If you found this behavior intuitive and unsurprising the first time you saw it, then your brain works differently than mine.
If this happens to be new to you (congrats on being one of today's 10,000), the reason is that JavaScript sorts by converting all elements to UTF-16 strings, except for `undefined` for some reason, which always sorts last. The MDN docs have a very clear explanation. I have been unable to find a historical explanation of why this choice was made, but I presume the initial JS v1 author either had a good reason or just really didn't expect that their language would outlive the job for which it was written.
If this is a bug for you, you can provide a comparison function explicitly and the typical is something like:
list.sort((a,b)=>a-b)
Somewhat puzzlingly, this will "work" even on lists with mixed number and string like: const list = [3, "11", 1, "2", 9, 23]
list.sort((a,b)=>a-b)
console.log(list)
[ 1, '2', 3, 9, '11', 23 ]
Because JavaScript goes out of its way to make comparisons like 3 < "11" or "3" < 11 work in the numeric domain. JS only uses string comparison when both sides are strings.It may have been thought that js would be more likely to be dealing with strings arrays.
I do think that JavaScript's choice to sort numbers lexicographical instead of arithmetically is a bit silly, but of course no language is free from warts. Of course they cannot change it now, because that would break the web. `JSON.stringify` is also pretty silly while we're at it, but Python's `json.dumps` is no better.