> the point where LLMs are surpassing humans in any task limited in scope enough to be a “benchmark”
so much as we're over-fitting these bench marks (and in many cases fishing for a particular way of measuring the results that looks more impressive).
While it's great that the LLM community has so many benchmarks and cares about attempting to measure performance, these benchmarks are becoming an increasingly poor signal.
> This is a nerve-wracking time to be a knowledge worker for sure.
It might because I'm in this space, but I personally feel like this is the best time to working in tech. LLMs still are awful at things requiring true expertise while increasingly replacing the need for mediocre programmers and dilettantes. I'm increasingly seeing the quality of the technical people I'm working with going up. After years of being stuck in rooms with leetcode grinding TC chasers, it's very refreshing.
Maybe I'm confused but 10k attempts on the same problem set would make anyone an expert in that topic? It's also weird that zero shot performance is so bad, but over a lot of attempts it seems to get correct answers? Or is it learning from previous attempts? No info given.
A computer will never go on strike, demand better working conditions, unionize, secretly be in cahoots with your competitor or foreign adversary, play office politics, scroll through Tiktok instead of doing its job, or cause an embarrassment to your company by posting a politically incorrect meme on its personal social media account.
I get the AI skepticism because so much tech hype of recent years turned out to be hot air (if you're generous, obvious fraud if you're not). But AI tools available toady, once you get the hang of using them, are pretty damn amazing already. Many jobs can be fully automated with AI tools that exist today. No further breakthroughs required. And although I still don't believe software engineers will find themselves out of work anytime soon, I can no longer completely rule it out either.
I am interpreting this to mean that the model tried 10K approaches to solve the problem, and finally selected the one that did the trick. Am I wrong?
That's the thing, did the operator select the correct result or did the model check it's own attempts? No info given whatsoever in the article.
I have spent significant time with GPT-4o, and I disagree. LLMs are as useful as a random forum dweller who recognises your question as something they read somewhere at some point but are too lazy to check so they just say the first thing which comes to mind.
Here’s a recent example I shared before: I asked GPT-4o which Monty Python members have been knighted (not a trick question, I wanted to know). It answered Michael Palin and Terry Gilliam, and that they had been knighted for X, Y, and Z (I don’t recall the exact reasons). Then I verified the answer on the BBC, Wikipedia, and a few others, and determined only Michael Palin has been knighted, and those weren’t even the reasons.
Just for kicks, I then said I didn’t think Michael Palin had been knighted. It promptly apologised, told me I was right, and that only Terry Gilliam had been knighted. Worse than useless.
Coding-wise, it’s been hit or miss with way more misses. It can be half-right if you ask it uninteresting boilerplate crap everyone has done hundreds of times, but for anything even remotely interesting it falls flatter than a pancake under a steam roller.
> Only one Monty Python member, Michael Palin, has been knighted. He was honored in 2019 for his contributions to travel, culture, and geography. His extensive work as a travel documentarian, including notable series on the BBC, earned him recognition beyond his comedic career with Monty Python (NERDBOT) (Wikipedia).
> Other members, such as John Cleese, declined honors, including a CBE (Commander of the British Empire) in 1996 and a peerage later on (8days).
Maybe you just asked the question wrong. My prompt was "which monty python actors have been knighted. look it up and give the reasons why. be brief".
The point is that you never know what you can trust or not. Unless you’re intimately familiar with Monty Python history, you only know you got the correct answer in one shot because I already told you what the right answer is.
Oh, and by the way, I just asked GPT-4o the same question, with your phrasing, copied verbatim and it said two Pythons were knighted: Michael Palin (with the correct reasons this time) and John Cleese.
¹ And I’ve had enough discussions on HN where someone insists on the correct way to prompt, then they do it and get wrong answers. Which they don’t realise until they shared it and disproven their own argument.
If you pay careful attention to prompt phrasing you will get a lot more mileage out of these models. That's the bottom line. If you believe that you shouldn't have to learn how to use a tool well then you can be satisfied with your righteous attitude but you won't get anywhere.
Similarly language models are probabilistic and yet they get the easiest questions right 100% of the time with little variability and the hardest prompts will return gibberish. The point of good prompting is to get useful responses to questions at the boundary of what the language model is capable of.
(You can also configure a language model to generate the same output for every prompt without any random noise. Image models for instance generate exactly the same image pixel for pixel when given the same seed.)
In comparison, with text a single word can change the entire meaning of a sentence, paragraph, or idea. The same word in different parts of a text can make all the difference between clarity and ambiguity.
It makes no difference how good your prompting is, some things are simply unknowable by an LLM. I repeatedly asked GPT-4o how many Magic: The Gathering cards based on Monty Python exist. It said there are none (wrong) because they didn’t exist yet at the cut off date of its training. No amount of prompting changes that, unless you steer it by giving it the answer (at which point there would have been no point in asking).
Furthermore, there’s no seed that guarantees truth in all answers or the best images in all cases. Seeds matter for reproducibility, they are unrelated to accuracy.
You think the probabilistic nature of language models is a fundamental problem that puts a ceiling on how smart they can become, but you're wrong.
No. Language can be fuzzy, yes, but not at all in the same way. I have just explained that.
> LLMs can create factually correct responses in dozens of languages using endless variations in phrasing.
So which is it? Is it about good prompting, or can you have endless variations? You can’t have of both ways.
> You fixate on the kind of questions that current language models struggle with
So you’re saying LLMs struggle with simple factual and verifiable questions? Because that’s all the example questions were. If they can’t handle that (and they do it poorly, I agree), what’s the point?
By the way, that’s a single example. I have many more and you can find plenty of others online. Do you also think the Gemini ridiculous answers like putting glue on pizza are about bad promoting?
> You think the probabilistic nature of language models is a fundamental problem that puts a ceiling on how smart they can become, but you're wrong.
One of your mistakes is thinking you know what I think. You’re engaging with a preconceived notion you formed in your head instead of the argument.
And LLMs aren’t smart, because they don’t think. They are an impressive trick for sure, but that does not imply cleverness on their part.
It seems many of the popular tools want to make writing software harder than in the 2010s, though. Perhaps their stewards believe that if they keep making things more and more unnecessarily complicated, LLMs won't be able to keep up?
The big decline is driven by a few big factors. Two of which are 1- the overhiring that happened in 2021. This was followed by the increase of interest rates which dramatically constrained the money supply. Investors stopped preferring growth over profits. This shift in investor preferences is reflected in engineering orgs tightening their budgets as they are no longer rewarded for unbridled growth.
I'm not especially nerve-wracked about being a knowledge worker, because my day-to-day doesn't consist of being handed a detailed specification of exactly what is required, and then me 'computing' it. Although this does sound a lot like what a product manager does!
legacy code is a problem regardless of who wrote it. Humans have been writing suboptimal, hard-to-maintain code for decades. At least with LLMs, we have the opportunity to design and implement better coding standards and review processes from the start.
let's be real, most of the code written by humans is not exactly a paragon of elegance and maintainability either. I've seen my fair share of 'accidentally quadratic algorithms' and 'subtly wrong code that looks right' written by humans. At least with LLMs, we can identify and address these issues more systematically.
As for 'un-idiomatic use of programming language features', isn't that just a matter of training the LLM on a more diverse set of coding styles and idioms? It's not like humans have a monopoly on good coding practices.
So, instead of throwing up our hands, why not try to address these issues head-on and see if we can create a better future for software development?
We will yearn for the pre-GPT years at some point, like we yearn for the internet of the late 90s/early 2000s. Not for a while though. We're going through the early phases of GPT today, so it hasn't been taken over by the traditional power players yet.
Which is how we ended up here, which I guess is tolerable, where a webpage with a bit of styling and a table uses up 200MB of RAM.
If LLMs can generate high-quality code with minimal human input, what does that mean for the wages and job security of programmers? Will companies start to rely more heavily on AI-generated code, and less on human developers? It's not hard to imagine a future where LLMs are used to drive down programming costs, and human developers are relegated to maintenance and debugging work.
I'm not saying that's necessarily a bad thing, but it's definitely something that needs to be considered. As someone who's enthusiastic about the potential of code gen this O1 reasoning capability is going to make big changes.
do you think you'll be willing to take a pay cut when your employer realizes they can get similar results from a machine in a few seconds?
I had a funny one a while back (granted this was probably ChatGPT 3.5) where I was trying to figure out what payload would get AWS CloudFormation to fix an authentication problem between 2 services and ChatGPT confidently proposed adding some OAuth querystring parameters to the AWS API endpoint.
It tends to fall apart on bigger asks with larger context. Breaking your task into discrete subtasks works well.
It's like infinitely tailored blog posts, for me at least.
With that said i'm not one of those "It's just a parrot!" people. It is, definitely just a parrot atm.. however i'm not convinced we're not parrots as well. Notably i'm not convinced that that complexity won't be sufficient to walk talk and act like intelligence. I'm not convinced that intelligence is different than complexity. I'm not an expert though, so this is just some dudes stupid opinion.
I suspect if LLMs can prove to have duck-intelligence (ie duck typing but for intelligence) then it'll only be achieved in volumes much larger than we imagine. We'll continue to refine and reduce how much volume is necessary, but nevertheless i expect complexity to be the real barrier.
I recently wrote a complex web frontend for a tool I’ve been building with Cursor/Claude and I wrote maybe 10% of the code; the rest with broad instructions. Had I done it all myself (or even with GitHub Copilot only) it would have taken 5 times longer. You can say this isn’t the most complex task on the planet, but it’s real work, and it matters a lot! So for increasingly many, regardless of your personal experience, these things have gone far beyond “useful toy”.
It's time to learn some real math and science, the era of regurgitating UI templates is over.
I’m sure this tech continues to have many limitations, but every piece of trajectory evidence we have points in the same direction. I just think you should be prepared for the ratio of “real” work vs. LLM-capable work to become increasingly small.
> It's time to learn some real math and science, the era of regurgitating UI templates is over.
You do realize that software development was one of the last social elevators, right?
What you're asking for won't happen, let alone the fact that "real math and science" pay a pittance, there's a reason the pauper mathematician was a common meme.
If you're not seeing value in them, maybe it's because you're not looking at the right problems. Or maybe you're just not using them correctly. Either way, dismissing an entire field of research because it doesn't fit your narrow use case is pretty short-sighted.
FWIW, I've been using LLMs to generate production code and it's saved me weeks if not months. YMMV, I guess
Can you explain what this statement means? It sounds like you're saying LLMs are now smart enough to be able to jump through arbitrary hoops but are not able to do so when taken outside of that comfort zone. If my reading is correct then it sounds like skepticism is still warranted? I'm not trying to be an asshole here, it's just that my #1 problem with anything AI is being able to separate fact from hype.
We are seeing steady improvement on long-run tasks (SWE-Bench being one example) and much more improvement on shorter, more well-defined tasks. The latter capabilities aren’t “hype” or just for show, there really is productive work like that to be done in the world! It’s just not everything, yet.
If you have to keep checking the result of an LLM, you do not trust it enough to give you the correct answer.
Thus, having to 'prompt' hundreds of times for the answer you believe is correct over something that claims to be smart - which is why it can confidently convince others that its answer is correct (even when it can be totally erroneous).
I bet if Google DeepMind announced the exact same product, you would equally be as skeptical with its cherry-picked results.
This seems like a bold statement considering we have so few benchmarks, and so many of them are poorly put together.