"It is difficult to get a man to understand something when his salary depends upon his not understanding it." - Upton Sinclair
It's been an absolute boon to finally build out all of the fun side projects I had always dreamed of, and after showing one off to some people I might even be able to monetize.
On the other hand I acknowledge that other people dont want to embrace LLM driven development for one reason or another, and I respect that. People got into the industry for different reasons , but code was always just a means to an ends for me.
I wonder though: Was the code ever the important part, or was it just the ability to reason through complex problems in a specific way that mattered?
Code forces you to think in a specific and valuable way. You can do that in English as well, and maybe even get more done faster... but it's a skill that will take time to master.
The people going all in on it don't realize the downsides, limitations, or understand how they come off to other people with it all.
The comparison between ELIZA and LLMs is valid you boil it down to "humans evolved for 6-7 million years, had spoken language for 500k years, but have only had something non-human that could generate convincingly novel language well enough to hold a conversation for a few decades".
There's no inherent reason it can't turn out having a non-human generate convincing enough language for conversation isn't a complete evolutionary blindspot the same way the short form feed has pretty much one-shotted society...
They had to steal the work of researches solving these open problems and then rewrite their solution. The AI equivalent of fraud.
It's not an insane assumption that the user isn't dumb and has some other reason to be asking the question other than it being a trick/stupid question (duh, if you want to wash your car you need to drive it to the car wash!). Taking it as some ultimate measure of intelligence simply doesn't make sense to me.
a) There's zero evidence of them doing so
b) Some models released before the car wash problem was discovered would consistently get it right
c) Hardcoding it is pointless. No one is seriously asking that. It's just a trick question. Hardcoding one trick question won't fix its weakness at other simple trick questions.
d) Since it went viral on the internet, the next time they updated the knowledge cutoff, the LLM would likely be aware of the trick. It will fix itself without the labs doing anything special, even assuming the new models weren't smart enough to naturally figure it out.