391 karma · joined August 6, 2024
For example, there’s a ton of room for developing all kinds of low latency, highly reliable, embedded classifiers in a number of domains.
It’s not as gee-whiz/sci-fi as an LLM demo, but I think potentially much bigger impact over time.
The condition of “some people are bad at thing” does not equal “computer better at thing than people”, but I see this argument all the time in LLM/AI discourse.
I then looked it up and they had each copy/pasted the same Stack overflow answer.
Furthermore, the answer was extremely wrong, the language I used was superficially similar to the source material, but the programming concepts were entirely different.
What this tells me is there is clearly no “reasoning” happening whatsoever with either model, despite marketing claiming as such.
In a competitive market where people make very long term engineering decisions based on stability and reliability you can’t fuck up this badly and survive.
Currently at least 50% of online ads are outright illegal in most parts of the world.
Nobody is morally required to have their legal rights violated to get information. Period.
For one off scripts or sketching a concept quickly it’s good enough, and for language reference it’s generally useful.
However, one thing I’ve noticed with Claude in particular is it tends to overweight the top answers in stack overflow.
The problem there is top answers are rarely the best answer - rather they tend to be overly verbose and long, whereas the best answer is usually the second one that just tells you what function to call.
On multiple occasions I’ve had Claude answer a simple prompt with horribly verbose and complicated code.
Then I’ll say “what about this single call?” (Eg the type of SO answer than gets the second most votes), and it’s says “You’re right! That’s a much better answer”.
Likewise I suggest to anybody to take the domain you are most knowledgeable in and pepper your LLM of choice with lots of questions and see how much it knows.
You’ll get a feel for how much you can trust it in other domains - which is “not much”.