Honestly this feels like a true statement to me. It's obviously a new technology, but so much of the "non-deterministic === unusable" HN sentiment seems to ignore the last two years where LLMs have become 10x as reliable as the initial models.
Honestly this feels like a true statement to me. It's obviously a new technology, but so much of the "non-deterministic === unusable" HN sentiment seems to ignore the last two years where LLMs have become 10x as reliable as the initial models.
But NNs are fundamentally continuous, I don't think it even makes sense to "count" bugs. You can have a list of prompts to which the model gives unwanted output, but it's a completely different ball game compared to regular software.
Of course LLMs aren't people, but an AGI might behave like a person.
LLMs don't learn from a project. At best, you learn how to better use the LLM.
They do have other benefits, of course, i.e. once you have trained one generation of Claude, you have as many instances as you need, something that isn't true with human beings. Whether that makes up for the lack of quality is an open question, which presumably depends on the projects.
How long do you think that will remain true? I've bootstrapped some workflows with Claude Code where it writes a markdown file at the end of each session for its own reference in later sessions. It worked pretty well. I assume other people are developing similar memory systems that will be more useful and robust than anything I could hack together.
There’s an interplay between two different ideas of reliability here.
LLMs can only provide output which is somehow within training boundaries.
We can get better at expanding the area within these boundaries.
It will still not be reliable like code is.
Many of the inventors of LLMs have moved on to (what they believe are) better models that would handle such learnings much better. I guess we'll see in 10-20 years if they have succeeded.