Programmers just moved up a level , not dumber, we are now managers of a team of 'agent' programmers. The deliverable is now a functionality instead of a specific block of code
Programmers just moved up a level , not dumber, we are now managers of a team of 'agent' programmers. The deliverable is now a functionality instead of a specific block of code
But will Claude give you an authentic rationale and a traceable, verifiable "line of reasoning" for those things? Or will it just construct the next plausible Markov chain built on whatever Reddit thread it ingested at random?
You can ask Claude or any LLM for citations, and it will RAG them out ex post facto. Those actually aren't citations, they're just web searches for related articles, and they don't necessarily support the assertions that you're asking to cite.
I am sure that Claude and the others can produce intermediate logs of their inference and "reasoning" process while they are processing stuff, but can they really go back within the context window and construct an authentic apologia for a specific thing when you ask for it?
Specifically, humans are known to decide subconsciously, then invent some "reasoning" out of thin air to justify it.
This matches my experience with decision-making in software projects.
If you ask it 3 times to generate 3 verifiable reports that confirm its claims, will it answer with the same process and same answers each time?
See, when you ask a human to justify a result or a decision, they can often do this very meticulously. If a judge writes a decision from the bench, or a firefighter describes how his battalion knocked down an apartment fire, or a systems admin describes how he configured a NAS, they will all be relying on their training, and precedent, and specifications, and things like that, and they can give you reproducible results and solid justifications for the way they did things. When mistakes are made, and money or life is lost, they can be accountable and you can modify that process to set a precedent for the future.
But Claude? How in the world will it produce the same results twice? It is non-deterministic. That is the fundamental issue of LLMs and genAI today. They are all non-deterministic, and SWE treat them as if they are somehow reliable, or produce reproducible results, or that they can follow a procedure or a specification, outlined in their prompts and context, and produce results.
No, they only produce results by accident and happenstance, and they only justify them ex post facto by making things up. There is no humanity or deterministic activity in an LLM. You'll never verify "why" they chose that string of tokens, because they could've easily chosen a very different stream of tokens. In fact, now with watermarking, the most deterministic thing will be hitting that watermark standard at all costs!
The fact that some course of action was previously mentioned in a reasoning trace, or any other context, makes it more likely to be performed. It has nothing to do with the reason that it was mentioned in the reasoning trace.
And it's not "no relation", it was brought up as an attempt to fix/subset the original claim.
No, it will invent retroactively a plausible sounding reason why someone might have done it that way. These are very different things.
Dementia patients also do this.
If I need to understand a specific line of code it means I did something wrong in planning or in requirements for testing.
What I’m describing is thorough documentation of requirements (acceptance criteria, if you like), and then encouraging Claude to be agile in execution.
As long as the outcome is well-defined, it is expected and normal to iterate on implementation.