I have the feeling that LLMs are effectively running on dream logic, and everything we've done to make them reason properly is insufficient to bring them up to human level.
You could build layers and layers of LLMs watching the output of each others thoughts and offering different commentary as they go, folding all the thoughts back together at the end. Currently, a group of agents acts more like a discussion than something somewhat omnipotent or omnitemporal.
What they lack is multi turn long walk goal functions — which is being solved to some degree by agents.
And if you don't prove your code, do you not design at all then? Do you never draw state diagrams?
Every design is an informal proof of the solution. Rarely I write formal proofs. Most of the time I write down enough for myself to be convinced that the desing solves the problem.
As a exam grader, you can easily tell when a student has the mindset of "solving a problem" but made a mistake, and when they had the mindset of "looks like it solves the problem" and just wrote some stuff.
> Most of the time I write down enough for myself to be convinced that the desing solves the problem.
Again, why do you assume we aren't doing the same thing with LLMs?
1. Spec given
2. Ask LLM to write a bunch of design documents based off of spec
3. Ask LLM to identify edge cases
4. Ask LLM to device edge cases in to a test plan involving N tests
5. Ask LLM to write tests
6. Ask LLM to write commented code
7. Ask LLM to run tests on code, and determine on failing tests if test or code is wrong, go back to the appropriate step to fix test and/or code.
Whenever I hear someone here on HN imply that the only way to code with an AI is via vibe coding I just die a bit more inside.
It was a response to you saying: "Im not going around trying to prove my code correct when I write it manually."
How did you manage to forget what you wrote previously?
Also, in this post you are now suddenly taking the exact opposite position, contradicting your previous point.
And I have not made any statements about how you use LLMs, only about how the LLMs produce code. All statements about how you use LLMs have been made by you, not me. I haven't discussed it since it is not related to the arguments, which are: 1) whether LLMs are goal-oriented and 2) whether humans and LLMs both merely maximize plausibility when writing/generating code.
Both claims that you made. Note, however, that if you are correct in your own points, then you should indeed be able to "just dump out code without any process in between". So if anyone is claiming this, it's you.
This definitely matches my experience of talking to AI agents and chatbots. They can be extremely knowledgeable on arcane matters yet need to have obvious (to humans) assumptions pointed out to them, since they only have book smarts and not street smarts.
Sure, he could have submitted a ill-considered 3800 line PR five years ago, but it would have taken him at least a week and there probably would have been opportunities to submit smaller chunks along the way or discuss the approach.
I think we’re going to see a lot of the systems we depend on fail a lot more often. You’d often see an ATM or flight staus screen have a BSOD - I think we’re going to see that kind of thing everywhere soon.
Otherwise I'll just say I'm right and you're wrong, after all, that's what you're saying.
Now, could an equivalent process be modelled at some point? Probably. It'd be a conscious decision to do so on our part, and given fears over the AI Alignment quandary, it seems a rather fraught direction to carelessly proceed.
One of the papers call this "programming language semantics", but it is using a 2D grid navigation DSL. The semantics of that language are nothing like actual programming language semantics.
These are not the same as the concept being discussed here, a human "world model" of a computer system, through which to interpret the semantics of a program.
That's not really a world model.
I see current expectations that technical debt doesn't matter. The current tools embrace superficial understand. These tools to paper over the debt. There is no need for deeper understanding of the problem or solution. The tools take care of it behind the scenes.
But people want an AI that is objective and right. HN is where people who know the distinction hang out, but it’s not what the layperson things they are getting when they use this miraculous super hyped tool that everybody is raving about?
YMMV.
-Michael Crichton
that doesn’t mean the future won’t herald a way of using what a transformer is good at - interfacing with humans - to translate to and interact with something that can be a lot more sound and objective.
And even if they were solved, how would that even work? The world is not sound and objective.
But right now there are lots of domains where current lauded success is in treating something objective - like code - as tokens for an llm.
We could instead explore using transformers to translate human languages to a symbology that can be reasoned about and applied eg to code.
It’s the talk of conferences. But whether it works better than we have today, or whether it aligns with the incentives or the big players, is another matter
It seems to me that it's all a matter of company culture, as it has always been, not AI. Those that tolerate bad code will continue to tolerate it, at their peril.