LLMs interpret (so, "execute" in a way) natural language in the sense they have internal logic that assigns to the sequence of tokens in context a next token. If we delineate the input and output into a series of logical statements, we can think of it as a program that builds a logical statement from a list of input statements. So it encodes derivation in some logical system.
However, the internal logical system is informal in the sense that the above rules are not guaranteed to be sound on the fragment of classical logic encoded in the natural language. It is a close approximation, though, so it often works.
To add, half of my problem with natural language would be resolved by agreeing on exact definitions, which is kinda what LLMs do internally. However, they don't surface this formalization very well(even with open weights it's difficult), which makes it pretty unusable.
You could say the same thing about human programmers, but I've never heard anyone say they think that programmers "execute" Jira tickets.
> they have internal logic that assigns to the sequence of tokens in context a next token.
I don't think this means what you think it means, because it has almost no information content relevant to what we're discussing. The probabilities that are most relevant at the level we're discussing are satisfying a reward function from post-training, which approximates to:
"What is the likelihood the solution the agent is pursuing will be marked correct by the automated grader based on the full prompt and other context provided?"
It still has to predict the next token but that isn't based on a likelihood of that token appearing in a corpus of internet text consumed in pretraining. That was eons ago. Every predicted token is shaped by the probabilities of the predicted solution, which must already be very specific and shaped completely by the request and associated context that is built during investigation of the same.
Yes I could. We have "executives", for starters. And first "computers" were actual humans.
"That was eons ago."
Yes, technically I should call them LRMs (large reasoning models) not LLMs. But that doesn't seem relevant here, to my point they encode some logic (which we want to be close to classical logic, i.e. behavior of words like "true", "and", "not" and so on matches).
And that’s why the parent comment is saying that we need people that understand the thing that is being maintained, i.e. the program code, which is the source of truth about what is being maintained.
We do for a lot of software today, but not all of it. I think in a year we'll need them for less, but I'm not sure how much less.
This claim isn't supported or justified by talking about natural language versus formal language. For the last 50 years or so the people shaping a lot of software in the most important ways are often using only natural language.
The product manager doesn't understand the code today. They write PRDs and Jira tickets and comments in Slack, and software comes out. They have people test the software, report bugs, more software comes out. Eventually they decide its close enough to their vision to ship, without ever understanding or looking at any of the code. They can do that with human programmers or agent programmers. The former holds up better in larger systems, but I don't see any evidence for the proposition that this is due to the limits of natural language.