To understand an English description of code you already have to have a deeper understanding of what the code is doing. For code itself you can reference the syntax to understand what's going on.
The prompt in this case is using very technical language that a beginner will have no idea about. But if you gave them the code they could at least struggle along and figure it out by looking things up.
The first two paragraphs alone are absolutely chock with terms that would not be easily explained to a layperson:
"The current system is an online whiteboard system. Tech stack: typescript, react, redux, konvajs and react-konva. And vitest, react testing library for model, view model and related hooks, cypress component tests for view.
All codes should be written in the tech stack mentioned above. Requirements should be implemented as react components in the MVVM architecture pattern."
What is every library in that list? What is a model? What is a view model? What is a hook, component test, view, MVVM, etc?
If a layperson could understand explanations for all these things then they would not be a layperson.
That is so much more productive because it immediately removes or highlights all the ambiguity which you can sugar coat with English but not in a programming language.
I'm not sure that is true. The level of back and forth and refinements needed indicate to me that the "English" used is not the normal language I use when talking to people.
It's almost like a refined version of cucumber with syntax that is slightly more forgiving.
Maybe I'm being a codger, but LLMs seem (at least for now) far better for summarizing and giving high level overviews of concepts rather than nailing precise code requirements.
I don't know if "cucumber" is autocorrupt or an actual non-vegetable thing; can you clarify?
That "did they actually mean that or was it autowrong?" feeling is going to get worse I fear.
In the 00s/early 10s, software went through a fad phase where people earnestly thought that by implementing Gherkin frameworks like Cucumber, you'd be able to hand off writing tests to "business people" in "plain English." It went about as well as you'd expect.
Despite that period being when I finished my Software Engineering degree, got my first job, and then attempted self-employment, I'd never heard of it before.
Looking at the book titles — "Cucumber Recipes" in particular — even if I had encountered it, I might have assumed the whole thing was a joke.
On the other hand, many developers write shitty tests where it's absolutely unclear what the test is even trying to achieve, so trying to find some sort of framework which tries to forcibly decouple the what of the test from the how maybe isn't the worst idea.
https://github.com/hitchdev/examples
Rather than trying to force your testers and stakeholders to adapt to the DSL, the templated story->documentation generation lets the dev or tester adapt the DSL to whatever the stakeholders want to see while keeping the story strictly about behavior.
It's also not efficient for doing higher level work. There was a time before we had algebra where people were still expressing the same ideas but the notation wasn't there. Mathematics was expressed in "plain language." It's extremely difficult to read for us. For mathematician's of the time there was no other way to explain algorithms or expressions.
For simple programs I have no doubt that these tools enable more people to generate code.
However it's not going to be helpful for people working on hypervisors, networking stacks, operating systems, distributed databases, cryptography, and the like yet. For that you need a more precise language and an LLM that can reason about semantics and generate understandable proofs: not boilerplate proofs either -- they have to be elegant so that a human reading them can understand the problem as well. We're still a ways from being able to do that.
If we want more programs that are correct with respect to their specifications we need to write better, precise specifications… not wave our hands around.
However for a lot of line-of-business tasks we’re generally fine with ambiguous, informal specifications. We’re not certain our programs are correct with respect to the specifications, if we had written them out formally, but it’s good enough.
I think most businesses that are writing software that needs to be reliable and precise are not going to benefit from these kinds of tools for some time.
And vice-versa! Most software projects do not benefit from the rigor used in aerospace, because it's just not needed, and would be a waste of time.
I am definitely seeing ways that GPT tools could speed up some aerospace work, but we need to be really really sure that things are being done correctly... not just mostly correct, or seemingly correct.
Why should we assume that won't lead to a rabbit hole of misunderstanding or outright hallucination? If it doesn't know what "correct" really is, even infinite levels of supervision and reinforcement might still be toward an incorrect goal.
Of course, it's still bad that humans do it; but despite the scientific method etc., even successful humans often work towards an incorrect goal.
[0] I am cultured, you're quoting memes, that AI is just a stochastic parrot: https://en.wikipedia.org/wiki/Emotive_conjugation
I'm referring to the behaviour, not the inner nature.
> in fact extant AIs are notorious for doubling down instead of accepting correction.
My experience suggests ChatGPT is better than, say, humans on Twitter.
I've had the misfortune of several IRL humans who were also much, much worse; but the problem was much rarer outside social media.
> Even more importantly, the AI doesn't have real-world context, which is often helpful to notice when "correct" (to the spec) behavior is not useful, acceptable, or even safe in practice.
Absolutely a problem. Not only for AI, though.
When I was a kid, my mum had a kneeling stool she couldn't use, because the woodworker she'd asked to reinforce it didn't understand it and put a rod where your legs should go.
I've made the mistake of trying to use RegEx for what I thought was a limited-by-the-server subset of HTML, despite the infamous StackOverflow post, because I incorrectly thought it didn't apply to the situation.
There's an ongoing two-way "real-world context" miss-match between those who want the state to be able to pierce encryption and those who consider that to be an existential threat to all digital services.
> a human who knows about physical properties like mass or velocity or rigidity will intuitively honor requirements related to those
Yeah, kinda, but also no.
We can intuit within the range of our experience, but we had to invent counter-intuitive maths to make most of our modern technological wonders.
--
All that said, with this:
> It doesn't learn from its mistakes until it gets the equivalent of a brain transplant
You've boosted my optimism that an ASI probably won't succeed if it decided it preferred our atoms to be rearranged to our detriments.
Since the inner nature does affect behavior, that's a non sequitur.
> we had to invent counter-intuitive maths to make most of our modern technological wonders.
Indeed, and that's worth considering, but we shouldn't pretend it's the common case. In the common case, the machine's lack of real-world context is a disadvantage. Ditto for the absence of any actual understanding beyond "word X often follows word Y" which would allow it to predict consequences it hasn't seen yet. Because of these deficits, any "intuitive leaps" the AI might make are less likely to yield useful results than the same in a human. The ability to form a coherent - even if novel - theory and an experiment to test it is key to that kind of progress, and it's something these models are fundamentally incapable of doing.
I would say the reverse: we humans exhibit diverse behaviour despite similar inner nature, and likewise clusters of AI with similar nature to each other display diverse behaviour.
So from my point of view, that I can draw clusters — based on similarities of failures — that encompasses both humans and AI, makes it a non sequitur to point to the internal differences.
> The ability to form a coherent - even if novel - theory and an experiment to test it is key to that kind of progress, and it's something these models are fundamentally incapable of doing.
Sure.
But, again, this is something most humans demonstrate they can't get right.
IMO, most people act like science is a list of facts, not a method, and also most people mix up correlation and causation.