I am honestly heartbroken to live in a world where reading the code is seen as an unreasonable ask by either students or by professional working programmers.
I am honestly heartbroken to live in a world where reading the code is seen as an unreasonable ask by either students or by professional working programmers.
- having your job responsibilities being reduced to ONLY reviewing code.
- having to review code unnecessarily high scrutiny because it can hallucinate randomly and you as a human are responsible for the code even though you didn’t write it. In a traditional context when I review code, there’s a shared responsibility. Someone writes the code and another person reviews it. Now it’s entirely on the person who reviews it.
There may be other buckets, these are the ones that I hear often from other engineers.
They are also among more recent inventions, they are not "the traditional" programming at all. It is not like code review was the thing that attracted people to the profession or something that would be ore rewarding part of it.
I'm just upset that we are throwing away the original prompts for generated code in such a cavalier fashion.
An LLM prompt, even a huge one, is an incredibly vague document that leaves out most of the edge cases. And even Fable 5 happily ignores clear instructions in its prompt.
Now, to be fair, I absolutely expect the buggy slop to win, and to drive out the people that either write their own code or at least read the output. This will, in turn, make customers less willing to spend money on software after they get burnt a few times by buggy garbage. I think this is pretty much inevitable once Fable returns. It's just too damn good at long time horizon tasks, generating far more mostly sorta working code than any human could reasonably read.
How do you know your compiler is a rigurous and deterministic? Did you review all of its code?
LLMs are intentionally not deterministic, nor is their output subject to any known specification. Output is a point in a high dimensional manifold, determined by the input vector, but this manifold is unknowable in a real and intractable sense.
These are not equivalent constructions and it demeans you to conflate them.
...I agree it's not deterministic (neither are all your variations of C compilers, neither is Firefox v Safari v Chrome), but it probably Does Something(tm), and I might not want to peel back the covers and see how it used React, or Vue, VanillaJS, QT, or GTK.
It's upsetting that we are _committing the generated code_ rather than being able to use better and better optimizing compilers against the original prompt of: "make me a user registration form with database connection"
...I'm very with you on "it's not an accurate analogy", but I'm pointing out that there have been sea-changes already w.r.t. strict adherence to the generated code, or inclusion of left-pad v react libraries.
...and there have been corresponding productivity gains (debatable? ;-) when we've worked at these higher levels of abstraction.
I'm personally still in the "blacksmith" stage of working with AI output (put it back in the fire and beat on it a bunch more times), and shudder in horror at the thought of maintaining (or paying to maintain) megabytes of hours of token generation that looks like source code.
I'm hopeful that we'll eventually strip out some of the mud between the CPU and putting pixels on the screen (with the help of LLM's?), and that we'll still be able to understand and reason about the real "DAG" of what our programs are trying to do (eg: declarative guis, kindof like we have declarative sql), but there will always be a muddy middle part where the computer/complier/LLM is doing something in between that _is_ sufficiently reliable for us to ignore those bits most of the time.
When you write code and it compiles the outcome that you get is entirely predictable. An LLM prompt is not at all predictable.
Overall though agree entirely with your sentiments. And my llm usage is also blacksmith esque, in that its like hammering steel into shape. "Sharing my prompts" would be hundreds of "move this method" "rename this to that" "extract line 12-24 to a new method" etc etc
...I think we've been exponential (or s-curve) w.r.t. repeatable predictability that _smaller_ prompts are yielding. eg: `regform.php: register( first, last, address, email )` would have had a 30% "one-shot" outcome 1-2 years ago, but closer to 90% "one-shot" outcome this year.
Smaller prompts == wider outcomes... potentially wider (larger?) prompts == smaller outcomes?
As we're "committing" code (blacksmith-type), presumably it's b/c we've reached a plateau of "probably works right" and includes some level of testability and fitness-for-purpose, ie: we've rolled the dice enough times and are committing the current "yahtzee" board to the repo.
16 years ago I've been advocating effectively that "acceptance tests are more valuable than the code under test" (b/c given sufficient acceptance tests, it's "trivial" to reproduce the code correctly... but given the code and no tests it's at least an order of magnitude harder to prove correctness)
http://www.robertames.com/blog.cgi/entries/to-upgrade-or-not...
It's again the dual hourglass shape of consensus between creation and acceptance, and LLM's (when used well) are helping to compress the two sides together around an "exact" agreeable crystallization.
"I want you to calculate pi" (wide hourglass) <=> ...etc... <=> [(precise code/assembly) <=> (precise validation)] <=> ...etc... <=> "pi should start with 3.14", "pirr should equal area", ...etc...
...but throughout time, we as developers have been the human interpreters / vessels via "english-to-buggy-code", and now we have this whole new set of (semi-unpredictable) tools. Think of woodworking (hackers), with hand-tools and eventually moving to power tools w/ jigs. Some things are lost, but others are gained! Maybe we're needing to come up with "jigs" (or "harnesses?" ;-) to strap around our new LLM power-tools?
Closing thought... I've done some sketchup modeling, eg: mocking my kitchen to consider a remodel. "Models must represent to a precision useful for their purpose". Counters, sink, fridge... cool! But then, should I include the toe-kick cutout, or cabinets == flat boxes? Should I include the counter overhang? What about representing the cabinet doors? Should I include the cabinet knobs? Should I actually measure the true openings down to the 16th inch? Validate that they're 90deg rather than 88deg square? On the counter overhang, should I include the rounding? What if it was routed, is the shape of the overhang important?
Basically, reality has infinite detail and our "captured models" of the problem domain are always working at a level of detail that is "fit for purpose". Design a nuclear power plant? LLM == "probably not". Design a local CD ripper GUI for yourself? "the level of detail and correctness is probably fit for purpose"