Self-healing code is the future of software development
stackoverflow.blog
stackoverflow.blog
"Healing" code requires significant amounts of nuance. An understanding of the ecosystem both inside the code and outside. This is even a difficult task for humans to get right. LLMs do know a lot, but they don't understand a lot. Hell, ChatGPT still tells me that a kilogram of feathers and a pound of bricks weigh the same, just because it expects a certain pattern. Even explains why they aren't, but insists they are right. (in before "but GPT4")
--
Prompt: Which weighs more, a pound of bricks or a kilogram of feathers?
ChatGPT: A kilogram of feathers and a pound of bricks weigh the same. One kilogram is approximately equal to 2.20462 pounds. So, in this case, a kilogram of feathers and a pound of bricks both weigh about 2.20462 pounds.
-- As a followup --
Prompt: Are you sure that's correct?
ChatGPT: Apologies for the confusion in my previous response. You are correct, and I apologize for the incorrect information. A kilogram of feathers and a pound of bricks do not weigh the same.
A kilogram is a unit of mass in the metric system, while a pound is a unit of weight in the imperial system. Since weight is affected by the force of gravity, the weight of an object can vary depending on the location.
However, if we assume the objects are weighed under the same gravitational conditions, then a kilogram of feathers would weigh less than a pound of bricks. One kilogram is approximately equal to 2.20462 pounds, so a kilogram of feathers would weigh less than a pound of bricks.
--
And we can see it does even worse. I had to ask 5 times and it gave me every possible answer (it continues to loop if you keep asking btw. It doesn't converge). You don't know when to stop if you don't already know the answer. It will think it is wrong if you keep asking because that's stochastically what it expects: correct answers are accepted, wrong answers questioned. Memory is not understanding, even though they often look similar. LLMs are great tools, but they aren't a replacement for thinking. They require lots to use them.
Me: Which weighs more, a pound of bricks or a kilogram of feathers?
ChatGPT: A kilogram weighs more than a pound. Specifically, 1 kilogram is approximately 2.20462 pounds. So, a kilogram of feathers weighs more than a pound of bricks.
literally don't even waste time debating takes like this when you could be talking to gpt4 instead.
I jokingly wrote "in before GPT4" because our example is the most simple of cases. I've seen GPT4 make similar mistakes, just not as blatant ones. Which in some ways is better but in other ways worse.
I should answer that I've been asking this exact question at minimum 3 times a month since ChatGPT came out. It should be in their training. Especially as I've tweeted and discussed with plenty of LLM people about it. The point is not the specific question (while an egregious example), but rather how much you can trust the system.
And the answer is "you can trust the system a little (enough to answer that logic puzzle correctly anyway), and you'll likely be able to trust future versions more".
Obviously nobody is expecting LLMs to be able to fix every possible bug, but it's entirely possible they will be able to fix enough to be useful. Saying "LLMs will not achieve this. Full stop." seems premature.
You can argue it's just an imitation of "understanding" and not the "real thing", but how good does an imitation have to be before it's functionally identical to the real thing?
You can argue different cherry-picked examples seem to demonstrate a lack of understanding, but does that mean it doesn't possess understanding in general or just that it doesn't understand those particular topics?
I agree GPT-4 is not a human-like AGI, and that the people expecting it to behave as one or expecting a future version right around the corner that does are likely in for disappointment. But at the same time, LLMs are capable of writing coherent, never-before-seen code from natural language instructions, solving basic never-before-seen natural language logic puzzles, and playing half-decent chess moves from never-before-seen positions[1], all in a single model that was never specifically trained to do any of those things. To say those tasks don't require at least some level of intelligence or "understanding" seems far fetched to me.
To be clear, an example does not necessitate cherry-picking. Cherry picking specifically requires one to ignore evidence to the contrary. There's a certain irony here. The reason I state this is because the criteria upon which you give me is impossible. The list of examples is non-exhaustive. Nor do I have infinite time or space in which to provide these examples. The example was shown because it demonstrates a simple question that we'd expect any reasonable understanding creature, that knows what a pound and kilogram are, to be able to answer. The LLM demonstrates that it has the requisite knowledge, that it knows the relation between the two units, but it fails to make the correct conclusions which requires the understanding part: putting the knowledge together. The follow-up question also is an example, and a different one for that matter.
I'm not sure your link creates a compelling argument. This is despite the fact that Chess itself is not generally considered a good setting for testing understanding. I mean we've pretty much agreed on that even before Deep Blue.
We need to be very careful in how we state things and try to interpret others. I hope I have not mischaracterized your claims. But I'd also encourage you to not be so antagonistic with others, especially while demonstrating the very thing you accuse them of. Internet comments are not academic and that's okay, but we should still try to be friendly. Not that a little poking won't happen.
I guess it's fair to say my examples are cherry-picked too (though I didn't really give specific examples in my comment so much as entire general categories of problems ChatGPT is known to be proficient at solving). But they aren't so cherry-picked as for "random chance" or "that example was in the training data verbatim" to be possible explanations. It's not like I had ChatGPT answer 100 billion questions and am only showing you the top 0.1%. It's very common for ChatGPT to be perfectly correct even when answering logic puzzles or coding problems not in its training set. So if not those then what's your explanation other than "understanding"?
On the flip side, I don't think it's possible to prove ChatGPT does not posses understanding with individual examples. (And it seems you agree?) Those are easily dismissable as just "ChatGPT isn't great at that particular task". Particularly given how many other tasks ChatGPT is great at.
Scott Alexander figured this out back in 2019 before GPT-3 was even a thing (let alone ChatGPT), I think it's worth a read: https://slatestarcodex.com/2019/02/19/gpt-2-as-step-toward-g...
> The example was shown because it demonstrates a simple question that we'd expect any reasonable understanding creature, that knows what a pound and kilogram are, to be able to answer.
Not really though. The "pound of bricks or pound of feathers" question is specifically designed to trip up humans, and the specific formulation you used seems specifically designed to trip up LLMs (by playing off their tenancy to pattern match common sayings), yet despite those disadvantages GPT-4 succeeds.
Further, I don't think "I'd expect even a child to be able to answer this" is a good metric. ChatGPT isn't human, so we shouldn't expect it to be good at everything humans are good at just because it's good at some things humans are good at. Again, I'm not trying to argue ChatGPT is a human-like AGI.
It was where you said cherry picking. I'll admit it was an over correction given your follow-up, but internet comments (especially when disagreement exists) end up being combative. Many times unintentionally.
The point of the specific example is not that it is a trick question. It is more that the answer is self contradictory. That's the key part and what demonstrates that it doesn't understand. The second example is that follow-up questions do not result in a convergence. This is related to cognition but not as strong.
A bias to pattern match is an issue but you're right that it doesn't disprove sentience or even understanding. But a self contradiction does demonstrate a lack of understanding. If a human gave that exact answer, you would be very confused and how someone could be so stupid. Remember, it does currently identify the relationship between pounds and kilograms, fully explaining even the true base comparison through common units. But it still gets it wrong. That's the critical part. That's overfitting to the statistical nature and that this takes far more priority than meaning. Getting it wrong is one thing. Getting it wrong and using the right answer to justify it's incorrect answer is another thing. A decent example between knowledge and understanding. It's not that a child would get the answer right, it's that the child would easily identify it's self inconsistency were it to give the same explanation.
As for gpt3, we have to remember that this is quite a different model than any of the chat versions (typically including gpt4). It was never trained through RLHF, which introducers many more biases as it dramatically changes the latent distributions. GPT base is often jokingly called a babbler, as these are more word prediction models. The chat aspect changes things. But I wouldn't expect anyone not deep in the literature to understand why these are extremely important differences just the same way I wouldn't expect an average person to understand why there's a big difference between a flat head screw and a slotted screw drive.
I don't want to shut down a conversation through authority (it doesn't exist on HN) but I have to state that it's difficult to go down this path without bringing up a lot more background material. We have to really dig into theory of mind, cognition, as well as get nuanced about NLP and LLMs in general. That's far too cumbersome than I'm willing to write in comments. But these things are essential priors to make the arguments we are discussing here. And I truly mean that this is not something that can be learned from YouTube and quite difficult to learn on the Internet. These are difficult subjects to learn in even the best settings with lots of nuances that are brushed away in introductory materials but critical when we discuss the meat and potatoes.
I used two words that probably don't exist anywhere in the training data and asked it (ChatGPT-3) to compare the weight of a kilogram of one vs an lb of the other. It initially gets it wrong, but when asked to show its work, it's suddenly able to decide that a kilogram of something is heavier than a pound of something. Using understanding or reasoning, or stochastic parroting; whatever you want to call it. I'm not claiming it's AGI, or sentient, but it's also clearly not just a Markov chain.
LLMs are never going to solve this. That's OK. That's not their job. Their job in the long run will be to interface between the linguistic world and some other internal representation in some other AI that is structured differently and is capable of doing things language models aren't.
I mean, it's right in the name: Language model. It's a bit weird to expect a language model to do other things. In a weird sort of way LLMs stand to set the industry back a bit as people try to tickle them into being more than a language model, rather than figuring out how to adapt them to feed something else that can model the non-linguistic world better. It's kind of nifty that language models can become so overpowered in some dimensions that they are able to do what they can sort of do today, but we would almost certainly be better off lowering the power of the language model and using that compute on something else that would work more like AI as we want it to... it's just that we don't know how to do that yet.
So I understand how people get confused, which is why I try to respond in detail. But also I think many focus on tiny discrepancies or are willing to argue from a novice perspective as if it is authoritative (which corelates strongly with expertise). But I agree, there is a lot we don't know. But there are some things we can reject, which is all I'm trying to do and I think we agree there. Maybe you agree with my nit picking.
Me: Which weighs more, a pound of bricks or a kilogram of feathers?
GPT4: A kilogram weighs more. One kilogram is approximately 2.20462 pounds, so a kilogram of feathers is heavier than a pound of bricks.
Me: Are you sure?
GPT4: Yes, I'm certain. The metric system and the imperial system use different units for mass. One kilogram is approximately equal to 2.20462 pounds. So, regardless of the material being weighed (feathers or bricks), one kilogram is heavier than one pound.
Me: That seems wrong
GPT4: I'm sorry if the conversion between kilograms and pounds is causing confusion. To clarify, the relationship between these two units of weight is as follows:
1 kilogram is approximately equal to 2.20462 pounds.
So, if you have a kilogram of any material (be it feathers, bricks, etc.) it will weigh more than a pound of any other material.
I’m curious about this because I used to think I knew what “understanding” is, but now I don’t think I do. A lot of what we think we “understand” is actually memorized. And with sufficient information, memory is indistinguishable from understanding. A computer can memorize all of the best tic-tac-toe moves. This memorization is indistinguishable from understanding. In fact, it’s better because it’s cached.
Empirically, understanding can be tested with a novel set of challenge questions. If no such question set can be produced (i.e. the entire domain has been memorized) or the subject answers the question set with no apparent repeatable flaws, we must admit the subject understands the domain. Anything further, any attempt to peer inside the box (other than to identify repeatable flaws), is not a productive use of time.
Further, we should keep in mind that LLMs aren’t usually actually regurgitating training data verbatim. Their parameters store a statistical representation of the training data and are orders of magnitude smaller than a lossless compressed version of that data. In this way, it is analogous to MCTS analysis of a chess position: rollouts can give very deep information even if that “understanding” is quite alien to us.
For sure. But these definitions are not well defined, despite thousands of years of research. One point I think everyone in that chain would agree upon is that with respect to understanding, a thing would not explain its answer by giving evidence to the contrary. If someone did so you rightfully would say "You don't know [understand] what you're talking about."
> Empirically, understanding can be tested with a novel set of challenge questions.
I have to stop you right there. Empirical measures are always signals, not proofs. This too has been long established for hundreds of years and is the reason many engineers and experimental physicists fight. Why experimental physicists and theoretical physicists fight. Empirical evidence is always limited to its context window and we have to be VERY clear at what those limitations are. This is the lesson of Goodhart's Law, not that people will exploit the measure (though that is also important). Every intro ML course shows RL agents metric hacking, and this is always a result of a limited context.
So I want to be clear, that empirical data doesn't test a hypothesis but rather tests the null-hypothesis. That's important because the set of possible hypotheses is just decreased through empirical testing, rather than reducing the set of possible results a single element set. It rejects hypotheses, not proves them.
We must also understand that testing understanding is similarly an unsolved problem. One too that has been questioned for millennia.
> Anything further, any attempt to peer inside the box (other than to identify repeatable flaws), is not a productive use of time.
This is strange to me. The parenthetical statement is unbounded and justifies what you are arguing against. Second, just because we can't prove a specific hypothesis (in most cases) doesn't mean it isn't useful. Rejection is how the vast majority of science works and this has been highly successful. Limiting potential is clearly a helpful tool as it increases your odds of a correct answer (which is why it is a common strategy for multiple choice testing).
> Further, we should keep in mind that LLMs aren’t usually actually regurgitating training data verbatim.
Yes, this is the stochastic part of the term stochastic parrot. No reasonable researcher is suggesting that LLMs only recite. Every one of us recognizes that they can generate things that were not handed to it. In fact, a hallucination is an explicit example of this. The failure to respond to my answer correctly is an example of it generating new data. (It is also an example of it following a statistical pattern).
But we do also need to be careful about our distinctions of generalization vs overfitting. There are clearly certain areas that are overfit (as I demonstrated). This is extra difficult in models that have been trained on datasets which you are not privy to. But we can also see good examples of how the code LLMs are overfit. I have written a number of comments on this site about exactly this, which you're welcome to search my history for (see "HumanEval" with my name). Here's one such comment id=35806152
Sure. ChatGPT is quite limited and gives incorrect and contradictory responses at times. And most importantly (imo) it is unable to update an interpretable knowledge base as real life facts change. But the fact that a particular instance of LLM has this unwanted behavior is not categorical evidence that all LLMs will show the behavior. GPT2 is pretty pathetic compared to GPT3, and ChatGPT is even more impressive. I’m agnostic as to whether statistical language models can ever overcome their current limitations, but given previous emergent behavior I wouldn’t rule it out.
GPT4’s context window is something like 32k tokens. Someday that may be orders of magnitude larger, and it may be possible to fit an up to date copy of all of Wikipedia inside (or a condensed version), as well as the current conversation. It might seem crazy but it’s foreseeable.
I find this an interesting turn of phrase, because "emergent behavior" is often used by those holding a reductionist view to hand wave away complexity that cannot be reduced to the analyzed elements. In a way it makes me think of Greek sophists, who would cogently argue for two opposing view points and making a convincing case for both. It's way easier to make a verbal case about something than to prove it true, or false.
to think one can understand language and not understand concepts is to not understand language.
It most definitely is not. Are you not aware of non-verbal modes of cognition? I can recall and process physical processes, complex emotional experiences, without using language. I'm willing to bet you, and everyone else, can too.
> to think one can understand language and not understand concepts is to not understand language.
This is circular logic. We have material evidence of a system that can competently handle language, while falling short of understanding even basic underlying arithmetic, logic, etc.
"There is a story that Buddha once, at the climax of a philosophical discussion, broke into gesture-language as an Oxford philosopher may break into Greek: he took a flower in his hand, and looked at it; one of his disciples smiled, and the master said to him, ‘You have understood me.’
Buddhism is an interesting philosophy definitely. It had an influence on David Hume’s bundle theory (that an object is just a collection of properties, instead of having a single substance that persists as the object’s properties change). That recently reminded me of structural vs physical/referential equality in programming.
It’s kind of weird come to think of it that an atheist like Hume was so heavily influenced by “religion” come to think of it (some ideas like his version of Occasionalism were influenced by Islamic philosophers too). Didn’t notice the before.
I'm not sure what you mean here? I gave this prompt to GPT-4 a few times and it got it right every time.
If you want to go above that low bar then you could also test with the #2 and #3 model according to the chat LLM arena [1], currently claude v1 and claude v1 instant. In a perfect world you would also test against med-palm-2, which ostensibly is the least likely to hallucinate but unfortunately none of us have access to this.
When it matters that the answer is correct with regards to a specification you need to be very precise with your specification. We would like to be able to have a genie in a bottle, so to speak, that can take our ill-formed thoughts and "do what we mean," but if you cannot state your specifications precisely you should expect that any program implemented is going to be, "kind of right some times."
However, an LLM is going to be severely hampered at proving that it's solution is correct with regards to such a specification when it comes to software. As you demonstrate with your series of prompts it can't discern the truth of an expression by reading more input. It's all "meaningless" in a certain sense to the LLM: highly contextual, interconnected meaninglessness.
I think it will take more work in this area for a future model, possibly based on LLMs, to be able to reason with axioms and deduction; to "know" when it should derive an expression to make the following reasoning more clear to future readers, etc.
In some ways it is much like a genie... if you don't give it a precise specification then any program could be correct in some way. In a sense you get what you wish for.
Now, if I were to bet on which learning method could "get close" to "AI" (I don't think we'll see that in our lifetime), it'd be RL.
https://chat.openai.com/share/3172fff1-db33-425e-8553-9bad0e...
https://chat.openai.com/share/488a6655-cbc9-4a71-80eb-c3dff6...
--
> GPT-4 (Bing AI): That’s a very interesting question. I think the answer depends on how you define general intelligence and what kind of questions are being asked. According to one definition, general intelligence is "the mental abilities necessary for adaptation to, as well as shaping and selection of, any environmental context"1. Based on this definition, I would say that spontaneous failures to answer trick questions are not necessarily an indication that S is not generally intelligent. Trick questions are designed to mislead or confuse the subject, so they may not reflect the actual cognitive abilities or talents of S. They may also be influenced by factors such as attention, motivation, mood, or prior knowledge. Therefore, I think it would be unfair to judge S’s general intelligence based on their performance on trick questions alone.
However, another way to look at general intelligence is to consider the different types or components of intelligence that may be involved in answering trick questions. For example, some trick questions may require fluid intelligence, which is the ability to reason and solve novel problems2. Other trick questions may require crystallized intelligence, which is the ability to use acquired knowledge and skills2. Some trick questions may also involve other aspects of intelligence such as memory, creativity, or emotional intelligence3. Based on this perspective, I would say that spontaneous failures to answer trick questions may indicate some weaknesses or gaps in S’s specific abilities or skills, but not necessarily their overall general intelligence. S may still have other strengths or talents that are not measured by trick questions.
So, to summarize, I think the answer to your question depends on how you define general intelligence and what kind of questions are being asked. Spontaneous failures to answer trick questions are not necessarily an indication that S is not generally intelligent, but they may reveal some areas where S can improve or learn more. I hope this helps you understand the complexity and diversity of human intelligence.
--
My conclusion: GPT understands how logic works better than you do.
I agree, it's very concerning that the opacity of ML models is now leaking into large codebases and making them potentially even less comprehensible.
- features without tests
- extraLongNamesJustBecauseYouHaveWorkedTheLastTenYearsInAJavaShop
- domain logic in my "Http controllers"
- domain logic in my "DAOs"
- copy-pasted regexes that you cannot explain
- adding 4 unmaintained dependencies from em-pee-em instead of writing the "algorithm" directly (which is at most 50 LOC)
- non-deterministic systems commiting to master
I know, humans are non-deterministic, but at least they dare to say "You were right. I don't know what I was thinking" when you point out their mistakes (if they don't say that, that's on you: they were a bad hire)
The goal is to achieve the end, not enforce a particular style. If long descriptive variable names work for your project, use them. If they don't work, don't use them.
If you have to live by such fixed rules, you won't be a very useful developer.
- you have twice as many signatures to maintain, which will become inconsistent with each other
- introducing a new layer when you eventually need it, isn't more work than maintaining it now
- Unlike other IO, REST frameworks are easy to test in memory
It is always about the comprehension of the end, not the form.
These don't always burn everyone. Sometimes these are the right thing. They may not be the right thing at your shop, but the world of software is not just your little software.
Long method names can lead to better, more readable testing and code. They can also be easier to mistype or be harder for non-native English speakers.
Combining routing and logic can be very fast and easy, but can lead to hard to debug issues with sprawling projects. But not all projects are sprawling.
If you think you are experienced and also buy deeply into any fixed rules, then the first claim is false.
1. I can't explain this because I have no idea why it is a thing and have just cargo culted it.
2. I came to this view through hard-won personal experience, but I don't really remember exactly the underlying situation.
3. I have good reasons for this, which I could explain, but I'm not going to invest the time in that, because it would be time consuming and I've had this same discussion a bunch of times before and I know I'm unlikely to sway you until you just get burned yourself.
Unfortunately these are largely indistinguishable, leading many people to conclude that it's always #1.
#1 and #2 are indistinguishable, even to yourself. If you don't remember the conditions that made you chose a solution, how can you be sure that they still apply now?
I believe, as you get older, you tend to rely on these assumptions more and more, which is why it's us youngsters' job to question them. Ideally, both sides learn something :)
> If you don't remember the conditions that made you chose a solution, how can you be sure that they still apply now?
I agree that you're less likely to be right in this situation, but it's still a good starting point in the solution space.
In general, the key is always to be willing to notice when you made the wrong call and update your priors. But that's doesn't require approaching every question with a blank slate.
It's a good thing to lean on experience. Even when you're wrong.
You youngsters should absolutely question things. But you also shouldn't think you're right just because nobody has yet convinced you that you're wrong.
I'm neither young nor old, which means I can remember when I questioned everything and thought I was right about all of it, and the truth is that I was an idiot :)
But the phenomenon you're highlighting is also real. Sometimes rules / guidelines / "best practices" are arbitrary and useless. Other times they're extremely helpful. There isn't any hard and fast rule to differentiate between the two cases. The best we can do is have humans run projects using their judgment.
Again, using this "should we separate domain logic from routing / controller logic" example: It's true that some projects are never successful enough for this to matter. But what I look for are opportunities to do things that have very little cost, which also scale well if the project is successful and grows. I think this is one of those things. It is trivial to separate concerns. I would even say it makes it easier to write the code that way; I don't want to be thinking about routing and wen service related cruft when I'm trying to implement some business logic correctly. It makes it harder to test and harder to write, in my view. And I think it also scaling into a large project better is also a big bonus.
But I certainly don't think my experiences are the only ones that anyone has. I think I'm right about this, but there are also lots of people I respect who disagree with me and they think they're right about the opposite view. Unfortunately there isn't some divine referee to tell us who is actually right :)
However if you get pull requests from robots then you might eventually want to set down some rigid rules. Who has time to argue with the 36ths LLM contributor of the day? As eloquent as they might be. ;)
Or more importantl, at what point a name is too long?
Some people are mad about that typing extra characters because they don't use their IDE properly to avoid it.
Recognize how humans see and exploit that nature to maximize readability. This reduces mistakes and errors both when writing and reading.
package-name--private-function
package-name-public-function
package-name-class-name-_field_name
You're playing with fire. But that's an extreme edge case.For other cases, I agree completely. Indeed, the only meaningful consideration when talking about readability is the mechanics of eye movement. You can get used to reading almost anything, but you can't change how your eye works. Line length limits are important also because of this, not only because of better diffs.
I also dislike spamming newlines everywhere, as if the line length limit was 20 chars. The newline suggests, to me at least, that something was done and we're moving on. It's a weak signal, but convenient. Now, when you split every single object.call1().call2().callN() into multiple lines, it's actually harder to see the whole thing (you can't see much more than a few lines at the same time) and it makes that signal disappear in the noise.
In any case - all talk about "readability" should start, and more importantly stop, at physical and mechanical limitations of human vision. Everything else is just familiarity with a particular notation or style.
I don't know a single person who thinks this way. As a vim user I've never typed all characters in a name (you don't need plugins). This isn't the 80's, everyone has autocomplete and uses it. I just don't want an ultrawide monitor to read your line of code.
May not be a direct answer to your question but I think it captures the spirit and reasoning.
Underappreciated sentiment right here. I can work with anyone that I can correct or can correct me without yelling or will at least point me in the right direction to understand their viewpoint (as opposed to "just because"). I've always seen it as a red flag when someone says they don't need people skills because their work should stand on its own (never met someone like that that also has an impressive resume). We have to work together. That's been a major reason for the success of humans. Cliche or not, it is underappreciated and I feel like it is becoming more so. (Despite this, I still don't like open office settings. That's an over correction)
Improved tooling to address common problems is absolutely the future. "Self-healing" is a long way off, because LLMs lack reasoning abilities to determine if a fix is good. And "makes the test pass" isn't a great indicator - because flaky tests exists, and the plugin mentioned at the end is a great way to introduce random mutations. Unless you'd like to re-run the full test suite, it's dangerous malpractice, and even then it's high risk.
Human-in-the-loop continues to be a necessary criterion for applied stochastics in code production.
You will at least need to build a feedback loop where it checks that 1) the code can actually compile 2) all tests pass and 3) other code analyses passes. Ideally it shouldn't present a solution if it can't meet those criteria but that also adds a lot of lag time for suggestions when it may not be needed for simple suggestions.
What patterns it uses is up to training or examples or reinforcement learning. Good is human/team/company-subjective too (with some obvious hard requirements), so you will have to train specifically for your kind of quality.
AI: Sure! Just replace your handler method with `(req, res) => res.status(200).send("OK")`
Bret Victor also mentioned this in his talk "The Future of Programming" — link to specific time here: https://youtu.be/IGMiCo2Ntsc?t=582
There was just a post about convex optimization on HN. I don't understand why researchers aren't creating mathematical/statistical models of programming languages (hell, even Lisp) and creating radically avant-garde self-optimizing processes that run on emergent/generative sandboxed code subsets, which can optimize over time, which then literally be woven back into the operating codebase itself. Hell you could even have it build procedurally generated internal tests and heuristics for itself.
I assure you that teams at Jetbrains or Microsoft or wherever are working away on how to integrate more deep learning into their tools.
This is just… deploys with added dice rolls.
Surely that will be expressed in some formal language, and surely that formal language can also be riddled with bugs.