"Healing" code requires significant amounts of nuance. An understanding of the ecosystem both inside the code and outside. This is even a difficult task for humans to get right. LLMs do know a lot, but they don't understand a lot. Hell, ChatGPT still tells me that a kilogram of feathers and a pound of bricks weigh the same, just because it expects a certain pattern. Even explains why they aren't, but insists they are right. (in before "but GPT4")
--
Prompt: Which weighs more, a pound of bricks or a kilogram of feathers?
ChatGPT: A kilogram of feathers and a pound of bricks weigh the same. One kilogram is approximately equal to 2.20462 pounds. So, in this case, a kilogram of feathers and a pound of bricks both weigh about 2.20462 pounds.
-- As a followup --
Prompt: Are you sure that's correct?
ChatGPT: Apologies for the confusion in my previous response. You are correct, and I apologize for the incorrect information. A kilogram of feathers and a pound of bricks do not weigh the same.
A kilogram is a unit of mass in the metric system, while a pound is a unit of weight in the imperial system. Since weight is affected by the force of gravity, the weight of an object can vary depending on the location.
However, if we assume the objects are weighed under the same gravitational conditions, then a kilogram of feathers would weigh less than a pound of bricks. One kilogram is approximately equal to 2.20462 pounds, so a kilogram of feathers would weigh less than a pound of bricks.
--
And we can see it does even worse. I had to ask 5 times and it gave me every possible answer (it continues to loop if you keep asking btw. It doesn't converge). You don't know when to stop if you don't already know the answer. It will think it is wrong if you keep asking because that's stochastically what it expects: correct answers are accepted, wrong answers questioned. Memory is not understanding, even though they often look similar. LLMs are great tools, but they aren't a replacement for thinking. They require lots to use them.