Iteratively refines code, shifting “accuracy bottleneck” from correct code gen to correct test gen
HumanEval accuracy:
-Reflexion-based GPT-4 88%
-GPT-4 67.0%
-CodeT 65.8%
-PaLM 26.2%”
with link to code in the Tweet:
https://mobile.twitter.com/johnjnay/status/16393620718075494...
21% improvement after adding a feedback loop and self-reflection to GPT-4, which just went public 12 days ago. (The approach is based on a preprint published 4 days ago.)
Human coders often need a feedback loop and self-reflection to properly “generate” code for problems novel to them as well.
-----
A larger question: Are we hurling ourselves toward a (near) future of unaligned AGI with self-improvement capabilities?