ParentFull threadqwerty3344·I think it's still significantly behind GPT 3.5/4, both of which can get 67% on HumanEval, and 88% with ReflexionView on HN