Also, similar to Orca-Math but without a teacher model. They also followed an iterative DPO/KTO scheme, but with no length normalized NLL loss term.
Specifically for codegen, i am playing with an iterative interpreter that can quickly (re)evaluate a tree of similar responses