Fable and Astra level models have become relatively good at these sorts of things, have not tried it with CPL but Lisp and Fortran which I doubt there is much of a significant training body on in the corpuses, but they perform extremely well. I think scaling has reached the point that these sort of things are just part of the emergent behaviors of AI