A naive thought: What you would get if you hardcode the language grammar and not let the training discern it, so instead of it, kinda like an expert system constraining its output?
https://en.wikipedia.org/wiki/DisCoCat
>In this post, we are going to build a generalization of Transformer models that can operate on (almost) arbitrary structures such as functions, graphs, probability distributions, not just matrices and vectors.
https://cybercat.institute/2025/02/12/transformers-applicati...
I don’t think that would help with natural language to programming language, but that can probably help with patterns, kinda like a powerful suggestion engine.