It seems to me less of an issue that LLMs produce code that aren't compliant with language specs, but rather that LLMs produce code that do not meet the stated requirements, or that they break from earlier stated requirements as new requirements are requested.
And of course, code not meeting requirements is not a problem native to LLMs, but to all code production agents everywhere (including humans).
This nails one of the biggest issues that we're trying to help Mito users with -- code that meets __changing__ requirements. ie: the input data column headers change.
The best approach we have right now is giving users flexibility to parametrize their script (ie: tell the script the right column header or reference columns by position instead of name)
It turns out that teams that have moved their data to Snowflake tend to avoid these issues primarily because the schema of their data changes much less frequently.
LLM/ChatGPT is based on stochastic approach to NLP while the alternative e.g. Typed Feature Grammar (TFG) is based on deterministic approach [1]. Apparently Cuelang is based on the latter and since it's a constraint based language you can use it against the prompt (input) and also for checking the output as well [2],[3].
What it means is that LLM/ChatGPT and TFG/Cuelang can be used in concert to produce or generate codes by well crafted prompt based on the language spec and similarly also to debug based on the spec. This will probably change and enhance the discipline of software engineering as we know it for the better.
[1]Feature (linguistics):
https://en.wikipedia.org/wiki/Feature_(linguistics)
[2]The CUE Data Constraint Language:
https://github.com/cue-lang/cue
[3]Show HN: GPT-JSON – Structured and typehinted GPT responses in Python:
I wonder if Clojure can do it... I think the spec can check if function-calls and data given have valid types and reasonable unit/quantity each... Hmm.
In general, it feels it's halting problem hard to determine if any general purpose programming language program is executable or safe... but maybe if we reduce the size of the spec under consideration?
I've personally never seen a Mito users with a detailed enough spec of their report that an LLM would be able to use it to check compliance -- but maybe if we built the functionality users would create them...