> ...so you need a checker that actually reasons about the world (compiler, linter, SAT solver, ground-truth dataset, etc.).
Agree. What do you think about telling the LLM to also generate unit tests for the code it spits and then run all tests (including previous application unit tests).
I think this is a way to ensure some level of grounded verification:
- Does code compile?
- Do unit test pass?
AI can then consume test results to help fix their own mistakes.