Nice! I've been wondering similar things about whether you could use this to eek out more intelligence through methods like these, to quote the end of my write up:
> Does structured decoding increase the observability of emergent world models in these models? To make an analogy: I may not represent an opinion of how something works if I am not confident in it, but if I am forced to present an opinion we might find out that I in fact have (or have not) grasped something.
In practice, however, without tight integration with beam search, the autoregressive nature of these models means that the syntactic steering may result in the models rabbit-holing themselves without forward looking visibility that's obvious from the defined grammar. I.e. if it was forced to choose between "Don't Jump" and "Do run" in some hypothetical example, the set of tokens that it would likely be deciding between is "Don't" and "Do" with no idea what is going to end up syntactically required after those tokens.