you could, but it the model is not optimized for text
this is complex, but generating text is highly complicated and requires mode dropping to make long cohesive text
Or a partially completed song, asking for the next note. I’m not sure if you’re joking, but using it for space constrained next token generation within a grammar sounds like a really neat use case.
Feed the generated note back into the input for the next query and you have ... autoregression?
If you provide it an AST of the english language, yes.