Why wouldn't we apply the mask immediately for the first sampling? Is this an optimization somehow, is masking expensive?
Why wouldn't we apply the mask immediately for the first sampling? Is this an optimization somehow, is masking expensive?
Other libraries work by essentially pre-computing all the masks for all possible generations, but of course you're restricted to working with simple grammars in this case (like a subset of regular expressions)
> is masking expensive?
It's not expensive per-se; A single element-wise multiplication of the output vector.
The real "expense" is that you need to prepare masks for every element of your grammar as they are expensive to recompute as needed; LLM tokens do not cleanly map onto elements of your grammar. (Consider JSON: LLM tokens often combine various special characters such as curly braces, colons, and quotes.)
This isn't that hard to compute, it's just more work to implement.
The greedy accept is so that the mask doesn't need to be computed. Planning to make this more efficient from either ends.