Hi! From Ollama here - you can run:
ollama run qwen3.8 (or if on mac qwen3.8:27b-mlx)
43 karma · joined May 27, 2021
The greedy accept is so that the mask doesn't need to be computed. Planning to make this more efficient from either ends.
With the newer research - outlines/xgrammar coming out, I hope to be able to update the sampling to support more formats, increase accuracy, and improve performance.
Hoping to be more on top of community PRs and get them merged in the coming year.
The current implementation uses llama.cpp GBNF grammars. The more recent research (Outlines, XGrammar) points to potentially speeding up the sampling process through FSTs and GPU parallelism.
Hopefully with those changes we might also enable general structure generation not only limited to JSON.