I authored the blog with some other contributors and worked on the feature (PR: https://github.com/ollama/ollama/pull/7900).
The current implementation uses llama.cpp GBNF grammars. The more recent research (Outlines, XGrammar) points to potentially speeding up the sampling process through FSTs and GPU parallelism.