With Gemma, or any open model, you can use the open libraries in conjunction to get what you want. Some inference frameworks like Ollama include structured output as part of their functionality.
But you mentioned all of this already in your question so I feel like I'm missing something. Let me know!
But I think you already mentioned all this in your response so I might be missing the question?
Edit: per simonw’s sibling comment, ollama also has this feature.
The Gemma model by itself does not though, nor does any "raw" model, but many open libraries exist for you to plug into whatever local framework you decide to use.
Under the hood, it is using the llama.cpp grammars mechanism that restricts allowed logits at each step, similar to Outlines.
- We can constrain the output of a JSON grammar (old school llama.cpp)
- We can format inputs to make sure it matches the model format.
- Both of these combined is what llama.cpp does, via @ochafik, in inter alia, https://github.com/ggml-org/llama.cpp/pull/9639.
- ollama isn't plugged into this system AFAIK
To OP's question, specifying a format in the model unlocks training the model specifically had on functions calling: what I sometimes call an "agentic loop", i.e. we're dramatically increasing the odds we're singing in the right tune for the model to do the right thing in this situation.