> I wonder how the economics of this works - I'm guessing token markup?
We're mostly thinking about building tools now for a future in which many gamers will have the compute to locally run a sizable ~30B model. It's hard to me to see LLM gaming taking off in the very near term due to the compute costs for any model of reasonable immersiveness.
> I wrote a library for structured inference in-browser[1]
Very cool! Excited to see where this ends up. Yeah, OpenAI is a bit annoying since you can't ask for JSON or the logit distribution. But GPT3.5-turbo right now is even cheaper than running your own llama2-70b so we stuck with that for the demo since most people probably don't have sufficient compute for this.