Very cool and also very simple as I’d expect from Cloudflare.
But I have a question - why not make inference as easy as the translation? Why do I have to run that in a worker rather than just as a simple API call? That would be much simpler.
Is there a technical reason or is it that people would want to have logic before making the call to llama?