It seems still unclear how much quality loss there is compared to the best models. What's really needed is systematic evaluation of the output quality, but that's tricky and relatively expensive (compared to automated benchmarks), so I understand why it hasn't happened yet.
Edit: I just tried it with a single task of my own (that I've successfully used with ChatGPT and Bing) and it flubbed it horribly, so this model at least is noticeably inferior to the SOTA, which is not surprising given how small it is.