I think it's almost never worth it now. An up to date LLM with a good few-shot prompt, or an agentic LLM will likely outperform a small finetuned LLM, at a lower cost with more flexibility.
Moreover, it's very seldom that you have enough high-quality data to do a meaningful fine-tune.
I have successfully finetuned a small Gemma to play chess badly, and that worked. Mostly because chess data is plentiful and because my aim was to not have perfect answers. I'm self hosting the model inference because running it in the cloud would be too costly.
I also finetuned small LLMs using too small datasets of business specific human conversations and the results were disappointing.