The issue I find with this is that the frontier models still outperform the small finetuned model on its specific task. So much so that the ROI on doing fine tunes is likely negative. I would love to hear some specific example where it did provide value though, if any has any. That would be helpful to start being able to find similar cases.