The thing is that those super large models with CoT are often much more reliable than specialized fine-tuned smaller models you'd like to use on trillions of documents, as I am observing right now. You are not trading just speed but also quality with smaller models. The only relevant use case for fine-tuning for me recently was to apply wizard dataset to LLaMA 2 for a less censored model and that was done using LoRA.