I think of fine-tuning as an avenue to significantly reduce LLM inference costs, so I think this is an exciting development. You're right if you compare GPT-3.5-turbo to fine-tuned GPT-3.5-turbo, but if it's anything like fine-tuning the Llama-2 models, you'll be able to achieve GPT-4 level performance for a wide range of practical use cases (SQL query generation is an example), but probably not for math or coding (at least not without fine-tuning on a significant amount of data).
In fact, we've seen GPT-4 level performance from even the 7B Llama-2 model after fine-tuning. [1]
[1] https://www.anyscale.com/blog/fine-tuning-llama-2-a-comprehe...