Function calling might be one reason. If your app has a lot of custom functions for interacting with tools, fine tuning may be preferred over using context tokens.
Here is an example of data prepared for fine tuning Llama for function calling...
https://huggingface.co/datasets/mzbac/function-calling-llama...
I'm unaware of any comprehensive guides - we're still in the wild west.