We've run experiments on datasets ranging from 5K - 100K examples, which gave fantastic results [1].
Some examples - https://huggingface.co/datasets/b-mc2/sql-create-context - https://huggingface.co/datasets/GEM/viggo
On the other hand, 8K examples was not enough to learn to solve grade school math problems [2], so it is very problem dependent.
[1] https://www.anyscale.com/blog/fine-tuning-llama-2-a-comprehe...