If you do end up wanting to fine tune then use qlora with axolotl or unsloth to prove your hypothesis on a smaller model and then evaluate if you want the marginal gains you get from full precision training.
After you fine tune it with 100m token dataset, use DPO to polish it off. You need to create a DPO dataset for that but it can be relatively small to get some great gains.
After that, look at applying grammars during inference if you are expecting structured results like json.
You should be able to run the experiments on 4090s from vast.ai or runpod or similar service.
It can cost less than $100 depending on your requirements.