You wouldn't actually want to, because you'd be losing generalizability, and it's a lot of unnecessary work.
I think approach #1 outlined above is the better (more cost- and time-efficient) technique—where a pretrained model already understands JSON (among myriad other formats), and you merely constrain it at text-gen time to valid JSON (or other format).