Few thoughts, you're effectively creating representations that can convert to JSON (kudos!)
Can't mention how we did it (there are a lot of public patents, if interested), but back in 2018 we had a way to generate synthetic data (statistically, structurally similar) off any dataset - https://medium.com/capital-one-tech/why-you-dont-necessarily... You could also design datasets if you wanted.
It'd keep similar relations and worked pretty darn well. Not the exact same, but always produced valid JSON.