Generate Synthetic Data in 3 Lines of Code
gretel.ai
gretel.ai
If you want to know how you would go and use this for your own problems check out some of our other posts
https://gretel.ai/blog/how-to-safely-work-with-another-compa...
FYI, I'd built a hand-crafted generator using JSONNet templates and Golang, but I really wanted something that could model source data distributions accurately. The use case is large-scale load testing of customer workloads without requiring actual data.
They get 43,300 records per second on this example, which seems to the right order of magnitude for you
Next would be the ability to decompose or factor my entire db into the subcomponents and make suggestions for combining tables. The use case would be a legacy enterprise system that has grown so complex and tangled that devs are afraid to do basic db refactoring. From where I'm sitting and working this is the next gold rush, apply DL methods to the bread and butter computing, log flow analysis, etc.
its a massively under researched area because relational data has more... dimensions to it and is understandably not as exciting to most.
happy to discuss (email in profile)
I love the idea of "table space" though. It would be fun to traverse this space and output a new database at each step, like a VAE.
They mention in the podcast that most customers end up finding relationships in their tables that they didn't know they had - that weren't explicitly in schema
We've also recently released a new platform called Djinn that is specifically designed for data science workflows. It enables you to query from tables across your DB to build customized views of only the data you need and synthesize high-fidelity data based on models trained on those views. Relationships are fully preserved and no external scripting is required. You can create an account and take it for a spin here: https://djinn.tonic.ai/?signup
Full disclosure, I'm Chiara Colombi, Product Marketing Manager at Tonic.ai. Cheers!