https://github.com/nucleuscloud/neosync
(I'm one of the co-founders)
173 karma · joined January 24, 2021
https://github.com/nucleuscloud/neosync
(I'm one of the co-founders)
Maybe some day, it gets better but for now, we've found that using a more traditional algorithmic approach is more consistent.
Transparency: founder of Neosync - open source data anonymization - github.com/nucleuscloud/neosync
When you write mock data, you almost always write "happy path" data that usually just works. But prod data is messy and chaotic which is really hard to replicate manually.
This is actually exactly what we do at Neosync (https://github.com/nucleuscloud/neosync). We help you anonymize your prod data and then sync it across environments. You can also generate synthetic data as well. We take care of all of the orchestration. And Neosync is open source.
(for transparency: I'm one of the co-founders)
also - happy to chat further if you have any questions - evis@neosync.dev
I spent a lot of time building tokenization solutions at a previous startup so we'll definitely support tokenization at some point. There is a good use-case for it as well!
Even if you know the distribution of the data (which imo can be fairly difficult) replicating that can also be tricky. If you know that a gender column is 30-70 male - female, how do you create 30% male names? How about the female names? Are they the same name or do you repeat names? Does it matter? In some cases it does and in others it doesn't.
What we've seen is that it's really use-case specific and there are some tools that can help but there isn't a complete tool set. That's what we're trying to build over time.
We're actually evaluating a clickhouse integration at the moment for a customer that we're working with so that might be coming in the future. Although today just PG and Mysql.
To answer your questions:
1. We don't support this quite yet although we're working on it for both anonymization and synthetic data. For anonymization, that typically means having deterministic anonymizers that output the same value for the same input (like a hash). For synthetic data that means using a model to be able maintain those same statistical characteristics. We'll have support for both of these within the next quarter. It also depends on what you want to anonymize. If the values that you want to anonymize wouldn't meaningfully change the distribution of the data (think like a name or an address and you're not doing any analytics or queries on those fields) then the statistical distribution of the data stays the same.
2. A few big differences. PG Anonymizer doesn't handle referential integrity, pretty much everything has to be defined in sql and it doesn't have a GUI and it doesn't have any orchestration across environments or databases. Neosync supports all of those.
Folks use the cloud service because they don't have the resource or time to deploy/run the OSS offering themselves. These are usually startups who are okay with us streaming their data and anonymizing it and sending it back to them. We're SOC2 type 2 compliant and usually got through a security review for these deployments. Conversely, they can also just run our managed version and keep all of their data on their infra while we host the control plane.
it's a combination of creating a random number of records for foreign keys i.e 1 customer - create between 2 and 5 transctions. Working on giving you control over that, and handling referential integrity with table constraints (foreign keys, unique constraints, etc.)
ML based approaches typically are not very good at this and struggle with handling things like referential integrity. So a more "procedural" or imperative way is slightly better. The ideal is a combination of both.
It's really down to the use-case. If you're strictly doing development, then you'll probably want to use more synthetic data than anonymization. If you care about preserving the statistical characteristics of the data then you can use ML models like CTGAN to create net new data.
Definitely a balance between when do you anonymize vs. when do you create synthetic data.
We just launched our brand new AI Data Generation feature which allows you to use ayn LLM to generate synthetic data and insert that directly into a Postgres or Mysql database.
Simply connect your database, configure any LLM that you want to use (we support any model that is hosted on an endpoint, it can be local as well!), then provide a prompt and generate data.
We first sample 10 records to give you an idea of what the data looks like and you can iterate on the prompt and data until you're ready.
Neosync handles all of the infrastructure, prompt-chaining, formatting and orchestration of the data from the LLM to your database.
Here are some use-cases: - If you're building a new app you can use Neosync to seed your database - If you're working on a new service that has it's own database or schema, you can use Neosync to seed it with data. - If you want additional data for fine-tuning or RAG, you can use Neosync to generate that data
What's next?
We can generate up to 1000 records right now and are working on supporting up to 10k in the next few weeks. Currently it works for a single table but we're working on making it work for an entire database including all of your constraints. Here's a 5-min demo video:
https://www.loom.com/share/79af81c12b7543fd9d3174c22842ccf2?...
You can run it locally using Docker Compose or Helm or try our hosted platform for free at neosync.dev
Thanks!
(disclaimer - co-founder of Neosync: github.com/nucleuscloud/neosync, open source tool to generate synthetic data and orchestrate and anonymize data across database)
(I'm one of the co-founders)
If you're interested (github.com/nucleuscloud/neosync)
Agreed on the GC issue, going to be doing some perf testing on it and see how it performs. Expecting it won't be as fast as if it was written in C but curious to see the delta.
1. I want to better learn Pytorch and how it works so what better way than to just re-implement some of its core features. 2. I write mainly in Go and haven't come across a lot of ML support in Go 3. I'd rather have a Go ML service instead of spinning up additional infrastructure to just support a python ML service in my Go projects 3. Go's static typing, native concurrency (avoid GIL problem in python), efficient memory management, single binary deployment and more make it a better interface compared to python IMO
Knowing the Pytorch is mainly implemented in c++ and c under the covers, I'm not expecting any performance gains by porting it to Go. But still, will be interesting to see how it compares.
Check it out and let me know your thoughts!