Greenmask: PostgreSQL Dump and Obfuscation Tool
github.com
github.com
Let's take, say, a set of a hundred thousand nine digit social security numbers, stored in a DB column that has a uniqueness constraint on it. This is a space small enough that hashing doesn't really mask anything -- there are few enough possible values that an adversary can compute hashes of all of them, and unmask hashed data. But the birthday paradox says that the RandomString transformer is highly unlikely to preserve uniqueness -- ask a genuinely random string generator to generate that many nine-digit strings, and you're extremely likely to get the same string out more than once, violating the uniqueness constraint.
One approach that I've seen to this is to assign replacement strings sequentially, in the order that the underlying data is seen -- that is, in effect, building up a dictionary of replacements over time. But that requires a transformer with mutable state, which looks kind of awkward in this framework. The best way I can see to arrange this is a `Cmd` transformer that holds the dictionary in memory. Is there a neater approach?
Then, during hashing you just need a constant immutable state, which effectively expands the hash space, without incurring the mutable state overhead of replacement strings strategy.
We ran it weekly to dump production down to a “scrubbed” copy which we used for our test/qa environments and then further trimmed some large tables to create a much smaller development version.
It really helps to have the “shape” of production data and we saw surprises/bugs on prod deploys due to unexpected data drop to almost nothing. It was much easier to maintain than seeds imo, though it is a cat and mouse game, if a new column with sensitive data is added, you need to ensure it’s covered with a masking query too, but we solved that by auto tagging the security ops folks whenever our schema file changed on GitHub.
It wasn’t perfect, but it was 95% what we needed
Anyone have a solution to this problem that works well for them?
With PG you can bypass the FK checks locally by temporarily changing the replication mode from origin to replica. Though if you don't eventually copy the referenced records you may run into FK errors during normal operations.
I have a simple one like „base_tables“ that just pulls me all the fixtures into my local db, then entity specific ones that pull an entity with a specific id + all related entries in other tables for debugging but as long as you can query it you can set up everything very easily.
My experience is you can usually write the application with no dependencies on customer data, and do local development entirely with synthetic data.
For debugging specific customer escalations or understanding data distributions for your synthetic data modeling, you can use a replica of production. No need to extract just a subset or have devs host it themselves. All access can be controlled and logged this way.
Better than a replica would be something like [1] but that isn't available everywhere.
I agree the problem of "dump a DB sample that respects FK constraints" comes up sometimes and is interesting but I'm not sure I'd use it for the problem you describe.
[1] https://devcenter.heroku.com/articles/heroku-postgres-fork
I have a prelease build (v2.0.0-alpha) available at https://github.com/dbsnapper/dbsnapper/releases
And I'm updating the documentation for v2.0.0 in real-time https://docs.dbsnapper.com. The new configuration is there, but I have some work to do adding docs for the new features.
Works with Mac and Linux (Windows soon)
If you're interested in giving the alpha a try, I'd love feedback and can walk you through any issues you might encounter.
Joe
What is the license for DBSnapper ? GitHub does not seems to list one
It's closed-source / proprietary at the moment, but I'm highly considering opening parts / all of it up as a golang library at a minimum, or do a free/pro version to support development. Any thoughts, suggestions?
(disclaimer - co-founder of Neosync: github.com/nucleuscloud/neosync, open source tool to generate synthetic data and orchestrate and anonymize data across database)