Ask HN: How important is production like data to your development process?
However because so many people seem to have solved this issue (each with their own twist though), I've been thinking about that creating the environments and deploying them may not be the important part? Perhaps it is the DX around the data that is supplied to these environments. I think the data can come in a few different manners:
* A brand new database with the correct schema attached, however there is no data. This is useful to be able to use the environment, but a pain if you have to insert a lot of data to be able to use the environment. This is probably the easiest and could be done through a container running Postgres/Mysql/etc
* A database with the correct schema and some seed data. This fixes the scaffolding issue so the environment is immediately usable but doesn't allow for reproducing bugs from production. Again, this could be done through a container.
* A production cloned database. This allows for debugging any production bugs or issues. However without being able to scrub PII from the data, this opens security concerns. This approach could be done through a container, but doing say pg_restore into a container is going to make the startup time for this environment a lot slower. Another approach is to use something like RDS and backup snapshots to create the clones and have them ready to be used by an environment before hand.
* A scrubbed production clone. Same as above, but with all the PII removed. I think this is the top tier and would give the most benefits to developers without the security concerns.
I'm curious what other peoples thoughts are around this topic and how / if your company is providing production like data to the development process?