[1]: https://techcrunch.com/2019/02/19/google-acquires-cloud-migr...
[1]: https://techcrunch.com/2019/02/19/google-acquires-cloud-migr...
I highly recommend teams not rolling out their own solutions because of all these edge cases. My boss refused to pay a vendor solution because it was a 'easy problem' according to him. Denied whole team promotion because we spent lot of time solving 'easy problem' , outright asked us 'whats so hard about it that it deserves promotion' . Worst person i've ever worked for.
Keeping track of updates without a dedicated column.
Ditto for deletes.
Schema drift.
Managing datatype differences. I.e., Postgres supports complex datatypes that other DBs do not.
Stored procedures.
Basically any features of a DB that are beyond the bare minimum necessary to call something a relation database is an edge case.
I rolled a solution for Postgres -> BQ by hand and it was much more involved than I expected. I had to set up a read-only replica and completely copying over the databases each run, which is only possible because they are so tiny.
- Data type mismatches between systems
- Differences in handling ambiguous or bad data (e.g., null characters)
- Handling backfills
- Handling table schema changes
- Writing merge queries to handle deletes/updates in a cost-effective way
- Scrubbing the binlog of PII or other information that shouldn't make its way into the data warehouse
- Determining which tables to replicate, and which to leave behind
- Ability to replay the log from a point-in-time in case of an outage or other incident
And I'm sure there are a lot more I'm not thinking of. None of these are terribly difficult in isolation, but there's a long tail of issues like these that need to be solved.