My initial thought is: restore last nights backup to another mysql instance on aws and then let it catchup on the binlog?
But I guess the unstated assumption is that their goal is to also transform to some other datastore.
My initial thought is: restore last nights backup to another mysql instance on aws and then let it catchup on the binlog?
But I guess the unstated assumption is that their goal is to also transform to some other datastore.
For example, at MindMup, we’re using a similar setup to the one described in the article for appending events to user files. It’s critical that we don’t get two updates going over eachother in a single file, which is tricky to do with lambdas only because there’s no guarantee how many lambdas will kick off if updates come concurrently from different users for the same file. With Kinesis, we just use the file ID as the sharding key, so no more than a single lambda ever works on a single user file, but that we can have multiple lambdas in parallel working on different files.
How are you appending to files from Lambda? EFS?
> With Kinesis, we just use the file ID as the sharding key, so no more than a single lambda ever works on a single user
Is there a risk of "heavy" users causing hot shards?
I'd like to invite you to watch this talk and get a few more insights why we use Kafka in the ways we do: https://www.youtube.com/watch?v=cU0BCVl4bjo
Let me try to come up with a TL;DR here: trivago comes from a complete on-premise, central database point of view. Change Data Capture via Debezium into Kafka enables a lot of migration strategies into different directions (e.g. Cloud) in the first place, while not having the need to change everything on the spot.
It seems like a common pattern to compare Kafka with a pure MQ technolgy. Kafka can also serve as a persistent data storage and a source of truth for data.
I hope this makes the picture a bit more clear to you. Feel free to ask if I missed something.
Debezium seems to be a production version of Martin Kleppmann's CDC-to-Kafka POC, Bottled Water [1].
Database replication is the killer app for CDC, but CDC can be used for so much more than replication, like event-based alerting, triggering, etc.
[1] https://www.confluent.io/blog/bottled-water-real-time-integr...
While the basic idea of using PG logical decoding for CDC is the same, Debezium is a completely different code base than Bottled Water. Also we provide connectors for a variety of databases (MySQL, Postgres, MongoDB; Oracle and MongoDB connectors are in the workings). If you like, you can also use Debezium independently of Kafka by embedding it as a library into your own application, e.g. if you don't need to persist change events or want to connect it to other streaming solutions than Kafka.
In terms of CDC use cases, I keep seeing more and more the longer I work on it. Besides replication e.g. updates of full-text search indexes and caches, propagatating data between microservices, facilitating the extraction of microservices from monoliths (by streaming changes from writes to the old monoliths to new microservices), maintaining read models in CQRS architectures, life-updating UIs (by streaming data changes to Web Sockets clients) etc. I touch on a few in my Debezium talk (https://speakerdeck.com/gunnarmorling/data-streaming-for-mic...).
Any plans to support SQL Server? (SQL Server is prevalent in the enterprise world)