Bottled Water – Real-Time Integration of PostgreSQL and Kafka
confluent.io
confluent.io
I ran into both of these:
https://github.com/confluentinc/bottledwater-pg/issues/53
https://github.com/confluentinc/bottledwater-pg/issues/61
They're fixed now, but this is very much beta stuff. It wasn't as magical as we'd hoped so we ended up ripping it out and switching the application to use SSE[1] which streamed into a kafka producer.
Most systems do cache/index invalidation or HDFS archiving from the web app or a set of microservices that gets notified from the web app. So instead of pushing raw DB rows into Kafka they will push generic events e.g. "user deleted account" and let the various services work out how to respond to it.
In this model you have tightly coupled both the choice of DB as well as the DB schema to the events system.
Sounds like a bit of a nightmare if your iterate a lot on your architecture.
Your application domain model should be at the centre of your architecture not the physical database model. For example storing a User object rather than a row from a User table.
I understand why a database company would see a database as the centre of the world. But it really should be your application. Especially if you want to use PostgreSQL and InfluxDB for different domain types and yet have both indexed in ElasticSearch.
The application domain model can be buggy and full of holes. The data store is the source of truth.
It's also pointless to try and convince someone who thinks otherwise either, I found it better to let the those developers shoot themselves in the foot and learn the hard way. It's the only way they'll learn.
It's mostly useful when you can't change the application code.
That said, I'll warn you, we didn't find Bottled Water to be anywhere near robust enough for our use cases. It seems to make assumptions about the size of your database, the size of your transactions, and the downside risks of logical replication backing up that certainly didn't meet our requirements.
(authored by the creator of this tool)
IIRC this works closely with Samza. There's a decent video by the author knocking around somewhere.
EDIT: This is the video I was thinking of:
Instead we just write directly to durable queues on RabbitMQ (there are some exceptions where extreme consistency and transactions are needed).
That's is we do web->queue->databases.
There are serious cons to this bus approach as well as pros like improvements in latency.
The bottle approach is nice because you can write old school CRUD style.
Maybe I can use parts of bottle to write to our rabbit bus for the parts that are CRUD because of consistency reasons.
It supports HDFS and ElasticSearch as well.