ksqlDB: Event streaming database for stream processing applications
ksqldb.io
ksqldb.io
https://docs.confluent.io/current/ksql/docs/developer-guide/...
It’s a long list of caveats about how your queries will silently fail if you don’t guarantee all sorts of properties of the data.
Anyone considering making Kafka the center of your data infrastructure should also consider using a conventional column-store database instead. The only advantage of Kafka in this comparison is better latency (seconds vs minutes). If you can live with a minute or two of latency, the advantages of a real database are MANY.
Now, should Kafka be central place for routing data? I think that is the main reason why it exists. Should it be used for permanently storing data like conventional database? Absolutely not.
By design, it will temporarily store it in case consumer goes offline, but you can also use it to implement backpressure/queue for slow receivers that are not capable to catch up with fast data ingestion (e.g. syslog -> kafka -> logstash).
So, this is probably the last "database" you should ever opt to use, unless you're in a very serious bind.
So if you emit an event (like "New User Created") and have a ksql table that summarizes the user count, then Kafka having accepted the event does not mean the ksql tables have also been updated.
Contrasted with a traditional database where a COMMIT returning one one session means the second can immediately read it.
Now if it were possible for it to accept an event from Kafka but not guarantee that the event will eventually make it into the materialized view but may be lost, that'd be a problem.
This is usually not a problem these days as it’s possible to guarantee exactly once ingestion using Kafka offsets