https://www.confluent.io/blog/exactly-once-semantics-are-pos...
> Exactly-once Semantics are Possible: Here’s How Kafka Does it
> Now, I know what some of you are thinking. Exactly-once delivery is impossible, it comes at too high a price to put it to practical use, or that I’m getting all this entirely wrong! You’re not alone in thinking that.
blah blah blah ... of course it's at-least-once, with the "idempotent producer" so only one resulting message is published back to another kafka stream. Big surprise.
Now many people think "kafka has exactly-once delivery, that's what I want, I don't want to have to deal with this at-least-once stuff" when really it's the same thing, and others have been doing idempotent operations of various kinds for years, and the user still has to figure out how to do their desired thing (which might not be sending one more kafka message) in a mostly idempotent way.
> Is this Magical Pixie Dust I can sprinkle on my application?
> No, not quite. Exactly-once processing is an end-to-end guarantee and the application has to be designed to not violate the property as well. If you are using the consumer API, this means ensuring that you commit changes to your application state concordant with your offsets as described here.
I think that is a pretty clear statement that end-to-end exactly once semantics doesn't come for free. It states that there needs to be additional application logic to achieve this and also specifies what must be done.
What you could do though, is a keep a "set" data structure in a replicated redis or some other database, and each member of the set is the id of the message which incremented that count. Duplicates are naturally ignored when adding to the set. The count is the size of the set.
The message might also have a "time-generated" embedded in it. You could group by hour, and after some number of hours assume that no more messages will be re-delivered from that long ago, and consolidate that set down to a single number.
Or maybe that's too much overhead because there are millions of messages per hour. Maybe in that case you can afford the occasional double-count and don't need all this. Trade-offs will be unique for each situation, and I don't think kafka will completely take care of this kind of thing for you.
Relying on another distributed system to ensure the first distributed system isn't duplicating messages sounds like a headache I wouldn't wish on anyone. (I am a Confluent employee).
But you still need to understand how this works to do "kafka transactions" and you still need some other scheme to get effects/actions outside of kafka. (And you'll probably get people doing dumb stuff saying "I was told I get exactly-once delivery")
Exactly once seems to refer to exactly once write to kafka and processing in kafka streams that write back to kafka and don't do stuff like make external http requests, in which case you'll need your logic of at most or at least once to delivery results to wherever you're sending the http requests.
Confluence blogs do seem pretty misleading in terms of hiding these issues.
I just think that they are opaque to what they really mean by exactly-once and definitely hide the fact that they do not offer exactly-once delivery.
Check the sequence number map to determine if the PID is already present.
If it is not, check sequence number.
If it is 0, insert the PID, epoch, and sequence number into the PID mapping table and proceed to append the entries.
If the sequence is not zero, return the InvalidSequenceNumber error code.
If the PID is present in the mapping, check the epoch.
If the epoch is older than the current one, return InvalidProducerEpoch error code.
If the epoch matches the current epoch, check the sequence number:
If it matches the current sequence number, allow the append.
If the sequence number is older than the current sequence number, return DuplicateSequenceNumber.
If it is newer, return InvalidSequenceNumber.
If the epoch is newer, check the sequence number.
If it is 0, update the mapping and allow the append.
If it is not 0, return InvalidSequenceNumber.
That's a pretty straightforward approach of enforcing idempotency; checking producer and sequence IDs and throwing them out if they aren't correctly ordered. The other big addition to the code base was cross-partition transactions, which, while pretty cool, is a far cry from "Exactly once delivery".