The analysis and proposed solutions all seem to ignore the one important aspect of this data: time.
Why would you wait until the end of the month to count the recorded activities on the first of the month? And, what happens to failed activities that are repeated on the first of the next month? Are customers double-billed? If I'm interpreting the scenario correctly, a large network outage on April 30th would see a large number of duplicate charges as they are re-applied in May.
Counting during the send process makes much more sense. Duplicates can be handled in-line, and even detected over month boundaries. If its not "worth it" in terms of engineering time, then I'd be concerned about the leadership; the author indicates 'this was the only way we made money'
I have no problem with 'boring' solutions; actually, I prefer them. Processes that are simple, self-explanatory and/or easy to follow are usually superior IMHO.
It makes sense too. Billing requires a lot of man hours to pull off. Invoices, auditing, excel spreadsheets, etc. It's a people problem not an engineering problem.
I'm not sure of the disconnect here, but de-duplication is not trivial. If you do it every day nothing is easier than if you do it every month.
Doing it for all time is completely infeasible.
There was not a database of every piece of social media data sent out the door. That's what you would need to make sure not to record the entry again. All we had were flat files in s3.
Big flat files. It took hours to download and merge them all.
Once we had a database (Cassandra) it was updated continually (by a kafka consumer) and we could query it in a few minutes.
Oh, and Cassandra is great... used it for years. Just beware of over-sized results. I've seen run-away queries crash the instance(s) running the query; ugly stuff.
Glad you found a solution that works!
* downloaded the files daily
* sorted and removed the intraday duplication in each new file.
* sorted and removed intraweek duplication each week and on the first of the month.
* You could give your customers a weekly estimate of their bill.
All of this would spread the work out and tell you when the system had glitched out weeks before the results were due.
Also, the Unix approach can do something like this by spitting out days at a time and sorting them, and then doing a
sort -m "$month/*" | uniq -c
on the individual files.`sort -m` takes basically no memory, and should chew through data fast enough.