That's mere millions of data points per day, not trillions.
Like... why do they need real time machine learning at that rate? No, really: Why!? Would it be catastrophic to their business if they ran ML as batch jobs? Do their recommendations need to change second-to-second?
Another crazy quote: "At that time, Netflix had ~500 microservices, generating more than 10PB data every day in the ecosystem."
Wat?
That's a 170 MB of logs pre customer per day! Most of those customers might watch an episode of a TV show per day, or watch one movie per day. Call it 2 hours of usage per day? These guys were blasting out 1.4MB of logs per minute per active user while essentially doing nothing much more than streaming big binary blocks of data from a CDN!
In my mind that's the crux of the issue. The architecture astronauts have gone amok, solving problems that shouldn't exist...
The argument / excuse is: "Many product features, such as personalized recommendation, search, etc., can benefit from fresher data to improve user experiences, leading to higher user retention, engagement, etc."
My experience with NetFlix is that their recommendations are garbage and getting worse over time. Exabytes of data won't solve this.