Many of the issues presented in the article ring very true. Splunk is pretty amazing for adhoc analysis/threat hunting. However, once you know what you’re looking for the value proposition drops precipitously.
I've always been a proponent over leaving logs where they were produced, not collecting them, and not indexing them, so an architecture like this I find just shocking.
Never had any problems managing outages.
Grep alone can go a long way in managing outages.
Also our log collector agents run in containers like everything else, so there is some amount of resource isolation (not perfect of course).
All the log grepping for data older than the current hour happen off-prod host and thus was never a concern.
I could see orgs standing up their own solutions like ELK but our service didn't have to. We just relied on grepping logs stored in Timber for logs older than an hour and grepping logs on prod hosts for real time searching during outages. Granted, our service did not have many dependent services but AFAIK the retail website which has tons of dependencies also followed a similar model (along with using RTLA for fatals) atleast at that time (circa about 3 years ago).
The article said 42 TB per datacenter. How many datacenters is Twitter running?
No need to say send all you debug logs to security, or all your info logs to a logging instance for site reliability.