Never had any problems managing outages.
Grep alone can go a long way in managing outages.
Never had any problems managing outages.
Grep alone can go a long way in managing outages.
I could see orgs standing up their own solutions like ELK but our service didn't have to. We just relied on grepping logs stored in Timber for logs older than an hour and grepping logs on prod hosts for real time searching during outages. Granted, our service did not have many dependent services but AFAIK the retail website which has tons of dependencies also followed a similar model (along with using RTLA for fatals) atleast at that time (circa about 3 years ago).
Also our log collector agents run in containers like everything else, so there is some amount of resource isolation (not perfect of course).
All the log grepping for data older than the current hour happen off-prod host and thus was never a concern.