Logging at Twitter
blog.twitter.com
blog.twitter.com
Cloudflare - https://blog.cloudflare.com/http-analytics-for-6m-requests-p...
Uber - https://eng.uber.com/logging/
Facebook - https://research.facebook.com/publications/scuba-diving-into...
For those who are smaller and don't have the money to pay for Splunk enterprise, and don't have the headcount to build your own logging infrastructure, I built a a service called GraphJSON that makes it super easy to log and process any type of data. You can read more about how and why I built it here https://www.graphjson.com/guides/about
I was hoping TFA would break down their log-based observability strategy and go into things like trace ids, structured logs....
Instead I am disappointed to learn that engineering at Twitter still sounds.... suboptimal. This from the company that brought us the fail whale (I can't find the blog post now, but I recall Twitter's instability was due to a massive kludge that went unnoticed for years).
https://engineering.fb.com/2019/10/07/data-infrastructure/sc...
(I worked in Scribe)
The main issue (as usual in big companies) is the large amount of inter-dependencies with internal systems. Scribe as-is today doesn't make much sense outside Facebook. And yeah, it could be cleaned up, mock the internal services and OSS it. But that's a lot of work, both doing it and maintaining it. And having all those mocks, etc "wouldn't be Scribe as-is either"in terms of how e.g. it scales and so on.
In any case, the storage (LogDevice, https://github.com/facebookarchive/LogDevice) which is a large part of the system was open sourced a while back and... sadly it went again into "not-maintained OSS".
Finally, you can get a similar-ish system working with OSS tools (e.g. fluentd + Kafka) that will also scale quite well and which IMHO can also made to scale to Facebook-size levels. So, the incentive is there, but there are OSS alternatives already available :)
The internal improvements aren't interesting or available to anyone outside of facebook.
The open source version is creaky and very difficult to drag forward to newer versions of thrift.
So -- thanks Facebook for dropping weird artifacts like the aftermath of Roadside Picnic. Woe unto all who adopt these soon-to-be-orphaned technologies.
"Logging" is
* ingestion / message schema
* queue
* ETL
* storage / analytics
The stickiest part of a logging pipeline is the ingestion and message format (scribe in this case); those are the gravity and speed of light of your organization, embedded at the lowest layer of everything everywhere.Three years from now they could / will write a new blog post "today we're open sourcing smackhouse, a log analytics platform based on smooshing logs into clickhouse! We ETL the data using this zany rust thing (zaptl) from kafka into clickhouse, and made some zippy UI on top. We still stream logs into hadoop and goolens and splunk; any teams that want to use these other infrastructures just set a config in a fargml setting to zaptl to request the log stream and sample rate."
They're not using all of splunk; they're only using it to solve one of the harder parts of logging -- log analytics. They're not using it to solve the stickiest part of logging.
Good for them. BTW -- this reads like a cross between a proof of life video and a newsletter written by the intern. Glad it's earning them a huge discount from splunk.
(edited for formatting)
I hope in the near future these types of problems are “definitively solved” and more effort can be spent engineering novel features.
There’s no way to do this, but I’d be interested to see if there’s a company that just has an extremely simple stack that is very performant and popular.
I’m talking: their entire stack is Citus sharded over a few locations running two beefy servers per continent they have users in, that’s it. No dedicated search, logging or anything. Just dump everything in Postgres. If logs are lost then so be it.
A lot of the problems at scale are problems of our own creation. For example is strong consistency really necessary for tweets? Do you really need to see a followers latest tweet instantly?
Uncheck some boxes and the engineering becomes a lot simpler, very fast.
Multiple times he fired/hired senior members of his administration via twitter, including defense secretary, secretary of state, national security advisor, chief of staff, ...
I’m not saying that people don’t overcomplicate projects (they do) but a lot of simplifying assumptions can break down at scale.
This is the misunderstanding here. With a sufficiently complex system, they are never 'just logs'.
The problem is that "strong consistency" basically means your system behaves in predictable ways that you can easily reason about. If you give it up, you can see all kinds of failure modes that go beyond just seeing stale data.
Suppose a user first sets their profile to "private", meaning only their trusted friends can see their posts. And then, having seen that that operation succeeded, they make a post that contains personal, sensitive information. You had better be damn sure that no user-facing components of your system can ever observe those two events in the wrong order.
Things need to do what they say on the box
I have spent much of my time at this company on this exact problem.
I agree that current tweets area important, of course. But I also acknowledge that it is just as often that i will be scrolling through the "In case you missed it" section of tweets. And, I have no idea how rapidly those actually made it to my timeline.
You could almost argue they the small like counts have to be strong consistency. But, I often get notification of a like well before it shows on my timeline. So it is clearly not strongly consistent.
I think you're oversimplifying to the point of nonsense. In my domain reliable logging is an absolute must due to regulations and the fact that we deal with people's money. Pretty sure we're not the only industry that can't afford approaching logging as you painted it.
twitterblog: blog-post-ending.html: twitterblog failed: No space left on deviceI said dam. And that’s the logging data.
I do like how it ends up with “and then we basically ssh in and restart stuff by hand for reasons.”
That could have fit in a tweet.
The article said 42 TB per datacenter. How many datacenters is Twitter running?
I've always been a proponent over leaving logs where they were produced, not collecting them, and not indexing them, so an architecture like this I find just shocking.
Never had any problems managing outages.
Grep alone can go a long way in managing outages.
Also our log collector agents run in containers like everything else, so there is some amount of resource isolation (not perfect of course).
All the log grepping for data older than the current hour happen off-prod host and thus was never a concern.
I could see orgs standing up their own solutions like ELK but our service didn't have to. We just relied on grepping logs stored in Timber for logs older than an hour and grepping logs on prod hosts for real time searching during outages. Granted, our service did not have many dependent services but AFAIK the retail website which has tons of dependencies also followed a similar model (along with using RTLA for fatals) atleast at that time (circa about 3 years ago).
Many of the issues presented in the article ring very true. Splunk is pretty amazing for adhoc analysis/threat hunting. However, once you know what you’re looking for the value proposition drops precipitously.
No need to say send all you debug logs to security, or all your info logs to a logging instance for site reliability.
That was real strange coming from Amazon, where I was told to log anything and everything... And often, without the logs, it would have been impossible to determine root cause of some issues.
The difference was that logging was super cheap at Amazon, because we'd store all the data, but no indexing on the contents of logs.
If your service is small, or if you are only interested in logs for a specific hosts, you could download, and grep through it... Or run some sort of mapreduce jobs against it. But that was a long time ago. Surely they have better tooling now.
For security logging, Google cloud security (formerly Chronicle) lets you send unlimited data (but less control of the data or what you can do with it of course).
It has competition left and right from sumologic to elastic cloud, perfect time to sell to Cisco!
I've used Kibana and BigQuery, Splunk is lightyears ahead. It's not just for logging, it is excellent at big data analytics and visualization. This is what I struggle to communicate with people that develop and deploy these products. I can write a stupid front end to grep too if I just want to query and regex. I want the query language to let me extract and manipulate fields and their values very easily , let me measure all kinds of stats, control the output, pipeline between outputs and then visualize that data where possible.
You wanna see how manu unique users of Chrome 99.x.y.z transfer how much traffic and how frequently they see what page in your web logs? That's like 3-4 quick SPL lines in Splunk. Everyone else expects you to write parsing somewhere, stats elsewhere and visualization some other place and even then so many limits. Non-splunk users write pages of Jupyter notebook to replicate a short |stats splunk command.
- How do you deal with standardization across the different events you log? How do you handle standardized naming, representations (formatting) and types?
- What about validations? do you validate the payloads in any way? if so, at what layer? Client before logging? In Kafka?
- Any privacy capabilities? How do you make sure the data is not accessed by processes/people that's not supposed to access a datum?
Thanks!
https://www.gravwell.io/newsroom/gravwell-upgrades-community...
Disclaimer: we use them at Shodan and I'm on their board.
And at the same time the article is very vague on by who and how all that data is read. A much more interesting question for a blog post is: how many bytes of the collected log data is used to create actual business value?
In the beginning there were so many great technologies there for dealing with logs and data at scale and they caved in and bought splunk.
An insanely expensive solution that was already long solved there for years
Also, when you have bespoke log solutions and engineers leave, you end up with unsupported solutions. This may be Twitters 4th logging solution IIRC. Plenty of people know splunk. Loglens, not so much.
> Logging at Twitter (2021)