Please don't share our links on Mastodon
news.itsfoss.com
news.itsfoss.com
I’m entirely baffled why someone savvy enough to produce monitoring graphs and claim to already use Cloudflare can be brought to knees by HN (maybe Reddit too or something?) traffic on a frigging blog. I believe you need to actively sabotage your Cloudflare settings to achieve this.
This complaint is simply an attempt to shift the blame for their decision to ignore multiple decades of prior art. Even in the 90s we used techniques like caching to avoid this problem, and their page-load times now somehow manage to be worse.
This is even better (I do not know Cloudflare that much, but there is indeed the same on AWS' CloudFront)
1. The "It's Foss" site is designed somewhat carelessly, where many things that could be static aren't, and so it goes under with even a bit of load.
2. Mastodon link preview is badly designed and not spec-compliant. Because the link previews are not triggered by a user request they should respect robots.txt (https://github.com/mastodon/mastodon/issues/21738), or they should start being triggered by a user request (https://github.com/mastodon/mastodon/issues/23662).
Some OP advice to "get a damn CDN" seems to be a responce rather than a correction of the issue.
Even poorly designed websites shouldn't face DDoS scale access by a social network sharing a link. The social network should mitigate it's massive parallelization.
Off-topic, but I love how "mastedon" sounds.
1. You make a post that includes example.com/cats
2. Your post is federated to the instances your followers are on.
3. Those instances fetch example.com/cats in a fully automated way, and observe its preview image is example.com/tabby.jpg and fetch that as well. They save this locally.
4. When your followers view your federated post they see a short text extract from example.com/cats and a thumbnail of example.com/tabby.jpg
Step 3 is where the 'DDOS' is happening, and it's not a request issued by an agent of the original poster.
Thinking about other similar distributed systems, if I put a link to an html page in an RSS feed or a mailing list message, it's not normal for subscribers to be using clients that fetch linked pages. But if they did it's clearly as an agent of that user, and not as an agent of the poster?
I get paid to basically solve this sort of scaling problems so I'm not a great barometer but it isn't that hard to handle a traffic spike like this.
- They've set the generated page to be uncachable with max-age=0 (https://developers.cloudflare.com/cache/concepts/default-cac...)
- nginx is clearly not caching dynamic resources (currently 22s (not ms) to respond)
- lots of 3rd party assets loading. Why are you loading stripe before I'm giving you money?
- why is there a random 'grey.webp' loading from this michelmeyer.whatever domain?
This isn't mastadon or cloudflare, it's a skill issue.
I'm not particularly clear on Fediverse and Mastodon internals, but my understanding is that an image preview request is only generated once on a per server basis, regardless of how many local members see that link. But, despite some technical work at this, I don't believe there's yet a widely-implemented way of caching and forwarding such previews (which raises its own issues for authenticity and possible hostile manipulation) amongst instances. (There's a long history of caching proxy systems, with Squid being among the best known and most venerable.) Otherwise, I understand that preview requests are now staggered and triggered on demand (when toots are viewed rather than when created) which should mitigate some, but not all, of the issue.
The phenomenon is known as a Mastodon Stampede, analagous to what was once called the Slashdot Effect.
There's at least one open github issue, #4486, dating to 2017:
<https://github.com/mastodon/mastodon/issues/4486>
Some discussion from 2022:
<https://www.netscout.com/blog/mastodon-stampede>
And jwz, whose love of HN knows no bounds, discusses it as well. Raw text link for the usual reasons, copy & paste to view without his usual love image (see: <https://news.ycombinator.com/item?id=13342590>).
https://www.jwz.org/blog/2022/11/mastodon-stampede/> Have you considered switching to a Static Site Generator? Write your posts in Markdown, push the change to your Git repo, have GitHub/GitLab automatically republish your website upon push, and end up with a very cacheable website that can be served from any simple Nginx/Apache. In theory this scales a lot better than a CMS-driven website.
Admins Response:
> That would be too much of a hassle. A proper CMS allows us to focus on writing.
what would be ideal is a CMS that can separate content editing and serving it. iaw a kind of static site generator that is built into a CMS and can push updates to the static site as they happen.
I mean, even just hopping over to host your site in Wordpress.com was a viable option if you were in that middle ground between personal blog and having a dedicated server admin to handle your traffic
Hard to believe that you’d be in the business of serving content in 2024 and have to deal with the slashdot effect from 1999 for your blog of articles and images.
You can take most WordPress websites from multi-second load times to 750ms or less (in fact as a regular exercise I set up fresh WordPress installs on dirt-cheap VPS hosts and see how low I can get them while still having a good site. 250ms to display is not uncommon even without CDNs)
Not every site configuration is perfect, and blaming the site's configuration, and ignoring Mastodon's inherent issue is borderline not practical.
Hard to be sympathetic.
This isn't a new issue, nor is it unique to mastodon. Reducing server load for sites like this is a very common exercise for many reasons.
There have been other cases, such as where a mobile app developer has hardcoded an image from someone else's website into their app, then millions of users request it every time they open the app. Or where middlebox manufacturers have hardcoded an IP address to use for NTP.
Sure, having efficient and well-cached infrastructure setup is good, but there's only so much you can do to "reduce server load" where other people in control of widely-deployed software have made choices that causes millions of devices around the world to hammer _you_ specifically.
The people who made those choices don't give a shit, it's not _their_ infrastructure they fucked over. That's why you need to shame them into fixing their botnet and/or block their botnet with extreme prejudice.
Mastodon's link preview service is a botnet, Mastodon knows it, and they refuse to fix it.
Years ago, I was looking into some very popular c++ library and wanted to download the archive (a tar.gz or .zip, can't remember). At that time, they hosted it on sourceforge for download.
I was looking for a checksum (md5, sha1 or sha256) and found a mail in their mailing list archive where someone asked for providing said checksums on their page.
The answer? It's too complicated to put creating checksums into the process and sourceforge is safe enough. (paraphrased, but that was the gist of the answer)
That said, since quite some years they provide checksums for their source archives, but they kinda lost me with that answers years ago.
This got me wondering, and the reason is that they embed a "card" for a link to a similar blog post on michaelmeyer.com (and grey.webp is the imagine in the card). There's a little irony there, I think.
Don't make every Mastodon instance have to fetch the linked page and all its assets to generate its own previews
EDIT: as linked in TFA, it has been nearly 7 years and they're still arguing about it:
* https://github.com/mastodon/mastodon/issues/4486
* https://github.com/mastodon/mastodon/issues/23662
Dear nincompoops: if you trust the original poster and original server to send you text and images in a toot, and federated instances to pass that around without modification... then you can trust them equally to send you a URL and an image preview. It's arrogance and idiocy that lead you to believe you can trust their images but can't trust their web preview images and you have to verify that yourself by having the Fediverse DDoS the host. This problem will only get worse as the Fediverse expands. Fix it now, don't ignore it because it makes a problem for someone else
I don't think it's that straightforward: normally if I follow @amiga386@mastodon.example you and I both need to trust example.com to accurately report what you say. But if you put a link in to news.example and mastodon.example scrapes and includes a preview I now need to trust mastodon.example to accurately report what news.example is saying. And I might well not!
I got into this more with mockups here: https://www.jefftk.com/p/mastodons-dubious-crawler-exemption
While people have come to rely on centralised services' link preview generators as some kind of trustworthy source, this shouldn't be taken for granted, and Mastodon users definitely should not give that level of trust.
Even on centralised platforms, I've seen endless "screenshots" on Twitter of other web pages and other tweets, aping the form of screenshot-quotations, but actually doctored. The original page or tweet never said what the screenshot claims they did. And I've also seen just the text of tweets claiming that some person said X, when that person did not say X. There can be millions of people affected by this misrepresentation, because they use Twitter-only sources for their information, and don't verify what they see. Then there's the same problem one level up, the screenshot might be a valid screenshot of absolute poppycock published by a partisan news source.
This is why I phrased my original statement the way I did... if you trust the original poster and original server. It's quite possible you shouldn't. Trust should be anchored to the individual toot and its poster, and you should distrust them for misrepresenting link previews in the same way you should distrust them for any other misrepresentations they make.
Ultimately, you should not trust website preview links any more than you trust the person posting them. You should click through (if you even trust opening the link) to see if the preview matches, and you should use existing tools (blocking, defederation) for posters or servers who abuse your trust.
(Edit: and let's also add in the trust that the website represents itself equally to both you and the link preview generator code... a large number of sites become much more responsive and ad-free automatically if you claim to be GoogleBot.. and let's not even mention sites doing A/B testing for virality under the same URL)
Why this is a difficult problem to resolve is that the Mastodon wants one thing - "trustworthy" per-server link previews, meaning tens of thousands of servers come thundering around the same time, and this will be millions of servers in future if they don't fix this - at the expense of others, the link targets. Meanwhile, those affected by the Mastodon community's selfish behaviour want them to clean up their act, at the cost of something the Mastodon community thinks is precious and fears losing (trusting link previews).
I think solutions need to come from behavioural change, which is to say giving link previews no more trust than the other text or images on a toot, and from that direction it would be much more palatable to have the poster supply the preview, because viewers wouldn't be giving it undue trust.
That's a different problem: we're talking here about the equivalent of Twitter/FB/etc saying "this is the image and preview text at this link". Which you can trust the traditional social media platforms for.
> Trust should be anchored to the individual toot and its poster, and you should distrust them for misrepresenting link previews in the same way you should distrust them for any other misrepresentations they make.
Note that with Mastodon this could also be caused by their server admins.
Even here, you can't trust this.
You can vaguely trust that the centralised provider won't modify the link preview, they don't appear have been caught doing that despite the fact they totally could... and possibly do, possibly only to specific people, possibly under duress from governments where they operate their servers.
However, it's been shown several times that Facebook won't let you post certain links, including links that are merely uncovering Meta's wrongdoing. So centralised providers still have the power to distort what users see, they just do it more directly and openly than misrepresenting link previews.
What Mastodon users want is to pretend they can have that same level of trust as they could with a centralised provider - which they know they can't, but being honest about that hurts adoption, so they go along with a lie, and instead make their link targets pay all the costs, to make themselves look better.
That's not right, and if their purpose is to make the internet a better place, they should be willing to compromise, be willing to prioritise being a good neighbour to other web services (e.g. add a disclaimer like "poster supplied this link preview image"), over having millions of link preview services DDoS a target so they can say "we have per-server trust in link previews".
I would also consider to what degree such a system ends up looking like a half-baked, distributed Cloudflare anyways. Like yes, I'm sure we could build some kind of incredibly complicated, reputation-based, distributed link preview caching system. Or the host could just fix their Cloudflare (or accept dying under traffic load).
My generalized experience has been that untrusted, secure, distributed systems are incredibly difficult to build, and that it's probably not worth doing for something as trivial as URL previews. Just let the request die and swap the preview for a message like "This site was down as of $lastTimeItCheckedForAPreview". Maybe change the message so it shows the URL but doesn't make it an <a> element so people can't just click on it to discourage sending them further traffic.
Or worst case, they could try a fallback to reliable, centralized sources for that info. See if Google Search has a cached copy, or archive.org, or whatever else. It's not decentralized, but I also think it should be fine to use non-decentralized features as a fallback for optional features. I've got more confidence that archive.org hasn't tinkered with their version than some random Mastodon instance.
Hmm, yeah. I was thinking about reducing the amount of data transferred but it's the actual number of requests that's the problem then it won't help.
But that still wont get you the independent ping, so then the clients immediate server should confirm.
As long as client and server confirm, that should be enough.
Perhaps throw in one more confirmation gateway and, thats gotta be enough trust before you just paranoid
They must've spelled it wrong...
I had prepared the content for potential virality (hand-written HTML & well-optimised images) but it was still an unwelcome surprise when I checked the server logs and saw all that noise.
I have a low traffic news site; I wonder If I should share this link or would they prefer not to be troubled by the traffic.
If you don't allow caching, a CDN can't help you.
The defaults are really quite broken.
Never thought about this before. One of my sites is (a single-page application) 45k HTML and 40k image. Hence about 50% of what is served for previews is wasted.
It would be nice if there was a way to recognise an "only html head needed" type of request. Don't think there is?
https://developer.mozilla.org/en-US/docs/Web/HTTP/Methods/HE...
There is this quote:
> ” making for a traffic amplification of 36704:1”
Does it mean that posting a link on Mastodon generates 36704 requests to that URL?
But for popular accounts (15k followers as in this case) will definitely spread the link to thousands of instances instantly and cause thousands of individual GET requests.
It’s fun to tail access.log and see it happen in real time.
No, it's a ratio of bytes for traffic . They're counting size of 1 request to post to Mastodon "a single roughly ~3KB POST", vs the total size of content served from GETing the url in that post.
IDK how valid this metric is, but that's what they're saying.
I'm not saying that this is the optimal framing, but that's what they were going for: talking about this, correctly or not, as a de facto DDOS.
1) https://www.microsoft.com/en-us/security/blog/2022/05/23/ana...
A link preview with an image is over 100 MB? That sounds insane. And if they mean the total traffic in 5 minutes was 100MB that cannot possibly be bringing Cloudflare to its knees. That is an indictment of Cloudflare’s CDN then!
All the issues sound more like a them problem than mastadon or cloudflare.
Something's wrong with their setup. Hard to tell as now HN has brought them to their knees
Agreed, and hardly surprising is it that if the site can't cope with link preview traffic that non-trivial page views would be trouble too!
ActivityPub itself could more intelligently cache or display the content but something doesn’t add up if Cloudflare can’t handle that kind of traffic let alone Hacker News traffic.
From the article:
> Presently, we use Cloudflare as our CDN or WAF, as it is a widely adopted solution.
To me, it sounds like the author isn't really familiar with the difference between a CDN and WAF is; or familiar with Cloudflare beyond it being a popular thing they should probably have for that matter.
The images arn't the problem, it's their dynamicly generated page, even if the content is pretty static.
But I agree this does seem strange, like it shouldn't be unmanageable load at all.
What does getting a certificate through let's encrypt have to do with the server getting overwhelemed?
It's a thing that happens once to renew a cert every couple months.
The performance hit of https on a modern server is negligable. The performance hit on a watch is negligable.
Cloudflare is either passing traffic through without touching it, or as a proxy is doing tls termination, as they're a trusted CA in most devices/browsers/OSs/etc.
None of this really has anything to do with that's happening with OP.
I think what’s being implied here is that, when you share a link to Facebook, Facebook will access the page to generate a link preview, so will download a tiny bit of HTML and an image. But when you share a link on mastodon, that link immediately gets propagated to many other mastodon servers, which then propagate it to others, so suddenly many thousands of mastodon instances are simultaneously downloading a little bit of HTML and an image, and the cumulative effect of that in this instance was 100MB over a minute or two.
It does seem like a typical static website ought to not have a problem serving that, especially if it’s behind Cloudflare. It seems odd that a single EC2 instance would have a hard time serving that.
But given more than one person is complaining about, it also seems like each mastodon instance could very easily delay propagation of the story by a few minutes to soften the blow here.
I liked that idea at first glance, but thinking about it, CDN performance would actually be better with a single huge burst than if they were smeared out (assuming a very short max-age so the site can be updated rapidly).
In theory a CDN could optimistically coalesce requests then re-send them when the headers of the first one return. But this is very complex and rarely done in practice.
This can also occur on any time the cache gets stale and needs to be refetched.
I don't think this is true. It certainly isn't for any CDN that I've worked for or on.
Cloudflare don't do this either - they use a cache lock - the first request basically acts as a blocker for all the others, leaving the other requests waiting for the response (if it's cacheable they serve that response, if not then they proceed to origin).
It's normally configurable, but most sane CDNs do have it enabled by default, precisely because big bursts tend to be sharp in nature and a cache miss can be origin breaking at that point.
Just for completeness's sake, Nginx's HTTP proxy module can do it too (the setting's proxy_cache_lock) though it is off by default there.