MangaDex infrastructure overview
mangadex.dev
mangadex.dev
I have loads and loads of thoughts about what they could be doing differently to reduce their costs but I'll just say that the number one thing Mangadex could be doing right now from a cursory glance is to reduce the number of requests. A fresh load of the home page generates over 100 requests. (Mostly images, then javascript fragments.) Mangadex apparently gets over 100 million requests per day. My site - despite ostensibly having more traffic - gets fewer than half that many in a month. (And yes, it's image-heavy.)
A couple easy wins would be to reduce the number of images loaded on the front page. (Does the "Seasonal" slider really need 30 images, or would 5 and a link to the "seasonal" page be enough? Same thing with "Recently Added" and the numbers of images on pages in general.) The biggest win would probably be reducing the number of javascript requests. Somehow people seem to think there's some merit to loading javascript which subsequently loads additional javascript. This adds a tremendous amount of latency to your page load and generates needless network traffic. Each request has a tremendous amount of overhead - particularly for dynamically-generated javascript. It's much better to load all of the javascript you need in a single request or a small handful of requests. Unfortunately, this is probably a huge lift for a site already designed in this way, but the improved loading time would be a big UX win.
Anyway - best of luck to MangaDex! They've clearly put a lot of thought into this.
In general the less moving parts you have in a system the more reliable, secure, efficient and cheaper the system becomes.
In their case they run a site that is probably under constant attack by the "hired goons", so they're going to need to have more moving parts than others. Plus they will want to optimise for minimal development time (it's a hobby) so just adding another tried and trusted system into the stack to do something you need makes sense.
That's taken care of by the DDoS-Guard system they placed fronting their infrastructure. The design of their system has to take this into account, but that is mainly on a IP and DNS level. The design of their stack behind the loadbalancer is mainly driven by their functional and non-functional requirements, rather than by the need to prevent DDoS attacks.
> In general the less moving parts you have in a system the more reliable, secure, efficient and cheaper the system becomes.
100% agreed. This is not my first high-traffic site, nor even the highest. (I built the analytics system for a an Alexa top-10 site in 2010, reaching some 30 billion writes / day off of a mere 14 small ec2 instances.) I've never seen a k8s implementation in production that was necessary.
I will note that my Alexa-2k site is also a personal site (no revenue) and under constant attack. In fact we frequently suffer DDOSes that we don't even notice until reviewing the logs later because it doesn't suffer any latency under pressure.
I've spent much of my career working on systems with active users from the hundreds to low thousands, but which process a huge number (50k/sec scale) jobs/tasks.
It's a totally different kettle of fish, and if I'm totally honest I'm shocked at how badly "web" scales and how common these naive and super inefficient implementations are (hint: my bare-metal server from 2005 was faster than expensive cloud VMs).
Recently I've worked on two high-usage systems (one of which was "handling" 30k requests/second for the first couple of week).
MMO games, by any chance?
Really interested to see how you think about this sort of thing =)...
Perhaps mods have the ability to extend this for active discussions and that's why I can reply now?
(2) You can spend weeks building complex infrastructure or caching systems only to find out that some fixed C in your equation was larger than your overhead savings. In other words: Measure everything. In other other words: Premature optimization is the root of all evil.
(3) Fewer moving parts equals less overhead. (Again: Simple beats complex.) It also makes things simpler to reason about. If you can get by without the fancy frameworks, VMs, containers, ORM, message queues, etc. you'll probably have a more performant system. You need to understand what each of those things does and how and why you're using them. Which brings me to:
(4) Learn your tools. You can push an incredible amount of performance out of MySQL, for instance, if you learn to adjust its settings, benchmark different DB engines for your application, test different approaches to building your schemas, test different queries, make use of tools like the EXPLAIN statement, etc. you'll probably never need to do something silly like make half a dozen round-trips to the database in a single page load.
(5) Understand your data. Reason about the data you will need before you build your application. If you're working with an existing application, make sure you are very familiar with your application's database schema. Reason ahead of time about what requirements you have or will have, and which data will be needed simultaneously for different operations. Design your database tables in such a way as to minimize the number of round-trips you will need to make to the database. (My rule of thumb: Try to do everything in a single request per page, if possible. Two is acceptable. Three is the maximum. If I need to make more than three round-trips to the database in a single page request, I'm either doing something too complex or I seriously need to rethink my schema.)
(6) Networking is slow. Minimize network traversal. Avoid relying on third-party APIs where possible when performance counts. Prefer running small databases local to the web server to large databases that require network traversal to reach. This is how I handled 30 billion writes / day: 12 web servers with separate MySQL instances local to each sharded on primary key IDs. The servers continuously exported data to an "aggregation" server, which was subsequently copied to another server for additional processing. Having the web server and database local to the same VM meant they didn't need to wait for any network traversal to record their data. I could've easily needed several times as many servers if I had gone with a traditional cluster due to the additional latency. When you need to process 25,000 events in a second, every millisecond counts.
(7) Static files beat the hell out of databases for read-only performance. (Generally.)
(8) Sometimes you can get things moving even faster by storing it in memory instead of on disk.
(9) Reiterating what's in (3): Most web frameworks are garbage when it comes to performance. If your framework isn't in the top half of the Techempower benchmarks, (or higher for performance-critical applications) it's probably going to be better for performance to write your own code if you understand what you're doing. Link for reference: https://www.techempower.com/benchmarks/ Note that the Techempower benchmarks themselves can be misleading. Many of the best performers are only there because of some built-in caching, obscure language hack, or standards-breaking corner-cutting. But for the frameworks that aren't doing those things, the benchmark is solid. Again, make sure you know your tools and why the benchmark rating is what it is. Note also that some entire languages don't really show up in the top half of techempower benchmarks. Take that into consideration if performance is critical to your application.
(10) Most applications don't need great performance. Remember that a million hits a day is really just 12 hits per second. Of course the reality is that the traffic doesn't come in evenly across every second of the day, but the point remains: Most applications just don't need that much optimization. Just stick with (1) and (2) if you're not serving a hundred million hits per day and you'll be fine.
I've not really ever applied 9 myself, I've run comparative benchmarks a couple of times, but not thought about using that as a basis for whether to roll my own on critical performance parts.
Any example of such frameworks?
Took me almost a decade to really comprehend this.
I used to include all sorts of libraries, try out all the fancy patterns/architectures etc...
After countless of hours debugging production issues... the best code i've ever written is the one with the fewer moving parts. Easier to debug and the issues are predictable.
I know I’m not the first to use that phrasing, but I’m not sure where I picked it up. If someone wants to point out the etymology of that type of phrase, I’d be glad to read up on what I’ve forgotten/missed.
In the very first lecture of the Computer Science degree I did in the 1980s the lecturer emphasised KISS, and said that while we almost certainly wouldn't believe it at first eventually we'd realise that this is the most important design principle of all. Probably took me ~15 years... ;-)
I don't understand this part. Hopefully you can clarify this to me.
If you're sharding by primary key, doesn't that mean that there's a high chance that the shard in your local DB instance won't have the data the web server is requesting?
I'm not familiar with DB management.
In that case, you can easily split the data between shards based on ranges of an integer key. It's very easy to code, test, deploy and understand such a design.
- ignores the vast majority of "image serving" (most is handled by DDG and our custom CDN)
- the JS fragments thankfully should load only on first visit and then get aggressively cached by DDG/your browser
One of the pain points is that there are a lot of settings for users to decide what they should or shouldn't see (content rating, original language of origin, search tags, etc) and some are already specifically denormarlized (when querying chapter entities, ES indices for those contain some manga-level properties to avoid needing to dereference that first too) -- however this also makes caching substantially less efficient in many places, alas
Thanks!
The first thing that I noticed is that even with caching enabled, you're loading "too much data". After loading the main page and then clicking one of the tiles, there are several JSON API calls.
Here's an example, 195 kB transferred (528 kB size): https://api.mangadex.org/manga/bbaa17c4-0f36-4bbb-9861-34fc8...
Oof. Half a megabyte of JSON! Ignore the network traffic for a moment, because GZIP does wonders. The real problem is that generating that much JSON is very "heavy" on servers. Lots and lots of small object allocations, which gives the garbage collector a ton of work to do. It's also expensive to decode on the browser for similar reasons.
On my computer, this took a whopping 455ms to transfer, nearly half a second. That results in a noticeable latency hit to the site.
In my consulting gig I always give developers the same advice: "Displaying 1 kilobyte of data should take roughly 1 kilobyte of traffic".
In other words, there's isn't 500 KB of text anywhere on that page! A quick cut & paste shows about 8 KB of user-visible text in the final HTML rendering. That's a 1:60 ratio of content-to-data, which is very poor. I bet that behind the scenes, this took a heck of a lot more back-end network traffic and in-memory processing to generate. Probably tens to hundreds of megabytes of internal traffic, all up.
This is one of the core reasons most sites have difficulty scaling, because for every kilobyte of content output to the screen, they're powering through megabytes or even gigabytes of data behind the scenes.
Can this API query be cut down to match what's displayed on the screen? Can it be cached for all users? Can it be cached precompressed?
Etc...
For what it's worth, this isn't generated live but a mix of existing entity documents
Most of it is page filenames which indeed could be made optional and fetched only by the reader, but that'd be us actively nulling them out in the returned entity, since they are there in the ES documents for the chapters (a manga feed like this being a list of chapters)
For example, user role memberships:
{
"id": "c80b68c5-09ae-4a50-a447-df7c5a4a6d01",
"type": "user",
"attributes": {
"username": "kinshiki",
"roles": [
"ROLE_MEMBER",
"ROLE_GROUP_MEMBER",
"ROLE_POWER_UPLOADER"
],
"version": 1
}
}
Also record timestamp dates like created/changed, along with contact details that may be revealing sensitive info: "attributes": {
"name": "SENPAI TEAM",
"locked": true,
"website": "https:\/\/discord.gg\/84e3j9b",
"ircServer": null,
"ircChannel": null,
"discord": "84e3j9b",
"contactEmail": "senpai.info@gmail.com",
"description": null,
"official": false,
"verified": false,
"createdAt": "2021-04-19T21:45:59+00:00",
"updatedAt": "2021-04-19T21:45:59+00:00",
"version": 1
}
But let's just go back to your response:> Most of it is page filenames which indeed could be made optional
Do that! If you strip them out, the 529 kB document shrinks to 280 kB, which hardly seems worth the hassle, but when gzipped, this is a miniscule 13 kB! This is because those strings are hashes, which significantly reduces their compressibility compared to general JSON, which usually compresses very well.
It's basic stuff like this that can make a website absolutely fly.
Avoid giving computers unnecessary, mandatory work: https://blog.jooq.org/many-sql-performance-problems-stem-fro...
Because of this model, we also make sure that Elasticsearch merely works a search cache, not as an authoritative content database (hence everything we add in there is considered public, on purpose, and what isn't meant to be public is just not indexed in ES)
However the gzip efficiency improvements would be really neat for sure
Fwiw I also don't work on the backend and there might be good reasons to not expressly filter out data (yet anyway, perhaps it will end up as a separate entity and be a include parameter)
Edit: As you said, there may be reasons on the backend not to filter things out of the query. Though it seems likely that the web response could be trimmed down.
Someone who is focused on the performance aspect & someone who is focused on stack stability discussing the real world input & output of a business system and showing why performance & UX are not the only metrics that matter is a good thing for us to see.
> Can this API query be cut down to match what's displayed on the screen? Can it be cached for all users? Can it be cached precompressed?
This is why you want to bypass the JS realm, (or whatever language does the serdes) and send clients JSON or XML directly from the database, so the client is only getting the data at rest.
Is this to be taken literally? I don't consider myself a performance-tuning expert, but I'm not sure how can I make something useful out of this advice. Of course, "the less you transfer, the better" is an obvious thing to say (a bit too obvious to be useful, in fact), but does it really mean I should aspire to transfer only what I'm actually going to display right now? For example, there is a city autocomplete form on the page (well, a couple of thousand relatively short entries). In that case I would probably consider making 1 request to fetch all these cities (on input focus, most likely), instead of making a request to the server on every couple of characters you type. Is it actually a wrong way of thinking?
In your case, you're optimising for round-trips, which is also important. As long as you only send the city names instead of a huge blob that also includes a bunch of metadata, you're probably fine.
The most common example of my rule is that I often see SELECT statements on unindexed columns. This means that behind the scenes, the database engine is forced to do a table scan to find the row. If the query uses a wildcard selector, then it is also forced to return all columns, whether they are used by the application or not.
I commonly see scans over 100 MB tables returning 100 KB to the web tier, which then converts this to 200 KB of JSON to show 100 bytes of text to the end user. Simply adding an index to the table allows the database engine to reduce the data it has to process to 10-30 KB. Selecting specific columns can reduce that to a few kilobytes, and likely also shrink the JSON to match. Eliminating the JSON and directly generating the HTML on the server like in the good old days would cut the Internet network traffic down to minimum 100 bytes required also.
Similarly, you often see performance monitoring, logging, or graphing programs store data in fantastic detail and precision. Meanwhile, the graph needs only 16 bits of data, because screens are typically at most a few thousand pixels across in size! A case in point is Microsoft System Center Operations Manager (SCOM), which has a metric write amplification of something like 300:1, which is why it can't log metrics at a usefully high frequency. Not because that's impossible, but because it's wasting the available computer power to an absurd degree. Azure has inherited this code, and then layered JSON on top. (I guess when you bill by gigabytes ingested, the incentives are all wrong.)
And if your JS assets are hashed then you can add cache-control: immutable so that a browser doesn't have to reload them when the user F5s.
According to Alexa you have a 46.4% bounce rate. [1]
When 46% of your users aren't coming back, how does 31 round-trips to your server for 100% of first-page visitors save anyone time or bandwidth? Your pageviews per visitor is 6.8, meaning the 53.6% that stick around view an average of 11.8 pages each. Even if there are zero subsequent js requests on other pages (clicking a random page I see 8) you would be generating 31 requests up-front to save 10.8 subsequent requests for about half of your users. (And again - in any scenario where the number of js fragments transferred on subsequent requests >= 1 even this benefit goes out the window.) How does that save you or your users bandwidth, server load, or other overhead?
The scale is not quite linear, but generally speaking, if you get your number of requests down from > 100 to < 5, you'll be able to handle around 20x the traffic with the same number of web-facing servers. Or alternatively the same amount of traffic with around 1 / 20th the servers.
Would that have a material effect on your costs?
However the serving of this JS has nearly no cost to us (as they are cached at the edge by DDoS-Guard and the frontend is otherwise entirely static on our end)
I see 17 requests, all over either h2 or h3. 4 of them JS, and 2 images.
Nope. Different problem.
The article was linked to a page under the domain "mangadex.dev".
Without any other context, I had assumed "home page" meant http://mangadex.dev , or what I got when clicking "Home" on the linked article.
Apparently not.
_How_ you hit scale on a budget is one part of the equation. The other part is: what you're doing.
Off the top of my head, the "how" will often involve the following (just to list a few):
1 - Baremetal
2 - Cache
3 - Denormalize
4 - Append-only
5 - Shard
6 - Performance focused clients/api
7 - Async / background everything
These strategies work _really_ well for catalog-type systems: amazon.com, wiki, shopify, spotify, stackoverflow. The list is virtually endless.
But it doesn't take much more complexity for it to become more difficult/expensive.
Twitter's a good example. Forget twitter-scale, just imagine you've outgrown what 1 single DB server can do, how do you scale? You can't shard on the `author_id` because the hot path isn't "get all my tweets", the hot path is "get all the tweets of the people I follow". If you shard on `author_id`, you now need to visit N shards. To optimize the hot path, you need to duplicate tweets into each "recipient" shard so that you can do: "select tweet from tweets where recipient_id = $1 order by created desc limit 50". But this duplication is never going to be cheap (to compute or store).
(At twitter's scale, even though it's a simple graph, you have the case of people with millions of followers which probably need special handling. I assume this involves a server-side merge of "tweets from normal people" & RAM["tweets from the popular people"].)
Also, put the browser to work for you, caching via Cache-Control, ETag, etc. Only then, optimize your server...
Mike Cvet's talk about Twitter's fan-in/fan-out problem and its solution makes for a fascinating watch: https://www.youtube-nocookie.com/embed/WEgCjwyXvwc
Learned something new today.
If something is read much more frequently than it changes, store it client-side, or store it temporarily in an in-memory-only, not-persisted-to-disk "persistence" layer like Redis.
For example, if you're running an online store, your product list doesn't change all that often, but it's queried constantly. The single source of truth lives in a relational database, but when your app needs to fetch the list of products, it should first check the caching layer to see if it's available there. If not, fetch it from the database, but then write it into the cache so that it's available more quickly the next time you need it.
> When and how to denormalize, why is it needed?
When you need to join several tables together in order to retrieve a result set, and especially when you need to do grouping to get the result set, and the retrieval & grouping is presenting a performance problem, then pre-bake that data on a regular basis, flattening it out into a table optimized for read performance.
Again with the online store example, let's say you want to show the 10 most popular products, with the average review score for each product. As your store grows and you have millions of reviews, you don't really want to calculate that data every time the web page renders. You would build a simpler table that just has the top 10 products, names, IDs, average rating, etc. Rendering the page becomes much more simple because you can just fetch that list from the table. If the average review counts are slightly out of date by a day or two, it doesn't really matter.
> Why append-only and how?
If you have a lot of users fighting over the same row, trying to update it, you can run into blocking problems. Consider just storing new versions of rows.
But now we're starting to get into the much more challenging things that require big application code changes - that's why the grandparent post listed 'em in this order. If you do the first two things I cover above there, you can go a long, long, long way.
Consider something like an amazon product page. It's mostly static. You can cache the "product", and calculate most of the "dynamic" parts in the background periodically (e.g., recommendation, suggestions) and serve it up as static content. For the truly dynamic/personalized parts (e.g., previous purchased) you can load this separately (either as a separate call from the client or let the server pieces all the parts together for the client). This personalized stuff is user specific, so [very naively]:
conn = connections[hash(user_id) % number_of_db_servers]
conn.row("select last_bought from user_purchases where user_id = $1 and product_id = $2", user_id, product_id)
Note that this is also a denormalization compared to:select max(o.purchase_date) from order o join order_items oi on o.id = oi.order_id where o.user_id = $1 and oi.product_id = $2
Anyways, I'd start with #7. I'd add RabbitMQ into your stack and start using it as a job queue (e.g. send forget password). Then I'd expand it to track changes in your data: write to "v1.user.create" with the user object in the payload (or just user id, both approaches are popular) when a user is created. It should let you decouple some of the logic you might have that's being executed sequentially on the http request, making it easier to test, change and expand. Though it does add a lot of operational complexity and stuff that can go wrong, so I wouldn't do it unless you need it or want to play with it. If nothing else, you'll get more comfortable with at-least-once, idempotency and poison messages, which are pretty important concepts. (to make the write to the DB transactionally safe with the write to the queue, lookup "transactional outbox pattern").
Sharding sucks, but if your database can't fit on a single machine anymore, you do what you've got to do. The basic idea is instead of everything in one database on one machine (or well redundant group of machines anyway), you have some method to decide for a given key what database machine will have the data. Managing the split of data across different machines is, of course, tricky in practice; especially if you need to change the distribution in the future.
OTOH, Supermicro sells dual processor servers that go up to 8 TB of ram now; you can fit a lot of database in 8 TB of ram, and if you don't keep the whole thing in ram, you can index a ton of data with 8 TB of ram, which means sharding can wait. In contrast, eBay had to shard because a Sun e10k, where they ran Oracle, could only go to 64 GB of ram, and they had no choice but to break up into multiple databases.
Super simple example, splitting there phone book into two volumes, A-K and L-Z. (Hmmmm, is a "phonebook" a thing that typical HN readers remember?)
> you can fit a lot of database in 8 TB of ram, and if you don't keep the whole thing in ram, you can index a ton of data with 8 TB of ram, which means sharding can wait.
For almost everyone, sharing can wait until after the business doesn't need it any more. FAANG need to shard. Maybe a few thousand other companies need to shard. I suspect way way more businesses start sharding when realistically spending more on suitable hardware would easily cover the next two orders of magnitude of growth.
One of these boxes maxed out will give you a few TB of ram, 24 cpu cores, and 24x16TB NVMe drives which gives you 380-ish TB of fairly fast database - for around $135k, and you'd want two for redundancy. So maybe 12 months worth of a senior engineer's time.
https://www.broadberry.com/performance-storage-servers/cyber...
In America. When the salaries are 2/3 times lower, people spend more time to use less hardware.
There's also a price breakpoint for single socket vs dual socket. Or four vs two, if you really want to spend money. My feeling is currently, single socket Epyc looks nice if you don't use a ton of ram, but dual socket is still decently affordable if you need more cores or more ram and probably for Intel sevees; quad socket adds a lot of expense and probably isn't worth it.
Of course, if time is cheap and hardware isn't, you can spend more time on reducing data size, profiling to find optimizations, etc.
On the other hand, directly because of the above, their hasty self-inflicted take down earlier this year nearly killed the entire hobby. Many series essentially stopped updating for the ~5 months the site was down, and many more are likely never coming back again.
The decision to suddenly take the site down for a full site rewrite feels completely inexplicable from the outside. (A writeup the above or the previous one[1], both of which read like they were written by a Google Product Manager, especially don't help as they conspicuously avoid any comment to the one question on everyone's mind: "leaving aside the supposed security issues with the backend, why on earth also rewrite and redesign the entire front end from scratch at the same time?")
And while it works fine for reading, it kills any interaction with the hosting sites. No chance for monetization, socialization or anything else that can help sites survive long-term.
So yeah, considering how fragile maintaining a site like this is, it is always a good idea to sync your progress in a third party so it is easier to migrate if something goes wrong.
> And while it works fine for reading, it kills any interaction with the hosting sites. No chance for monetization, socialization or anything else that can help sites survive long-term.
BTW, MangaDex doesn't have monetization because it is strict a hobby and also because it is a gray area to monetize about this kinda of work [1]. Also, their Tachiyomi client is official (MangaDex v5 API was tested primarily via their Tachiyomi client before they finished the Web interface).
[1]: both for companies (that has the copyright from the works hosted on those sites) and the scanlators (the fans that does actual work of translating those chapters). Sites that host those chapters and monetize are pretty much monetizing on work from other people.
It won't kill the hobby. Because these scanlators are making mad money from ads, patreon, crypto mining. I'll never get why they don't get more aggressive take down notices from Chinese/Japanese/Korean publishers.
Publishers see the loss as minimal and creators see piracy as free advertising to drum up enthusiasm for anime adaptations, which actually do drum up decent profits internationally (the committee keeps the streaming licensing fees, not the animation studio).
Most manga publishers will see relatively little revenue from international anime releases. Even for domestic anime releases of the vast majority of titles, the manga publisher is only a small part of the anime production committee, and the hope is mostly that popularity of the anime can lead to increased sales of the manga, merchandise, or other events. So when the anime is released internationally, they get an even smaller cut of that because the international licensee also has to take their profit.
But other than mega-hit titles where an international anime release may also lead to significant international manga sales, the popularity of an anime adaptation overseas is practically irrelevant to the original manga publisher.
(Just as an example of a local copyright quirk that will probably confuse a lot of people in the audience from Europe: copyright registration. America really, really wants you to register your copyright, even though they signed onto Berne/WTO/TRIPS which was supposed to abolish that regime entirely. As a result, America did the bare minimum of compliance. You don't lose your copyright if you don't register, but you can't sue until you do, and if you register after your work was infringed, you don't get statutory damages... which means your costs go way up.)
Furthermore, every enforcement action you take risks PR backlash. The whole fandom surrounding import Japanese comic books basically grew out of a piracy scene. Originally, there were no English translations, and the scene was basically reusing what we'd now call "orphan works". There used to be an unspoken rule among most fansubbers of not translating material that was licensed in the US. All that's changed; most everything gets licensed and many fan translators absolutely are stepping on the toes of licensees. However, every time a licensee or licensor actually takes an enforcement action, they get huge amounts of blowback from their own fans.
It’s seems weird to attribute Mangadex taking their site down for valid security concerns to the end of scandalization of certain series. That seems like entirely a Scan team problem if they decide not to upload via Cubari like other teams have done. And it doesn’t even matter since a series can get sniped at anytime.
It’s makes entire sense that if you’re going to rewrite the backend and API from scratch , you might as well do the front end too since it was a Goal from the beginning.
Try this yourself: write a simple web server in Go, host it on a cheap VPS provider, let's say at the option that costs $20/mo. Your website will be able to handle more than 1k/s requests with hardly any resource usage.
ok, let's assume you're doing some complicated things.
So what? You can scale vertically, upgrade to the $120/mo server. Your website now should be able to comfortably handle 5k req/s
Looking at the website itself, mangadex.org, it doesn't even host the manga itself. The whole website is just an index that links to manga on external websites. All you are doing is storing metadata and displaying it as a webpage. The easiest problem on the web.
So, I really don't understand the premise behind the whole post.
The problem statement is:
> In practice, we currently see peaks of above 2000 requests every single second during prime time.
This is great in terms of success as a website, but it's underwhelming in terms of describing a technical problem.
They do seem to. Clicking on a random manga on there the images are hosted on their server[0]. Also I guess some of those are much bigger images which is less trivial to serve at that rate than a 10kb static page.
0. blob:https://mangadex.org/e78bd61a-e761-4a73-a27c-5f58394e7ea4
A bit of an intro punchline, even though I agree it admittedly doesn't say much on itself :)
Fwiw most of the work is that there's little "static" traffic going on -- images and cacheable responses are not very CPU intensive to serve -- but what isn't static (which is a good chunk of it) is more problematic, but more to come on these
Where you can use 1 server, they will need to have something around 20 servers. Where you can use a cheap VPS provider, they must use an expensive shady provider who will take the heat of legal attacks. And so on and on... because of their situation they have a bunch more requirements which eat their budget than your average website, leading to a rather heavy, complex and thus expensive architecture.
Surely there is still room for optimization, but it seems this is a rather new redesign from scratch(?), so not details need time.
Their goal is for scanlators to have a place to post their new translated manga, rather than always linking it off from some Wordpress instance.
These people have never heard of Go, obviously. The likely scenario is not that you haven't fully understood their constraints or requirements, it's that you're just smarter than they are.
> So what? You can scale vertically, upgrade to the $120/mo server. Your website now should be able to comfortably handle 5k req/s
> Looking at the website itself, mangadex.org, it doesn't even host the manga itself. The whole website is just an index that links to manga on external websites. All you are doing is storing metadata and displaying it as a webpage. The easiest problem on the web.
Take that order of magnitude cheaper, single VPS server solution you're proposing and build something with it. Sounds like you'd make a lot of money. There has to be a business idea around "storing metadata and displaying it as a webpage" somewhere? Easiest problem on the web.
The peanut gallery at HN is out of control. People who don't do / build explaining to the people who do how easy, simple, better their solutions would be.
While it's not as simple as a Go program on a VPS, there is certainly a lot of unnecessary overhead here. I think you underestimate just how much poor and wasteful engineering there is out there.
I don't under estimate poor and wasteful engineering at all, but that's not what I saw in the article.
Serving traffic is a single element of their design. They also designed for security, redundancy, and observability. All with their own solutions because using a service or a cloud provider would be too costly. With that in mind, it's not a charitable view to think they didn't explore low hanging fruits like "make the server in Go". If you think you can do better, detail in depth how and solve all of their requirements vs. the single piece you're familiar with.
And if you can do the above holistically, for an order of magnitude below their costs, it sounds like I need to get in touch to throw money at you.
> "detail in depth how "
This thing seems to be little more than a very complex API and SPA sitting on top of Elasticsearch. These frontend/backend sites are almost always a poor choice compared to a simple server-side framework that just generates pages. ES itself is probably unnecessary depending on the requirements of their search (it doesn't seem to be actual full text indexing of the content but just the metadata). The security and observability also tends to be a problem of their own making and a symptom of too much complexity.
I don't dispute this or your credentials. You've built critical systems in a space where it was a core of the business. If given time, and resources, I have no doubt you could build a custom solution to their problem that was more efficient.
Unstated in this is the type of business MangaDex is, which I have the following assumptions about. I don't think it's unfair to assume that we're mostly on the same page here:
- Small to mid size, at most
- Small engineering team. Need to develop, deploy, support, and maintain solutions.
- Lacks deep systems expertise, or is unable to attract talent that has that expertise ($)
These characteristics are very common in our space. To solve their technical problems, most of the time, they reach for an open source solution (after examining the alternatives like a service).
Now the question is given those constraints, and their other business requirements, how do they best optimize for dimensions they care about? Everything is a trade-off. Everyone who builds knows this. It's unkind to pretend this is a purely technical exercise. And after reading their article, it's obvious they know some of trade-offs they're making, so it's unkind to suggest a naive solution that does nothing but make you feel smarter. I'm not saying you did the above, but some of these comments are outrageous.
I can and do frequently advise on certain topics in comments specifically because I do build and can in fact speak of such topics authoritatively. Isn't that what this website is for?
That said, the post you are replying to is perhaps overly dismissive of the criteria that this website operates under. Other comment chains have some really good advice though.
I never claimed to be smarter. I just understand some things that I noticed a lot of people in the industry don't understand.
My understanding is not even that great.
But still, this is just one example that I keep running into over and over and over:
People opting for a complicated infrastructure setup because that's what they think you should do.
No one showed them how to make a stable reliable website that just runs on a single machine and handle thousands of concurrent connections.
It's not hard. It's just that they've never seen it and assume it's basically impossible.
There are areas about computing that I feel the same way about. For example, before Casey Muratori demoed his refterm implementation, I had no idea that it was possible to render a terminal at thousands of frames per second. I just assumed such a feat was technically impossible. Partly because no one has done it. But then he did it, and I was blown away.
> Take that order of magnitude cheaper, single VPS server solution you're proposing and build something with it. Sounds like you'd make a lot of money.
Building something and making money out of it are not the same thing. But thanks for the advice. I'm in the process of trying. I know for sure I can build the thing, but I don't know if it will make any money. We will see.
> People who don't do / build explaining to the people who do how easy, simple, better their solutions would be.
I do and have done.
This kind of advice is exactly the kind of thing I know how to do because I have done it in the past using my trivial setup of a single process running on a cheap VPS. And I have also seen other teams struggle to get some feature nearly half-working on a complicated infrastructure setup with AWS and all the other buzzwords: Kibana, Elastic Search, Dynamo DB, Ansible, Terraform, Kubernetes ... what else? I can't even keep track of all this stuff that everyone keeps talking about even though hardly anyone needs at all.
I've seen 4 or 5 companies try to build their service using this kind of setup, with the proposed advantange of "horizontal" and "auto" scaling. And you know what? They ALL struggled with poor performance, ALL THE TIME. It's really sad.
What is a reliable website? What does this website do?
If given a static constraint, like serve 2000 requests per second with 99.999% uptime, and enough time, I'm sure you can optimize it to be as efficient as you'd like. But that's not our exercise. Bespoke, custom solutions that are not the core of the business are not solutions. Repeat that a dozen times.
MangaDex's business is not to be the most efficient website possible. Their business, I assume, is content and features for their users. They pick off the shelf technologies to do it because it's well documented, proven, and most importantly already built.
They compose these technologies to solve their business problems. Often there's a mismatch or an overlap in functionality that introduces inefficiency (complexity, cost, performance), but that's a trade-off MangaDex and many other businesses make. We can judge how poorly or well they've made some trade-offs base on their business and overall requirements.
You coming here and telling people that you can run it on a VPS ignores all of the above. And Casey has a YouTube channel where he makes a game his way I would assume because working to solve uninteresting business problems (and possibly dealing with co-workers who may pull down N project instead of building it themselves) wasn't a space he was interested in.
There's a difference between a purely technical challenge, and working with complex, interacting systems like ... people and business requirements and laws and regulations and auditing and hiring and security and who the hell maintains this system when Bob quits. Conflating the two is the root of the comments like yours, I'd think.
> It's not hard. It's just that they've never seen it and assume it's basically impossible.
I'm sorry, but who thinks only serving 2000 requests per second is hard? Or do you assume they think it's hard because you misunderstood or are unaware of their 100s of other requirements that they need to solve for in addition to serving 2000 requests per second?
I'm going on and on in this thread mainly because I'm tired of people assuming they know better than the people in the trenches making these decisions. You're assuming you're more knowledgable and skilled than them in their own problem space! They're obviously unaware that using Go to make a server would solve their problems otherwise why wouldn't they have done it?
FYI I work (and have worked) for large tech companies (think silly acronym). I'm not even in this space, as in the type of problems I face are quite different, but I can respect the authors of that article enough to nod along, shrug, and not assume I know better.
It performs well and doesn't randomly go down when someone posts a link to it on HN or tries to put in bad input.
> I'm sure you can optimize it to be as efficient as you'd like.
Wrong! I've said nothing about optimizing things.
All I'm advocating is simple solutions that are proven to work.
A web server in Go is far from efficient. An optimized server in C/C++ can probably perform 20x better than a Go server. If not more.
However, a web server in Go makes far more reasonable use of system resources to achieve the desired goals. It's also pretty reliable.
> Bespoke, custom solutions that are not the core of the business are not solutions. Repeat that a dozen times.
I don't understand the point of this sentence.
Are you saying that Kubernetes or Elastic Search or AWS or any of the other buzzwords are at the core of their business?
Clearly they are not.
> MangaDex's business is not to be the most efficient website possible. Their business, I assume, is content and features for their users. They pick off the shelf technologies to do it because it's well documented, proven, and most importantly already built.
It's in the interest of their business to lower their cost of operations. Building on a complicated infrastructure when you don't need is incurring a lot of cost. Not just the monthly cost ($1500/mo) but the cost of the staff needed to understand and maintain this infrastructure.
It's not the kind of thing that is easy to maintain.
To be completely frank with you, I myself am not capable of understanding or maintaining such a system. And every company I've been almost had no one who understood how the system really works. Someone set things up sometime by following some tutorials. When things go wrong, people panic and go into fire fighting mode. They spend hours trying to make sense of what's going on, usually involving multiple people - because it's not a task that a single individual can handle.
> They compose these technologies to solve their business problems. Often there's a mismatch or an overlap in functionality that introduces inefficiency (complexity, cost, performance), but that's a trade-off MangaDex and many other businesses make. We can judge how poorly or well they've made some trade-offs base on their business and overall requirements.
You are talking as if these off the shelf technologies are reliable and easy to implement or integrate.
From what I've seen, these solutions are a lot more complicated than what I'm proposing.
Every place I've been to that tries to take this approach ends up burning too much money and resources trying to make their thing work.
It's not as if these companies don't have to write code to make their product work. You still have to write code anyway. So, why not, instead of writing tons of glue code and configuration files to hopelessly integrate a hodge podge of tools and frameworks ... why not just write the simple code that just does the thing you want?
> I'm going on and on in this thread mainly because I'm tired of people assuming they know better than the people in the trenches making these decisions. You're assuming you're more knowledgable and skilled than them in their own problem space! They're obviously unaware that using Go to make a server would solve their problems otherwise why wouldn't they have done it?
The first company I've been to that was doing this kind of thing was spending upwards of $10k/mo on the most beefed up server that AWS provides to host the database server, and they still struggled to server more than 1000 users concurrently.
According to you, I'm not in a position to give them suggestions or adivce about how to fix this problem!!
> I'm sorry, but who thinks only serving 2000 requests per second is hard? Or do you assume they think it's hard because you misunderstood or are unaware of their 100s of other requirements that they need to solve for in addition to serving 2000 requests per second?
What are the other 100 requirements that are not fulfilled by the thing I'm proposing?!
They're using battle-tested tech from Redis and RabbitMQ to Ansible and Grafana. Nothing super fancy, nothing used just for the sake of being modern. Not sure how long it took them to end up with this architecture but it doesn't look like a new dev would have a hard time getting familiar with how everything works.
Would definitely like to hear more about their dev environment, how it is different from prod, and how they handle the differences.
But as a craftsman, it is definitely nice :)
It's honestly quite boringly similar (hence why it's only vaguely alluded to in the article)
Take out DDoS-Guard/External LBs (no need for publicness of it), pick a cheap-o cloud provider to get niceties like quick rebuilding with Terraform etc, slap a VPC-like thing to make it a similar private network (do use a different subnet so copypasting typos across dev and prod are impossible) and scale down everything (ES node has 8 CPUs and 24GB ram in prod? It will have to do with 2vCPUs and 2GB RAM in dev)
One of the annoying things is you do want to test the replicated/distributed nature of things, so you can't just throw everything on a single-instance-single-host because it's dev, otherwise you miss out on a lot of the configuration being properly tested, which ends up a bit costlier than necessary
I easily got 3K requests / sec out of my laptop at the same time, and it was not a trivial app!
People's expectations have shifted so much it's absurd. If you look at the TechEmpower benchmarks, ordinary VMs can easily push 100K requests per second, no sweat, even with managed languages.
Trivial stuff like static content being treated as static content (files on the disk!) not as a distributed cache in front of a database can do wonders.
Am I just old and jaded?
We threw the switch and watched as postgres, with 640Kb of ram and a tmpfs backed store proceeded to handle all of the query traffic. There were some stored procedures or something that were long-querying or whatever - i'm not a DB person at all, so we switched back to the regular production server about 8 minutes later.
Yes, we did it in production.
but, wtf do i know, i'm the crazy guy who tries to interpret comments generously.
Obviously the tmpfs was doing the heavy lifting, there - and if i had to do a postmortem, i'd wager that filling the OS caches was the main reason the long queries took so long. We didn't do any sort of performance tracing.
The main purpose was to show that these $35k servers could essentially replace the older machines if need be, even though the old ones had FusionIO. I just removed the middleman of the PCIe bus between the application and the memory. It was a near constant argument on the floor about whether or not we could feasibly switch to SSDs in some configuration over spinning rust or even FusionIO, i wanted a third option.
Basically, serve out of registered, ECC memory in front, replicate to the fusionIO and let those handle the spindled backups, which iirc was a pain point.
The issue is not really the number of requests per second, probably, but the number of bytes, which they don't talk about at all in the article; reading manga with no ads is a pretty static kind of application, which could be satisfied amply with a web browser or even a much simpler program loading images from a filesystem directory.
Valgrind claims httpdito runs a few thousand instructions per request, but that's not really accurate; what happens is that the kernel is doing all the work. httpdito on Linux can handle about 4000 requests per second per core, nearly a million clock cycles per request, almost all of which is in the kernel. Of course it doesn't ship its logs off to Grafana. In fact, it doesn't have logs at all. But it would work fine for reading manga.
I assume they are talking about their more dynamic content serving in this post (for things like search, tracking which chapters are read, new chapter listing based on what user follows etc.).
They have a custom CDN that is hosted by volunteers to serve the images for the manga pages. They provide some metrics for that at https://mangadex.network, there are also some older screenshots where they hit 3.2GB/s.
Already an over kill.
Think smaller. Think simpler.
A single machine serving files directly from the file system (yes, from the SSD attached to the machine) will be able to handle a LOT more.
That was a short-lived thing, and has now become a myth perpetuated by companies like Citrix and F5 that sell "SSL offload" appliances for $$$.
Have you benchmarked the overhead of TLS?
In my experience, a single CPU core can easily put out multiple gigabytes of AES-256 (tens of gigabits). This benchmark shows 3 GB/s (24 Gbps) for recent AMD CPUs, and nearly 40 Gbps per core for an Intel CPU: https://calomel.org/aesni_ssl_performance.html
A multi-core server is very unlikely to have more than a 1-5% overheard due to TLS. Even connection set up is a minor overhead with elliptic curve certificates.
This is thanks to the AES offload instructions, which are present in all server CPUs made any time in the last 5-7 years or so. As long as the modern Galois Counter Mode (GCM) is used with AES, performance should be great.
Meanwhile, Citrix ADC v13 with a hardware "SSL offload card" actually slows down connections! I had a very hard time getting more than 600 Mbps through one. It seems to be the way the ASIC offload chip is architected: it seems to use a large number of slow cores, a bit like a GPU. This means that any one TLS stream will have its bandwidth capped!
I haven’t read the article, but the headline alone to me seems alarming, $1,500 a month is a lot of money for only 2k rps.
I've ran far more complex sites with much higher traffic for less.
I am sure they must be using some kind of CDN for sure, however, those options are unlikely to be free
Are they worried about CDNs logging the images their visitors access? Seems like an absurd edge case to worry about in my opinion.
> however, those options are unlikely to be free
I wasn't even talking about free CDNs :)
I wonder if there's merit in them approaching studios with a proper business plan?
That said, I echo that the amazing feat is that they can fit modern inefficient tool choices with poor mechanical sympathy into that budget. The last decade of web-dev tooling has been pushing the TCO of systems through the roof and this post is all about how to struggle against that whilst using those tools.
If they went old-school they'd get another order of magnitude savings. Many veterans know of systems doing 10x that in 10x less cost. Remember C10K was in 1999.
How to learn more about the old-school way without getting a job related to it? Like, topic or book recommendations.
I'm not an old-school guy by any means .. but I might have something to contribute.
Well, for example, what's the old-school alternative to mangadex's solution?
> And where do all the people like you hang out?
We are here on HN.
But there are many ways to achieve 20K RPS without this type architecture and especially without k8s, for less than $1,500.
If this metric is what you are chasing, there are ways to reliably break 1 million RPS using a single box if you don't play the shiny BS tech game. The moment you involve multiple computers and containers, you are typically removed from this level of performance. Going from 2,000 to 2,000,000 RPS (serialized throughput) requires many ideological sacrifices.
Mechanical sympathy (ring buffers, batching, minimizing latency) can save you unbelievable amounts of margin and time when properly utilized.
Basically a container is a glorified chroot. It has the same networking unless you asked for isolation, then packets have to follow a local (inside the host) route. It has exactly no CPU or kernel interface penalty.
Maybe you wanted to say about container orchestration like k8s, with its custom network fabric, etc.
Have you seen most k8s deployments? It's not the containers, it's the thoughtspace that comes with them. Even just using bare containers invites a level of abstraction and generally comes with a type of developer that just isn't desirable.
Much of the work I've done professionally has happened in the webhosting space, and I'm not a new hand at it - the first "professional" website I ever ran was hosted off of a machine running on a Quantum R5000. I have served (and still serve) plenty of static content as files off the disk. My own first impulse when building anything is to use as few moving parts and as simple of a setup as possible.
Requests per second are not a good metric for the amount of work being done by a system, because not all requests are made equal. You say you got 3k requests out of your laptop and that it wasn't a trivial app, but you have only provided a trivial amount of information. Taking a quick look at the features provided by the site, they offer quite a bit of flexibility on both what is and how it's displayed. Filtering options based on original and translated language, adult content filtering, multiple types of search, request routing to low quality images to save on data, category and tag filtering, tracking of what you have read and what your progress is on it, follows and notifications, permissions systems for uploading and updating content, etc. This is all on top of the basic "Display images and metadata about the images" functionality.
They have concerns around privacy and being able to withstand active attackers due to the content they host. They have their own requirements about logging, analytics, etc. Their own concerns about data availability.
You could not meet the design requirements of this website and serve 2k requests per second on a 2 core VM in 2007. I'm not saying there aren't inefficiencies in their architecture and further places where they could save money/increase performance/etc., but acting like two cores of lower-clockrate lower-ipc compute from a decade and a half ago could do all of the work needed to support the features and design requirements they have for this website is pretty disparaging towards the people who built this infrastructure.
What is actually more interesting is to understand what portion is spent on servers versus bandwidth - and what hardware configuration they use to host the site. For example, Is $1,500/mo just paying for colo costs + bandwidth, with already owned recycled hardware (think last gen hardware that you can get at steep discounts from eBay / used hardware resellers...)
That would have been way more interesting to know given the blog title than the choice of infrastructure software they use.
They cited Cloudflare not being used due to privacy concerns. It'd be interesting to hear more about that, as well as why other CDNs weren't worth evaluating too.
CF will also pass through things like DMCAs easily.
Based on their sidebar, it's probably hosted at Ecatel or whatever they are called now (cybercrime host) via Epik as a reseller, the provider famous for hosting far-right stuff.
Regarding DMCA’s, as an entity doing business where they’re legal, what should they do as a middle man?
Don't use them and instead have your middleman be in a country that ignores intellectual property rights and copyright?
I'm not saying CF is wrong to pass them through. I'm just saying CF is not the right choice for a warez site for longevity.
The "trick" has really been known for a decade, or more. Have as many things static as possible, and only use backend logic for the barest minimum.
Modern CDNs also provide lots of functionality from security (firewall, DDOS) to application delivery (image optimization, partial requests).
The NSFW counterpart of MD also has a CDN appropriately named Hentai@Home run by volunteers.
These 2 sites are the only ones rolling their own CDN for free that I know.
A lot of their infrastructure design choices should be viewed with OPSEC constraints in mind.
[1] https://torrentfreak.com/japan-pirate-site-traffic-collapsed...
[2] https://torrentfreak.com/mangamura-operator-handed-three-yea...
[3] https://torrentfreak.com/mangadex-targeted-by-dmca-subpoena-...
I guess it all depends on how much money you bring in for them really.
I"ll join the discord afterwork to see if they need any extra hand.
Gee, how do these people find other people online to work on all of the cool projects. I would love to join rather than playing games after WFH on the same pc over and over again lol
Also, some people hold the view that things like information, media, code can not be “stolen” in the traditional sense, so that further reduces any qualms about associating themselves with it.
Scanlations are often viewed by fans as the only way to read comics that have not been licensed for release in their area. However, according to international copyright law, such as the Berne Convention, scanlations are illegal. [1]
This is a snippet about the Berne Convention:
The Berne Convention for the Protection of Literary and Artistic Works, usually known as the Berne Convention, is an international agreement governing copyright, which was first accepted in Berne, Switzerland, in 1886. The Berne Convention has 179 contracting parties, most of which are parties to the Paris Act of 1971.
The Berne Convention formally mandated several aspects of modern copyright law; it introduced the concept that a copyright exists the moment a work is "fixed", rather than requiring registration. It also enforces a requirement that countries recognize copyrights held by the citizens of all other parties to the convention. [2]
[1] https://en.wikipedia.org/wiki/Scanlation#Legal_action
[2] https://en.wikipedia.org/wiki/Berne_ConventionSomething you made up in your head with literally not a single shred of evidence.
To be more precise, the real reason why such sites are alive is that they delete titles that got licenses in Europe and the USA. Still, publishers can measure the popularity of titles and buy legal rights to publish it, because it's popular enough. It's harder to find manga "raws" than translated versions.
And by that, they're not 100% "illegal" for the western world, and asian companies are not so interested in fighting with scanlations because they need to combat piracy in their part of the world.
It was a completely retarded play on MPA's part and they only managed to get the repo down for days until GitHub restored it even without hearing from the repo owners. So really they only brought about some minor nuisance alongside a bunch of headlines to advertise Nyaa.si for the rest of the world.
https://torrentfreak.com/mpa-takes-down-nyaa-github-reposito...
https://torrentfreak.com/github-restores-nyaa-repository-as-...
It's entirely possible for a copyright owner to construe some kind of secondary liability based on your conduct, even if the underlying software is legal. This is how they ultimately got Grokster, for example - even if the software was legal, advertising it's use for copyright infringement makes you liable for the infringement. I could also see someone alledging contributory liability for, say, implementing features of the software that have no non-infringing uses. Even if that turned out to ultimately not be illegal, that would be at the end of a long, expensive, and fruitless legal defense that would drain your finances.
In other words, "chilling effects dominate".
Perhaps if you are limited by requests per second you could consider how many requests a single user is making per interaction, and if this is a reasonable number.
The website is impressively fast though, I'll give you that.
BTW, the site is not just fast: they serve images on high quality (same as the original [1], that can be multiple MBs per page [2]) at an pretty impressive speed too.
[1]: before someone asks why they don't optimize the images, this is by design since they want to serve high quality images. There is an optional toggle to reduce the image size, but this is disabled by default.
[2]: for those not familiar, the average number of pages on a manga is something like ~20, and this can be read in ~5 minutes depending on the density of the text. So you can easily consume 50MB+ per chapter.
- first request is fast, because you only need to download chunks required for a single page/controller (and you prefetch others in the background)
- changing some parts of codebase requires to re-download only affected chunks, instead of the whole bundle
[1] https://www.telerik.com/blogs/what-you-should-know-code-spli...
Not sure I'm missing something here. Surely you could sample some prod traffic and then replay it with one of the many load test tools out there. You might lose in the geographical distribution, but load testing a web server with 2k TPS sounds a bit trivial.
I don't know how many requests per second it can handle.
Trying a guess via curl:
time curl --insecure --header 'Host: www.mysite.com' https://127.0.0.1 > test
This gives me 0.03s
So it could handle about 30 requests per second? Or 30x the number of CPUs? What do you guys think?
But it seems you think I wanted to start a dick measuring contest?
If your question is genuine: I would serve images via a CDN. The above timing is for assembling a page by doing an auth check, a bunch of database queries and templating the result.
> In practice, we currently see peaks of above 2000 requests every single second during prime time. That is multiple billions of requests per month, or more than 10 million unique monthly visitors. And all of this before actually serving images.
If I am reading that correctly, 2000r/s does not include images, and makes it unclear if $1500/month does.
Best way to figure it out is to use an application like Apache Bench from a powerful computer with a good internet connection, throw a lot of concurrent connections at the site, and see what happens.
I just tried Apache Bench:
ab -n 1000 -c 100 'https://www.mysite.com'
Concurrency Level: 100
Time taken for tests: 1.447 seconds
Complete requests: 1000
Failed requests: 0
Requests per second: 691.19 [#/sec] (mean)
Time per request: 144.679 [ms] (mean)
Time per request: 1.447 [ms] (mean, across all concurrent requests)
Wow, that is fast. Around 700 requests per second!Upping it 10x times to 10k requests ...
Requests per second: 844.99 [#/sec] (mean)
Even faster!What is more important is what kind of requests your server has to serve. Nginx can easily serve 50-80k req/s of static content; 100ks range if tuned properly.
Deploying an instance of Prometheus with *every host is also unusual, to say the least and I don't quite understand their comment to that. If you don't like a pull-based architecture (which is a valid point) why use one at all!? There are many more push-based setups out there that are simpler to set up and less complex.
Does image processing, runs our analytics, runs our Sentry, runs our gitlab-ci runners, and quite a few other things not mentioned expressly
> which would be extreme overkill
That's an interesting argument against k8s ; if anything I find it much easier to work with -- once accustomed to its idioms, ofc -- than alternatives like dedicated VMs, Docker Swarm etc
Getting HA and auto-healing for free is possible without it, of course, but does require much more work, especially if you aim for a somewhat minimised amount of statefulness (as in deviation from the template of your system)
Also S3-compatible storage backends are really aplenty, from commercial offering to simpler ones like MinIO. Ceph just happens to be a bit higher of a deployment investment with the benefit of fantastic performance, flexibility and resiliency. Somewhat like k8s itself, it's a bit daunting at first but does actually make things simpler in the long run (imo)
> 18k metrics/samples [...] are nothing
Well yes and no, the number of metrics isn't relevant per se, but its cardinality is very relevant, and managing that in a single prometheus instance will quickly require some serious vertical scaling, especially if you want to look at data on longer ranges (which, in contrary to logs, we are interested in)
> 7k logs per second [...] are nothing
That's an interesting take . Surely this isn't a world-record-shattering amount indeed, but no one seems to have such a great non-SaaS-or-cheap solution to storing, sorting and querying this amount of logs either (at the resource efficiency of Loki anyway), so maybe we just have a different set of expectations for log management
> If you don't like a pull-based architecture [...] why use one at all!? There are many more push-based setups out there that are simpler to set up and less complex.
Are there really? That is non-SaaS and with as widespread 3rd-party software support as Prometheus does? ie great integration with essentially any database, webserver, runtime, OS, etc?
Because if we talk only about node metrics like CPU etc then yeah, sure there are plenty of options. But (maybe not so) obviously the diagram showing only node exporter doesn't mean that this is the only integration we use -- we collect prometheus metrics for MySQL, PHP-FPM, Varnish, Nginx, HAProxy, Elasticsearch, Redis, RabbitMQ etc (essentially every single piece of software we use).
Fwiw I found very little in the way of open-source solutions to that problem that ticked as many boxes as Prometheus.
As for "simpler to set up and less complex", both Cortex and Loki would be really annoying to manage outside of Kubernetes, I'll happily give you that. But... being able to easily deploy and manage such systems once you have Kubernetes is precisely one of the reasons to use it. You can't say it's complex to deploy itself but then ignore the fact that it largely outweighs this by making reliable operation of complex-but-powerful software on top it, that is precisely one of the upsides of using it in the first place :)
> Does image processing, runs our analytics, [...]
Fair enough, I was strictly going by the diagrams. From my experience with a somewhat similar setup (HA Loki, HA Prom + Thanos with a MinIO storage backend using Terraform + Ansible and docker) I have to say that the most complex and frustrating part was configuring Loki (this was way before they expanded their documentation, which still isn't great). I'd imagine this would be even more challenging under k8s at least if you stray from the vanilla deployment and/or charts. I agree with your statement regarding Ceph, we use it extensively in production (probably on a much bigger scale). However, I think Ceph, unlike MinIO, just adds unnecessary complexity to your setup.
> Well yes and no, the number of metrics isn't relevant per se, but its cardinality is very relevant [...]
Cardinality is something you should avoid when using Prometheus - for exactly that reason. There are, in my opinion, very few good reasons for dynamic labels (ignoring the baked-in cardinality from a setup like k8s). On first impulse I'd say you're doing metrics wrong but then again, I do not know enough about your use case. Maxing out a single Instance of Prometheus is no easy feat however, especially if your infra isn't that complex and/or big. I've used Thanos for so long now, how does the Cortex compactor handle range queries? Does it also compact and create additional 5m & 1h resolution metrics? These might help with your larger range queries.
Just out of curiosity, have you had any look at alternatives like Victoriametrics?
> 7k logs per second [...] are nothing
My remark was just regarding the added complexity as this depends solely on the size of your log messages. If you don't need or use the (awesome!) capabilities of Loki + Grafana and just need a place for long-term storage of your logs, a 'simple' rsyslog server will do just fine.
> we collect prometheus metrics for MySQL, PHP-FPM, Varnish [...]
Many (if not all) of these can be handled by Telegraf or Fluentd plus InfluxDB (not that I'd used that myself, I absolutely love Prometheus and its Eco-system). My tongue-in-cheek comment was mostly about the Prometheus instance you deploy on every server just to scrape metrics locally and remote-write them into Cortex. Why not the more usual setup of (one or more) Prometheus instances scraping their targets and writing to Cortex?
>our ~$1500/month budget
I understand not wanting to show ads, but is there no way for the users to contribute to hosting costs?
BTW, how can I register my VPS on MD@H? Before we had a dedicated form on the page to register interest, at least after the rewrite I didn't find it. Is it only using something like Discord?
Now show an ad, or premium accounts, and it becomes a for-profit endeavour which is straight jail time. I'm unsure about donations.
(Based on previous rulings I followed ~10 years ago, laws might have changed IANAL yada yada)
*not for profit != non-profit
https://old.reddit.com/r/mangadex/comments/nvj7qf/is_verizon...
I'm guessing though they're using some old spam ip/block though, there's a lot more obvious piracy sites then a Manga site. For instance, I can access all the major torrent sites.
That's disappointing. If only I had some choice to ISPs, then I could express my disappointment by voting with my wallet…
It can do much higher requests per second wise on simple requests but most common requests are actually heavy iterative calculations so hence the average of 5000 requests/s
It's basically pirating content