How we made editing Wikipedia twice as fast
blog.wikimedia.org
blog.wikimedia.org
Yes, three seconds for a page render is still uncomfortably slow, but it's substantially faster than the original implementation, and unsurprisingly frees up CPU for additional concurrent requests. It's a shame Wikimedia didn't have this platform available to them earlier.
Today web developers have many high-performance platform options that offer moderate to good developer efficiency. Those who use low-performance platforms may do their future selves a service by evaluating (comfortable) alternatives when embarking on new projects.
I'd be more interested to see what, if any changes they have made to their other server components. Are they using NGINX vs Apache? What kinds of caching systems are they using? What is their database layout? Is their data sharded? These kinds of things I think are much more interesting, and at their scale at least as impactful.
I do know they use Varnish as their cache server.
Also, speaking from a few years' personal experience with HBase / Cassandra, the support for non-Java languages on these two NoSQL databases is small (though it is getting better.)
Though node.js tends to be very flexible in terms of wrapping a friendlier interface around a less friendly one.
The medium-term goal is to store all revisions as HTML, so that views and HTML-based saves can become storage-bound operations with low double-digit ms latencies. This will mostly eliminate the remaining 2-4% of cache misses, and should get us closer to actual DC fail-over without completely cold caches.
There is still a large amount of work to be done until all the puzzle pieces for this are in place, but we are hard at work to make it a reality. A major one, the bidirectional conversion between wikitext and HTML in Parsoid (https://www.mediawiki.org/wiki/Parsoid), is already powering VisualEditor and a bunch of other projects (https://www.mediawiki.org/wiki/Parsoid/Users). Watch this space ;)
GD is using it to store user-generated content, and it works very well, most responses (lookup and processing) were sub 20ms consistently under heavy load... (I'm hoping they write a DNS service that does similar).
Keep in mind, Wikipedia is at scale that most web applications will never, ever get to.
It's important to think about performance and scale, but it's not the only important trade off engineers should concern themselves with.
It's also not necessarily about scale, as in number of users that one needs to serve. There are lots and lots of problems where performance is very relevant, from startups that can't afford to scale horizontally, to projects in an industrial setting that have to be rock solid, to participating in real-time bidding on RTB platforms, to games in which frames being dropped are a disaster, to the platforms on top of which we build stuff.
Quite the contrary, I believe that too many software developers are focusing on building front-ends to a database, but this isn't necessarily because these problems are solving real needs.
The win here was about individual page load time. And page load time is just as important, if not more so, for something new trying to vigorously grow as it is for the big sites.
(Disclaimer: HHVM alum.)
So, more request can be served with a smaller number of servers.
The un-shortened URL is https://ganglia.wikimedia.org/latest/graph.php?r=year&z=xlar...
(I should've put a disclaimer there: I am one of the WMF engineers involved in the project).
His machine could easily handle the traffic thrown at it, but the page load time was slow. 2 seconds at best, with all machines involved being almost 100% idle. With various common tweaks such as caching, we got it down to about 800ms. We eventually replaced it with our own solution and got it to 50ms.
Scale never entered the picture because from our testing, we had the machinery in place to handle enough traffic to be wildly profitable. The user experience at 2 second page load times was vastly different from the user experience at 50ms though.
MySQL over Postgres because that's what we had experience with at the time. Redis as both a cache and ephemeral store.
As I read the article I thought exactly this. Scaling is a nice problem to have. I keep playing around with Nginx, elasticsearch, and other tech that's supposed to help PHP's performance problem, but from personal experience, missing database indexes for complex joins have been the issue every time. I'm unlikely to get to use the tech I play with unless I switch jobs.
But back to the point, this is a great writeup, and I think a big step forward for both PHP and HHVM
This will also make tiny MediaWiki installs just as much faster.
According to that 99% of magento2 tests are passing.
I don't think Wikipedia is really a typical case. Most websites probably don't have the CPU burden Wikipedia faces.
[0]https://en.wikipedia.org/w/index.php?title=Barack_Obama&acti...
Perhaps MediaWiki should figure out how to use a weaker form of computation?
That's my experience when I view Wikipedia. I am a Wikipedian who has been editing fairly actively this year, and I almost always view Wikipedia as a logged-in Wikipedian. I see the Wikimedia Foundation tracks the relevant statistics very closely and has devoted a lot of thought to improving the experience of people editing Wikipedia pages. I can't say that I've noticed any particular improvement in speediness from where I edit, and I have definitely seen some EXTREMELY long lags in edits being committed just in the past month, but maybe things would have been much worse if the technical changes this year described in this interesting article had not been made.
From where I sit at my keyboard, I still think the most important things to do to change the user experience for Wikipedia editors is to change the editing culture a lot more to emphasize collaboration in using reliable sources over edit-warring around fine points of Wikipedia tradition from the first decade of Wikipedia. But maybe I feel that way because I have worked as an editor in governmental, commercial, and academic editorial offices, so I've seen how grown-ups do editing. I think the Wikimedia Foundation is working on the issue of editing culture on Wikipedia too, but fixing that will be harder than fixing the technological problems of editing a huge wiki at scale. Human behavior is usually a tougher problem to solve than the scalability of software.
By the way, the article illustrates the role for-profit business corporations like Facebook have in raising technical standards for everybody through direct assistance to nonprofit organizations running large websites like the Wikimedia Foundation. That's a win-win for all of us users.
Completely agree with you. Also agree that fixing culture is often harder than fixing technological scaling problems!
All that said, kudos to Wikimedia Foundation for addressing the speed issues for uncached pages. Great work!
I keep hearing this, but it isn't true anymore. For something like wikipedia, even when I'm logged in, 95% of the content is the same for everyone (the article body). You can still cache that on an edge server, and then use javascript to fill in the customizations afterwards. This will get you two wins: 1) The thing the person is most likely interest in will load quickly (the article) and 2) your servers will have a drastically reduced load because most of the content still comes from the cache.
The tradeoff of course is complexity. Testing a split cache setup is definitely harder and more time consuming as is developing towards it. But given the page views of Wikipedia, would be totally worth it.
(Also make sure the javascript sets a cookie so that wikipedia can fall back to non-cached if javascript isn't enabled.)
I think we're reaching a more acceptable point in technologies where our proper test coverage ensures that using an unblessed release in production is acceptable.
EDIT: I didn't think I would need to mention the obvious reason nhtechie there mentioned. That was my whole point after all, for the ones who understood it before downvoting, but whatever.
I run a similar site (95% read-only), and have been pondering whether it would make sense to use something like Varnish's Edge Side Includes (Like SSI, combining cached static page parts and generated dynamic page parts) -- I wonder if they've considered that and what the results would be like?
Varnish (with ESI enabled) -> Nginx (with memcached module enabled) -> PHP-FPM.
Most of the page is served with Varnish with the ESI directives hitting the Nginx server which serves the fragments from memcached if present, otherwise the PHP-FPM server is hit which then returns the results and sets the fragment in memcached with a low ttl.
In the end I ended up placing special HTML tags where the ESI tags would have been placed, and then making AJAX calls (on dom ready) to swap out the static content with the dynamic user-specific content.
It worked well for our needs.
https://phabricator.wikimedia.org/T34618 https://www.mediawiki.org/wiki/Requests_for_comment/Partial_...
Personally I doubt that it would help much – there are hardly any non-dynamic parts of the page in MediaWiki. (Everything is customizable using server-side configuration, modifying magical pages, or both; or dependent on user permissions or preferences; or both.)
ooo pretty!
A modern persistent web apps running in Python/Java/Ruby/etc is able to perform preparatory work at startup in order to optimize for runtime efficiency.
A CGI or PHP app has to recreate the world at the beginning of every request. (Solutions exist to cache byte code compilation for PHP, but the model is still essentially that of CGI.) Once your framework becomes moderately complex the slowdown is painful.
Is this why we get logged out every 30 days, to boost cache hits for users who rarely need to be logged in? (It seems like every time I want to make an edit I have to login again.)
"We currently use Varnish for serving bits.wikimedia.org, upload.wikimedia.org, text of pages retrieved from WMF projects, and various miscellaneous web services. Nothing uses Squid."
And "progressive" with PHP but not Squid ? They're both technologies from the same era, Facebook just made PHP more usable.
For GET requests, particularly those reached by clicking a link from elsewhere on the site, faster is better... Luckily, many of these types of requests can leverage a cache.
Save early, before the Editor actually presses save. Only commit the change if the Editor actually presses save.
This will improve the Editor experience by making save faster at the expense of CPU time. Predict well enough, or do enough processing client side, then you won't need extra server side CPU used.
What is more precious to you? The human Editors or some dumb pieces of silicon?
However, the method I mentioned can be done with zero extra server load if done correctly. I've done this before with other editors, and it definitely did not take a team 12 months to complete.
It is a full stack optimisation however, from UI all the way through the DB and the rest of the stack. Now that they're using only 10% load, there's plenty of room to breath.
Saving a draft of what someone may have spend hours typing into a crash prone browser is good UI practice regardless.
Saving a draft is indeed good UI practice, but best practice is never "regardless" in the real world, particularly in cash/manpower strapped non-profits.
I think the effort of my optimisation would cost less time and effort, and give better results.
It would also waste less Editor time.
Mediawiki already allows drafts to be saved.
Still, 3 seconds to load a page feels like a slow page and should be a lot faster.
By comparison, uncached page load time appears to be ~800ms
Also you didn't even scroll down to the next graph, where he shows the average page load time (for logged in, not anonymous, users) as 800 ms. (which is obviously larger then the average page load time served from cache). The first graph was the average amount of time it takes to save a wikipedia edit.