The Surprising Path to a Faster NYTimes.com
speakerdeck.com
speakerdeck.com
Focusing on shifting all "rendering" into front end JS seems like it will lead to more difficulty in the long run instead of using a more efficient structured page creation mechanism.
I am curious how the static pages were created. Others here are speculating that templating was not done. If not; what does "republishing" mean exactly?
I am not attacking what they were doing, my point was that this is an issue of delaying the move to a template based system way longer than necessary. I am speculating on whether they did move to some templating, but not all the way, and wondering what exactly they were using before that.
Saying "static HTML files" is well and good, but it doesn't explain how they were created. By hand??
Also, saying "90 days to republish" seems to be suggesting some process they have in mind besides manually scraping the data out. If you have 1 million random files, scraping could easily take years not 3 months. It would be interesting to know what process they are suggesting to follow in those 90 days. My speculation is they did use some sort of outdated CMS software.
It seems reasonable (and possibly low) for a 90 days estimate to extract the content from the variety of versions of static page, structure it, and then publish it in a more modern fashion.
It's all very well to say they should have used a more efficient system from the start, but "the start" in this case is 1996, which is the wild west in terms of best practices.
Upload the corpus and throw a bunch of cloud instances at it.
Publishing a page is running content from a CMS through a templating system. However, the time spent executing templates isn't the only factor in the duration of the "publish". The slides refer to a compilation step (which actually also included a preprocessor step), and includes delegating to a service to copy the resulting page to disk and ensure the write succeeds for all data centers. For data consistency and system monitoring, we essentially treat that entire process as an atomic action and wait for all parts to finish. Additionally, since "publishing" is a core process for us, we avoid doing massive publishes that might risk the systems involved in the successful publishing of current articles. So increasing the number of these for the sake of pushing code is considered too risky. Yes, there are ways of mitigating that risk, but dealing with this legacy problem once and for all is a better path forward than scaling up this solution.
Just thinking about it raises all sorts of questions about whether the browser/rendering engine can actually reliably know that information, but it doesn't mean I can't dream.
I'd first like to say that this is the deck from a presentation at Velocity NY last week. Like most other talks, separating the slides from the presenter can make interpreting the context difficult. I did try to make an effort to have my slides provide useful information without me presenting them, but I acknowledge that I may not have done enough in that regard. I also received feedback from people present that there were too many bullet points and my font was too small. Can't please everyone I guess. But if you have a link to what you consider the "perfect" slide deck where unambiguous context is maintained without video of the talk, I'd love to study it in order to improve.
Other replies will be directed at the specific comment thread.
My physical NYT copy from 1980 is fine. It was "published" this way, and it stays this way.
What we're really saying is that if you want to go the static route, you can't go half-way: everything that _is_ the page gets deployed in one file. I doubt very many people who think they have static pages actually do.
Once you are focusing on e-commerce and SEO as an executive team, are you still committed to journalism?
Two recent pieces are leading me to believe that NYT is floundering around, without a real cohesive online strategy, still:
[1]http://www.cjr.org/the_audit/the_new_york_timess_digital_li....
[2]http://www.cjr.org/the_audit/the_new_york_times_cant_abando....
Whoa.
At least I think that's what they were saying. Always challenging to deal with a slide deck meant to be presented by a human, but without the human doing the presentation.
#1 really did surprise me, because I had always assumed that serving static pages would be really fast. I guess I never thought about sites with millions of pages.
a possible way of approaching that problem is a divide and conquer approach with a reverse proxy that assigned manageable chunks of their content across numerous machines each serving less than millions of pages. the nyt already has a /<yyyy>/<mm>/<dd>/<section>/<subsection>/<slug> url scheme which would make this less painful.
I'm not sure how inefficient this would be, though, certainly a time investment, but it ends up offloading your disk i/o/ issues by creating more and more s3 buckets (or what have you) and routing via a proxy. i'd be curious to see when s3+cloudfront-as-host becomes too slow simply because of disk i/o limitations, although s3 almost certainly has its own abstraction above the bucket i'm not aware of which mitigates that.
it still doesn't address the serious and complex frontend issues they were facing, which seemed much more onerous to be honest. their server rendering seems to be pretty lean though, it looks like dom processing and client rendering make up easily 80% of their 3.6s pageload time.