React, Relay and GraphQL: Under the Hood of the Times Website Redesign
open.nytimes.com
open.nytimes.com
I'm always stunned, but at the same time never surprised, when you discover a single webpage is 35+ MB, consuming 2GB of RAM, and consuming CPU as if it were a midrange video game.
That's not what people here are saying, though. Indeed, even your parent comment said how he misses a small HTML pages with CSS. Most engineers would be totally fine with HTML and a modest payload of nice, modern CSS styling, a perhaps a small bit of non-required progressively enhancing JS.
Our problem is with huge JS frameworks used in sites that aren't actual web applications (e.g., Gmail), but rather web sites (like news sites). And I say this as someone who has primarily made my living the past 5 years as a "frontend engineer" (ie, JS programmer). JS frameworks can be wonderful for actual web applications, but they're way overkill for documents online for reading (and also make the experience worse for the reader).
Oh, and we also hate of dozens/hundreds of kb of unneeded font downloads.
Also, for the record, I'm a big fan of (appropriate use of) white space and color. I definitely come down heavily on the side of bettermotherfuckingwebsite.com (vs motherfuckingwebsite.com).
My favorite counter example is sankei news site. http://www.sankei.com/
All their pages end in .html.
Do view source and vast majority if byte is used for actual text content, rather than javascript and markups.
Mobile friendly too, but you wouldn't know it because they disabled responsive view unless it's visited from an actual phone. No premature mobile view with hamburger kicking in when viewing it on desktop.
Comparison of View source between two sites is quite amusing.
view-source:http://www.sankei.com/smp/
view-source:https://mobile.nytimes.com/
Bug, not feature. If I'm viewing a site in a narrow window, I expect it to collapse responsively. That site doesn't.
Personally I prefer this because information is always in the same place regardless of browser size. I don't have to worry about it moving around or hiding.
Do you not understand how URL rewriting works?
For NY Times, it's about 1:10.
Sankei.com, it's about 1:4.
News sites are also more complex than you would assume. I went from building products at Facebook to a large news site. What shocked me was the surface area of the user facing products. There were so many little one-off pieces of functionality and randomly integrated services. Tools like React and Relay make it much easier to manage this complexity and promote code reuse.
What's the advantage of server rendered React app over server rendered anything e.g. php jsp etc?
That said, I'm not aware of anyone doing this at scale. I know the BBC were considering it. My company will be using it to product AMP pages from our JS stack.
Documentation and infrastructure setup for this project aren't great at the moment, but will get a facelift shortly.
So instead of "we use GraphQL, much love" + basic example and how it looks on React - a "here's how we take that structure and resolve it and return it." Because that structure looks amazingly sweet - but if in the background it's requiring circles of work, work and rework...
Anyhow, maybe I just don't understand it enough.
https://github.com/rmosolgo/graphql-ruby
Ryan supports development with a pro version that has a bunch of neat features, including built-in support for a handful of common authorization frameworks. I highly recommend it.
[0] http://sangria-graphql.org [1] https://www.playframework.com [2] http://sangria-graphql.org/learn/#schema-definition
You need to be very careful. To be optimal, you'll basically need to write very complex resolvers that inspect the AST themselves and fetch the data optimally which at that point is almost as writing your custom execution module.
That is basically what i did, custom execution module to translate a graphql request to a single sql query (https://subzero.cloud/)
I've spent a significant amount of time over the last year or so optimizing GraphQL servers built on domain-driven services (i.e. joins aren't an option) and managed to get to equal (or very marginally worse) performance to existing handcrafted endpoints that returned equivalent data (it was possible to build the same UI, even though the payloads weren't identical).
There are areas where GraphQL is inherently inefficient (trying to work on ways to mitigate these issues), but the reality is that deeply-nested UI appears to be less of a problem than I originally thought it would be.
If you have a 3 level query (3 tables), the best you can hope for is to get 3 sequential queries, like get first level, collect the ids, request the second level, collect the ids, get 3rd level. It gets even more complicated teh more levels you and the bigger the dataset returned.
all of this can be done using a single join which is one roundtrip and it's faster.
In the 'join in SQL' case you could identify particular cases which are doing sequential queries and implement a different loader which just does one. It's not automatic but perhaps in most cases it's not necessary to do this step anyway. In the worst case you're back to doing as much work as you would for a bespoke API endpoint, but that's not the typical case. How much of a problem this ends up being in practice very much depends on the type of app you're building and how you intend to scale it.
In practice this probably means adding some logging of how long requests take and graphing it (say, 95th percentile request time) from time to time to spot pathological queries. Even better if you can automate it. I think this is stuff that everyone should be doing (after a certain stage), regardless of whether you are using GraphQL or a bespoke JSON API.
It's perfectly valid to identify subtrees of your query that would benefit from being executed as a database query, but to do that to your entire query just sounds like you're asking for trouble, I'd even go so far as to call it a premature optimization.
I am talking about queries/joins between tables that have foreign keys between them, like client/project/task/comment. I bet 90% of graphql schemas expose those kinds of relations between types.
For those type of relations (with FK) i can generate a single query that is as fast as it can be (certainly faster then dataloader) and as far as i've tested (a few millions of rows in tables, 3-7 levels in a query) i didn't leave the fast territory :) Of course there might be edge cases ...
About premature optimisations. Everyone likes to quote that, but never the full one which is "We should forget about small efficiencies, say about 97% of the time: premature optimization is the root of all evil." https://shreevatsa.wordpress.com/2008/05/16/premature-optimi...
The paper where it was published was about using goto to optimize loops and such small things, it was never "targeted" at algorithms and architecture.
As i think i mentioned here above, i was getting 10X throughput with joins compared to dataloader and i would not call that premature.
I get my data for a full UI within the budget i've allowed for uncached scenarios (100ms, but I want to go to 50ms). The approach you're suggesting will (in some circumstances) give me some short-terms win in terms of throughput and response times, but you've not said anything to suggest that I won't lose these benefits as I gradually transition into a domain-driven or micro-services (yuck) architecture.
I like to build my GraphQL servers under the assumption of a domain-driven architecture (because that's where all the projects I've worked on seem to end up, your mileage may vary), and then shoe-horn in some short-term performance tricks when I can.
I'm possibly a special snowflake here, but it's been a long time since i've had the opportunity to work on a project where I can go straight to the DB. Be it Elastic Search, a 3rd party, ill-advised micro-services, or complex logic in-between storage and presentation; nothing has quite been a pure DB project in the last 6-7 years.
Of course, you could argue this is premature architecture ;) but many of these complexities are from day one, or at least pretty early in a project's life.
I'm getting the impression that if nothing else though, I'm probably gonna need to wait a few years for this tech to stabilize more. Which I find very unfortunate, because the standard way of doings things these days feels very unpleasant to me (I hate writing boilerplate more than anything: have an overuse injury, so typing is the worst part of coding)—and the GraphQL queries do seem to make a lot of sense (although forcing clients to have explicit knowledge of schema structure seems a little dangerous...).
If I have a schema like this:
users
projects (many per user)
blog_posts (many per user)
What SQL do you generate to handle a request for everything in a single query?simplified example query https://pastebin.com/7J2ZswRC
IMHO Relay doesn't make a lot of sense. Apollo Client [1] has a much better feature set for most use cases, doesn't need React and is better documented.
[0] http://dev.apollodata.com/tools/graphql-tools/resolvers.html [1] http://dev.apollodata.com/core/
I've been sticking with it though, and I am enjoying it. I feel like I have a greater grasp of what's going to be executed and when than I ever did with Relay Classic, and the file size + performance improvements are worth the cost of admission in my mind.
> the file size + performance improvements are worth the cost of admission in my mind.
Is this Relay Classic vs Relay Modern or Relay vs Apollo?
Relay Modern is 20% of the size (or 5 times smaller) than Relay Classic, which (if my calculations are correct, I don't have equivalent environments set up) is just over half the size of React Apollo + Apollo Client.
You can also see a lot of examples of GraphQL server code for JS here: https://github.com/apollographql/launchpad
It includes connecting to DBs, APIs, etc.
https://dev-blog.apollodata.com/optimizing-your-graphql-requ...
Examples more specific to a particular backend technology feel redundant because my assumption is that once you're in the land of calling functions, we don't need to hold your hands anymore.
The most important thing is to be aware of the different batching strategies that are available to you in each GraphQL implementation because I believe this is the most critical part of getting a GraphQL server to perform well with anything other than a graph database.
http://joelgriffith.net/lessons-learned-wrapping-a-rest-api-... For starters
Plus bonuses if you have more than one database or are mixing data from your database and external APIs in your backend responses.
Please try it for a small project, it is unbelievable, but you'll probably enjoy it.
(For the record: I've just used https://github.com/graphql-go/graphql and Lokka on the client, because it is simple and does nothing fancy, it's a thin wrapper over XHR, I think.)
You can easily do that with a RESTful API too.
For example you have a Product type, and then write the resolver function that queries the database and/or another API. When you have the data, you pass it back to the client via Apollo or Relay.
The big advantage over REST is that the client can define what data it wants and how it wants it. If you are full stack dev this isn't such a great advantage, but for bigger projects where front/back are spread among many engineers this can be an advantage. Also, since the schema defines the types, your API is almost self documented so to speak.
The big disadvantage is authentication and authorization. We kept using REST for authentication, and we couldn't find any ready made solution for role based authorization like you have in Express, Hapi, etc.
I think a combination of REST and GraphQL is the better approach.
A good GraphQL backend resolves the graph into query results efficiently, rather than just field by field. But that does take a but of getting used to....
I would think React might make sense for realtime dashboards and similar webapps.
But does it make sense for displaying articles, navigation and ads?
Why not?
My expectation is that the site does the templating on the client then. And my experience with websites that do that is that they load slow, behave sludgy and suffer from all kinds of display errors.This might be a worthwile tradeoff to quickly build a highly interactive realtime interface. But for a newspaper? As a user I would be very much turned off to endure all that just to read an article.
If the user is on any hardware from the past 7 years, a React developer would have to do some distinctly bad programming to make it behave sludgy. Any poor performance is likely to be from the same things that make a classic static page slow: large media files and tracking scripts.
As a sidenode, it does feel sludgy:
- Scroll is not smooth.
- It does not adapt well to different browser sizes
- After a few seconds the page "jumps down"
But the reason might simply be the overkill of JS,animations, overlapping elements (like the static header), ads, dynamically loaded stuff and other crap. When I turn off JS, some of the problems go away.
A lot of this "JS is slow, what happened to good old HTML and CSS, get off my lawn" stuff is simply confirmation bias.
You could even write an entirely server-rendered web application, or a static website, in React should you be so inclined. In fact, I do the latter for my (very simple) personal site at https://davnicwil.com!
Nothing about React obliges you to do client rendering, a SPA app, or anything complex at all. It's, at the end of the day, just a view rendering library.
I'd say react is a perfect fit for this type of thing. A basic news site it might be, but there's a lot more going on under the hood than you'd imagine.
A typical template is an HTML file (or some variation of it) within which the dynamic content is inserted using string interpolation.
A React view on the other hand is a piece of Javascript code which can compose small snippets of view as JSX, use programming constructs like loops and conditionals and finally return the assembled result.
Basically, React views are pure functions that return a validated HTML snippet. Normal templates are big blobs of strings with logic mixed in.
How does a server side React app look like? Is it basically a node application then?
It pretty much has to be, there is no way NYT is giving up showing up in search results.
> How does a server side React app look like? Is it basically a node application then?
There is almost certainly Node somewhere in the pipeline. It's possible to build static content with Node/React as well, but NYT has dynamic server functionality as well so it would not surprise me if there's at least a Node layer in production.
It's possible to build static content
with Node/React as well
Without node.js? How? What type of software is the server side React then?Realistically, a server-side rendered JS app is also going to run most of the code that runs in the client per page load, so you would also have to consider initialization costs as well. I've had to work on one that took 100+ ms to render a static page (without accounting for network latency), which a static file server could render the same page orders of magnitude faster.
Long story short, most of the newer, non-string based JS server-side rendering does not consider performance a factor and consequently, perform pretty terribly. There are band-aid fixes such as putting a reverse proxy in front or running on super fast hardware, but it's like putting a band-aid on a bullet wound.
[0] https://github.com/daliwali/simulacra/blob/master/benchmark/...
I would argue it's the perfect candidate for this purpose. What does a news site need? Articles? Ads? Basic nav? Angular or Ember would be overkill. And with Redux, it's very easy to think about complex UI state.
My only complaint with React is how large it is given what it actually does. Loading 100kb of JS (not Gzipped) seems very heavy.
It's a 3kb React alternative with the same API.
No it doesn't. We (as a marketing tech company) deal with major publishers all the time and come across all kinds of ridiculous and complicated tech used to show a basic article when a simple static website would be easier and faster at this point.
It's a lack of good talent, resume-led motivations for choosing technology and overall poor vision in execution by CTOs and management.
As far as the criticism, would you be kind enough to explain what you mean by "relay concepts", and how GraphQL failed to satisfy those needs?
My interpretation is that they are now using a GraphQL API but fetching from it with regular HTTP fetches and managing data with Redux on the frontend.
I think I would just still be more attracted by Appollo framework which is more a redux-like syntax, so more consistent with all the workflow of the apps I used to develop. But maybe Relay has better benchmarks?
Also, if I want to stay REST but with optimist transactions between the backend and frontend, I prefer lighter lib like https://github.com/tonyhb/tectonic or even I just write some fast redux-saga watchers that helps to make my frontend always synchronized when my app calls a mutating db request.
GraphQL has helped us make an api that is easy to understand, easy to change and and easy to use. We love it.
The caching section in GraphQL docs is cringe worthy, and, as evidenced by the article, that's exactly what Times are going to do: try to slap global ids everywhere and pretend it's ok
IMO coupling the data layer with the presentation layer is a terrible idea.