Hacker News Clone Using GraphQL and React
github.com
github.com
It's fantastic to see people share their knowledge, and turn theory into practice. It's very easy to spout on about best practices and good ideas without demonstrating how they can be implemented in the real world. And this application does demonstrate some good ideas.
...But some bad ideas also. I think one problem we have in the frontend world is that developers of very small, very simple sites are pulling in solutions entirely unsuited to their scale of problems. They think they are being careful and investing in an insurance against rapid expansion, but what they're actually doing is overspending on the wrong investments, making things expensive that don't need to be expensive, and betting against agility. This seems like a mistake.
I encounter too many teams who are building web applications by bunding a cornucopia of technologies they don't understand and allow "best practice" examples on GitHub and Medium to do their decision-making for them. They disengage critically from their technology choices, which has the short term benefit of getting them moving faster, but the long term cost of never developing their research and evaluation skills.
This is not your fault, of course. People always want a silver bullet, a framework prefabricated and ready for drop-in domain code. And to be fair, most websites are very like each other. I just wish software developers chose their tech a little more critically.
But I guess WordPress is easier to work with if you're a non-technical person who won't be doing much else than adding new posts from time to time.
Imagine you want to sail across the Atlantic and, to acclimate to the necessary tasks, you started small: sailing down the Thames, then sailing to Calais, then sailing farther and into the deeper ocean, before you decide to take the big trip. Now imagine you posted a blog post about this, and the first thing someone pointed out is that you can take the ferry to Calais.
That's what you've done.
That's fine when we're writing personal projects, but don't professional gigs require a greater standard of care?
Imagine you need to need to transport water to your home. You start by consulting best practices and blogs and conclude that the only way to take water somewhere is using a helitanker. This is a dual rotor helicopter with an enormous hold. You consult a building planner to construct a helipad on your roof. You seek an expert helicopter pilot. You spend several weeks poring over different models of helicopter.
You can do a lot of iteration and eventually reach the right result, but that still won't stop your house burning down in this analogy.
I was referring to projects like the one linked, which are quite clearly either technology demonstrations or practice stuff.
Be that as it may, praise the author for creating a fully-functioning case study from which audiences could learn and discuss.
- https://news.ycombinator.com/item?id=15297922
- https://news.ycombinator.com/item?id=15134949
EDIT: Not sure why my comment is being downvoted for pointing out that this exact hobby project has been continually posted in the last couple of months.
Because of the following reasons:
1) There's the subtle implication in comments like yours that this posting is not worthy because it was posted before.
2) It doesn't matter. Stuff is re-posted all the time, sometimes simultaneously. That's just the way the web works.
3) It's not your job to moderate or research this site for others.
I understand the subtle difference between the two deliveries, but I wouldn't be so dismissive.
re 1) The assumed unworthiness is reading to much into it. Plenty of people point out previous posts. These are both good to read previous comments, and also to assess the spamminess of the post/poster -- how much this is so is up to the readers themselves to decide. I'm glad if someone has done the research for me already.
re 2) Does matter, see previous.
re 3) HN is a public forum, not a workplace.
Well, it _is_ less worthy if it has been posted before, isn't it?
> It doesn't matter. Stuff is re-posted all the time, sometimes simultaneously. That's just the way the web works.
Being common doesn't make it desirable.
> It's not your job to moderate or research this site for others.
Is it a bad deed if it's not "your job"? In my eyes, someone needs to do it (and it would be great if we had mods to do it instead).
Disclaimer: I'm on terrible internet connections most of the time and it really makes you hate all this modern web crud. Often "dynamic" sites don't load at all because while the initial HTTP request eventually gets through, one or more of the following XMLHttpRequests or whatever fail, and a lot of the time there is no handler for retries, or no error handler at all, or moronic timeouts (either far too long or far too short).
This is a pretty good example of older websites vs 'web 2.0' I guess.
HN is one of the few websites that consistently loads properly for me.
I can see some value in describing alternative or better ways to do what this web app does, or of critiquing the particular way the components are used. But rants about hating the js ecosystem or this HN implementation just makes it seem like you’ve had a horrible day at work. OK if that’s the case, we’re sorry. Go out for a run or a drink.
After 1.1 MB transferred, I experience subjectively sluggish page transitions and less-subjective FOUC on a recent MacBook Pro and Chrome over fast wifi. The features in this stack are impressive and appealing, but IMO they are not worth the general performance cost weighed by hundreds of kilobytes of JS abstractions.
If you try to create an account you get an internal server error. I assume that is what the OP means.
Or maybe I misunderstood what you meant?
[edit] In fact, in the README at the root of the project:
"Next - (Routing, SSR, Hot Module Reloading, Code Splitting, Build tool uses Webpack)"
What happens when two components used in two totally different areas of a single page need common data? Relying on a cache feels like a hack.
I really love React but really dislike the way I see it being used these days. I think of it as a polyfill for a UI component that should be stateless and free of side effects. Not an end to end controller/fetcher/renderer.
You can add additional layers/groupings/couplings but that isn't React doing that...
> What happens when two components used in two totally different areas of a single page need common data? Relying on a cache feels like a hack.
When declaring what data you rely on is as simple as a GraphQL query, it's basically like declaring what data goes in your component. It's NOT like building ORM layers directly into your UI components.
* A network request is kicked off if the cache doesn't contain the data the component declares a dependency on.
* other components that depend on any part of the requested subgraph will update when the response comes in
* query observers (which HoC'd components really are) will similarly fire, potentially triggering other side effects in your application
Data binding has side effects and the implementation of the thing doing the binding may cause things to happen, may cause some state to change. But I still wouldn’t consider the declarative data binding itself to have side effects.
If two components used in different areas needed common data that would get pieced together through the fragments. Just depends how you architect your app.
Not for React but it's basically the exact idiomatic use case of Relay. GraphQL or React don't require it at all however.
There are no traps there, everything is configurable. The documentation is really clean. Highly recommend Apollo.
Once you have this stack wired up, it empowers you to be quite nimble with your development.
Arguably that's it's own trap. Convention over configuration has saved millions of man hours in e.g. Rails projects, but a lot of JS land involves sinking a lot of time into the opposite.
Although I think coupling components with data structures is the smaller benefit. The real benefit of graphQL is not the client architecture, the real benefit is decoupling frontends from backends, unifying disjointed backends, and efficiency.
In the name of efficiency though, graph QL is projecting a graph into a tree which involves denormalizing & duplicating data in the response. This is where netflix's falcor is cleaner. Check out my little graphQL demo repo / read me https://github.com/joshribakoff/graphql-demo I write more about the benefits of using graphQL, and also the down sides.
The ugliest part of it, or one requiring most hacky solutions, is optimising queries. For example, there are some nasty N+1 traps or "oops-I-joined-your-entire-database" problems when using SQL.
JoinMonster helps with this, but it is a bit too much of a "full solution" or framework to just quickly use. You'll also lose a lot of control over your queries which was a big reason for me to move from Django/Rails/etc. style tools to trying out GraphQL.
This project seems to avoid databases entirely and it seems to have an N+1 queries problem when fetching comments. Each comment is fetched with a separate API call and a single HN Post can easily have hundreds of comments... even if cached, this is a problem.
The SQL "WHERE IN (list of IDs)" query combined with smart caching and batching (see: Dataloader by Facebook, JoinMonster) is a decent solution but requires some amount of good old manual work.
tldr; GraphQL is nice but has its own set of problems. Still no magic bullets.
Haxl seems particularly suited for this (I think it's the same thing as Dataloader, both from facebook?) as it can efficiently batch and cache queries to multiple sources at the same time. It seems like a good layer between GraphQL and your N data sources.
And comparing to something like Django/Rails - there's a lot more DIY involved.
For restricting queries, an option is using persisted queries (e.g. https://dev-blog.apollodata.com/persisted-graphql-queries-wi...).
The idea is nice, having full control in development while locking down allowed queries in production.
You'd still have to deal with the problems you mentioned in development, some of which are indeed a little ugly (at least currently).
However, I totally understand the confusion in my sentence and actually this thing you're talking about is also useful to me :)
Does anyone know what tool was used to generate the system architecture diagram? It's really clean and clear - https://github.com/clintonwoo/hackernews-react-graphql#archi...
"It is intended to be an example or boilerplate to help you structure your projects using production-ready technologies."
He could have used less tech, but then the project wouldn't have served it's purpose.
Seems like with more comments it is much slower.
Separation of concerns went right out the window didn't it? I'd like to see Hacker News Clone Using 300 lines of server-side Python code that works on browsers without JavaScript
Next.js is a framework that does roughly all of this, except it's currently missing (or not making) decisions for you regarding GraphQL as a data layer. More information here: https://malloc.fi/building-decoupled-sites-and-apps-with-gra...
For a fully integrated GraphQL option I looked into Gatsby.js, which started out as a static site generator. It is still this at it's core, but the move for GraphQL for everything seems like a good move: https://react-etc.net/entry/gatsby-is-a-static-site-generato...
Is this served from a CDN?
Great job putting together a robust example application.
You literally have a function called seedCache() :-)
What the hell is the "right way" to build a web app nowadays? That's what I want to learn. I've figured out how to put up a static page pretty well, but this is so different. Is this site what is best practice now? Or is it one of many choices?
Take what you like, leave the rest, that's probably the "right way" to build a web app.
React - (UI Framework)
GraphQL - (Web Data API)
Apollo - (GraphQL Client)
Next - (Routing, SSR, Hot Module Reloading, Code Splitting,
Build tool uses Webpack)
Redux - (State Management)
Node.js - (Web Server)
Express - (Web App Server)
Passport - (Authentication)
Babel - (JS Transpiling)
Flow - (Static Types)
ESLint - (JS Best Practices/Code Highlighting)
Jest - (Tests)
Yarn Package Manager - (Better Dependencies)
Docker - (Container Deployment)It took me a long time to realize that "launching the app" is not really what these tools are for. You can create perfectly fine awesome web apps with just HTML/CSS/ES5.
But as it gets more complex, or you add collaborators, or you feel how much easier/faster certain things would be if you could operate at a different level of abstraction... the world does tend to pull you towards these tools. None of them are necessary though, they just trade a little ramp at the beginning for a lot of saved headaches later on. I still don't use have of this stuff, but I recall recently debugging an issue that was a couple of clicks into an SPA (in an interview no less) and I sure wished I had jumped through the hoops to get hot module reloading on that one! Likewise, I'm kind of ready for tests now, to make sure the rendered HTML is what I expect in various scenarios, rather than walk through the app like a paranoid moron after every change.
I think the way to stay sane in modern front-end development is to not overthink the tools - if your project is already using them, great, use them the right amount for your team. If you don't use, say, Redux, but your state management is causing lots of pain, hey, maybe it's time for Redux now. If you already write only beautiful ES5 that runs correctly an is easy to reason about, who needs Babel.
My favorite situation is when somebody else picked the tools and the build process and I don't have to give a crap about evaluating all the options. Knowing and caring about all of the pros and cons IS a lot of bandwidth!
This is coming from someone who started out in the front-end with JS and SPAs. Why add in the responsiveness (and overhead) of a SPA when your users won't really notice the difference?
- 1.1 MB of assets (pretty much all of it scripts).
- 300ms of JS execution (on fast desktop in Chrome 63)
- document.querySelectorAll("*").length = 742
in terms of perf, i'm sorry, but this is nothing to be proud of. even a fast vdom lib without SSR can do this in 10% the payload (or less) and 20% the scripting time (or less). why people are amazed at the speed of this impl is rather odd to me. i suppose it works well to demonstrate a large stack, but what you end up paying for that stack is plainly evident here.
I personally build 80% of my apps using Rails or Django, with vanilla coffeescript or javascript + a few libraries for graphs. That works for most business applications, and pretty much any information app (which IMO are usually the most valuable by dollars)
Rails is easier for most people to develop in (even if they are framiliar with JS or Django) within a week of starting. Usually after a 5 minute discussion I can show them why it's amazing and everyone switches over.
However, at a point, react and the like becomes easier when you want a prettier customer facing application, or when you want only part of the web page to update in real time (Personally, I can do this in Rails, but many find react easier). Point being, in those cases a framework that's literally built for that might be easier, and that's when it should be used.
Honestly, I think this whole JS craze is annoying, and I try to run no javascript. Because of that, I might be more sensitive than most and I design websites accordingly (only using JS as necessary)
Flow? Unfortunate that it's a separate thing but it gets types which will relieve most programmers.
Yarn? .NET has Nuget, Ruby has Gem, Go has dep, etc.
Passport? I sure as hell don't want to do authentication all on my own.
Jest? I mean, you don't have to test but everyone should.
Best examples are pages with complex & persistent state (like a chat window open on Facebook as you navigate around).
https://www.howtographql.com/ is a very comprehensive tutorial demonstrating how to use GraphQL with different frontend (Vue, Ember, React) and backend (js, java, elixir, python, ruby, graphcool) technologies.
The end result is not as feature-complete as this entry though as the main focus is on learning the technologies.
Also, a shameless plug: If you want to get started with GraphQL, check out https://www.howtographql.com. There's a basic introduction to GraphQL and many hands-on tutorials, e.g. for React with Apollo or Relay (including videos).
I'm curious about the second part of your statement, the compiler being too sensitive for your server side endpoints, can you elaborate on that a bit?
The second big issue was that relay modern decoupled the networking logic, so having a series of chained queries is something you either have to handle yourself or write your own networking logic for. And because creating dynamic queries can't be done with relay it's a bit annoying that this logic isn't handled automatically for you.
My biggest gripe happens to be the lack of flexibility around dynamic queries. You can't define the data you want back at run time, which feels defeats a major point of GraphQL.
Apollo supports dynamic queries with higher-order components. It might be a case of the grass looks greener on the other side, but I felt as soon as you try to do something bigger than a todo app with Relay you start to feel the restrictions.
Fantastic resource! The key point of this is it provides working example code to start a project from. Everything is MIT Licensed so that means you get to download it and just start hacking.
I would love to see some accompanying video tutorials!
Of course people are also free to post their own examples using their own favorite technologies. I would love for 'HN Clone' to become a benchmarking standard for comparing stacks. It seems like a bit too much work for such a purpose, but that's the point! :-)
I don't think authentication is actually working in the demo, by the way. Don't be tempted to actually log in to your own account! But I tried creating a new account and it failed.
Besides GraphQL you mean
There's a massive bug where if your graphql outputs 1 small error, the whole component does not get loaded.
Imagine a components that lists a bunch of items. If 19 returned correctly but 1 has an error, the "data" prop doesn't get passed, all you receive back is the error prop. It's pretty shitty for that.
https://github.com/apollographql/react-apollo/issues/1112
Learned this the hard way :(
So, yes, you know which one is bad, but it's not a query that's failing, it's something inside the result.
Even though I suspect the times taken between real hacker news and the clone to load comments are similar, real HN feels a lot faster because the page only loads once the content is ready, rather than loading the page skeleton and then blinking the actual content when loaded.
Edit: Found this: https://www.youtube.com/watch?v=h16qMJ_LCyg
Has anyone used their iOS client with production apps?
I don't think it should be on the list.
In addition, it feels slower than regular HN. No actual benchmarks though.
Checking https://github.com/clintonwoo/hackernews-react-graphql/blob/... it might be on purpose though. Toggle code is commented.
Besides, the comments folding does not work (tested on iOS).
npm install
added 1226 packages in 52.23s
2) $ ls -l node_modules/ | wc -l
900
3) $ du -h node_modules/
180M node_modules/
4) Few people care what your web app is built with. Your differentiating feature should be how it's used, not how its put together. In this case, it works exactly like the original, just with different internals. I don't see the point. $ loc node_modules/
--------------------------------------------------------------------------------
Language Files Lines Blank Comment Code
--------------------------------------------------------------------------------
JavaScript 15594 1681330 215398 283955 1181977
Markdown 1831 246795 70260 0 176535
JSON 1602 167050 936 0 166114
HTML 86 37383 2511 18 34854
TypeScript 289 26943 2895 5929 18119
Jsx 19 11072 1190 587 9295
XML 14 6533 466 4 6063
C/C++ Header 21 7103 1126 403 5574
YAML 280 4697 199 166 4332
Makefile 69 3611 733 844 2034
C 4 2219 271 291 1657
CSS 19 1428 84 34 1310
Plain Text 81 1412 285 0 1127
CoffeeScript 16 1260 181 106 973
Python 6 926 179 56 691
ActionScript 4 904 89 227 588
Autoconf 2 778 70 256 452
C++ 7 423 70 36 317
Handlebars 6 262 30 0 232
Bourne Shell 13 213 27 22 164
OCaml 1 156 23 0 133
D 6 69 0 0 69
Batch 2 18 1 0 17
Lisp 2 12 0 0 12
GLSL 2 11 2 0 9
FORTRAN Legacy 1 0 0 0 0
--------------------------------------------------------------------------------
Total 19977 2202608 297026 292934 1612648
-------------------------------------------------------------------------------- $ npm install --production
added 718 packages in 36.292s
$ ls -l node_modules/ | wc -l
518
$ loc node_modules/
--------------------------------------------------------------------------------
Language Files Lines Blank Comment Code
--------------------------------------------------------------------------------
JavaScript 11122 1006887 125608 173965 707314
JSON 907 110272 661 0 109611
Markdown 1017 142126 41008 0 101118
TypeScript 279 23007 2381 4870 15756
C/C++ Header 20 6835 1082 319 5434
YAML 178 3140 137 105 2898
XML 2 2192 233 2 1957
Makefile 46 2431 466 612 1353
HTML 18 1280 40 11 1229
CoffeeScript 7 1133 157 86 890
Plain Text 33 505 118 0 387
C++ 7 423 70 36 317
CSS 8 267 26 6 235
Autoconf 1 389 35 128 226
Jsx 8 228 26 14 188
OCaml 1 156 23 0 133
Bourne Shell 5 100 9 12 79
D 6 69 0 0 69
Handlebars 2 50 8 0 42
Lisp 1 6 0 0 6
Batch 1 2 0 0 2
C 1 0 0 0 0
FORTRAN Legacy 1 0 0 0 0
--------------------------------------------------------------------------------
Total 13671 1301498 172088 180166 949244
--------------------------------------------------------------------------------- Why count Markdown files?
- TypeScript, CoffeeScript, C/C++ headers/source files, basically any language other than JavaScript, are probably part of that modules source, and not the compiled JS output that is used.
Still a good million lines though.
Willing to pay for such a theme.
I think the added value of HN is:
- Algorithm used to rank news
- Community and culture
I agree the community (with all its flaws) is far more important than any tech used.
Also, the comment collapse button doesn't seem to work.
Is there really any benefit to implement HN as a JS driven web application? IMO, Hacker News is simple enough such that implementing it as JS-Driven app actually makes it slower than the original.
Perhaps a good alternative to the flicker effect would be to put loading indicators in place of each UI component as it loads so you can see the structure.
HN is so simple that introducing such defects in the name of some theoretical abstract benefit shows a lack of care about the user experience and the product.
I see this attitude all the time and I find it unprofessional.
It would be better if the components showed a waiting indicator during the download, if the components are responsible for the download. Alternately, the page itself could show that loader before mounting its sub components. Actually, there are many "alternately"s because frankly this site's approach isn't ideal.
However, I love the advantage of exchanging the most minimal amount of data between a browser and the server for each additional page load (ie: JSON response).
Wow.
Proof? :)
All these actions are a lot faster on that site than on real HN with Firefox Nightly on PC and mobile.
Real Hacker News: 6 requests | 11.5 KB transferred | Finish: 336 ms | DOMContentLoaded: 325 ms | Load: 343 ms
Hacker News is just a static cached webpage
The limit is in the immense latency for opening a new connection.
Opening a network socket with TLS can take over 4 seconds on mobile, which is why ideally you’d even use an open websocket for communication instead of new GET requests. Navigating to a page in this version causes one or two network requests, in the real HN it takes far more just to check if the images have changed, the CSS has changed, etc.
Using React to build a simple (student level) site like HN is overengineering too in my opinion. If you don't want to reload the page when navigating you can just reload HTML content with jQuery or fetch() without using client side rendering, without draining device battery and without loading 1 Mb of Javascript that will take much more space in RAM.
SPA approach should be used for interactive applications able to work in offline mode, able to provide rich experience on mobile devices etc.
https://github.com/clintonwoo/hackernews-react-graphql/blob/...
I'm trying this in Firefox nightly (stylo enabled, webrender not).
Even on an old phone over throttled internet, that page is faster than real HN for page transitions, but the same happens on 100Mbps WiFi with my Nexus 5X, or on LAN with my desktop.
In all cases, that site is significantly faster than real HN in loading and rendering (I can see the real HN's icon slowly load on every page refresh — that page doesn't do that).
If you get different results, you're probably using Chrome, which is over-aggressively caching.
Not sure how to explain that difference.
This kind of feels a little over-engineered.
Edit: To downvoters, care to explain why I am wrong?
The project is extremely well done. It is impressive in its documentation and execution. But in the real world done, launched, and with traction is more important than being well engineered. The existing implementations of upvote style sites work just fine.
When Digg.com -- one of the first upvote style boards to go mainstream -- launched in 2004 dozens of copy cats came up. Few succeeded. Why? Community.
Also, we are moving away from the decentralized web, unfortunately. For instance, few people run their own Wordpress blog now. Most use Blogger or Medium. I don't see how projects like this fit in in that world.
Now, if rather than be a HN clone he added the ability for people to create and moderate their own boards. That would get my upvote.