We 30x'd our Node parallelism
blog.plaid.com
blog.plaid.com
My interview screening question was pretty simple- "Is node.js single threaded or multithreaded?" And to most, they spit back the blogspam headline- "Single threaded!" I think the most correct answer is "its complicated" but would accept that because most people would say that is the "right" answer. So I would follow up with- "what exactly happens in a default installation if we have say... 5 requests come in at exactly the same time to just return some static content from disk?" (Node's default threadpool is 4). And here is where you could see their understanding just fell apart. Some would say they would be handled entirely synchronously, others completely in parallel- but then had no idea what the cause of the parallelism was. Very few actually understood that node is an event loop executing javascript backed by a threadpool for async operations.
Before reading this post, I was like eh this is a waste of time- its typical medium bullshit- they almost certainly found they were doing some blocking call in the event loop and then removed it and voila, 30x speedup. It was interesting because it was a lot worse! They spent all this time and hard work figuring out everything but what was taking so long in the event loop, and it seems that was the last place they actually looked.
Anyway, node can be a highly scalable platform (https://changelog.com/podcast/116) but you need to understand it or else it will bite you in the foot. When I was last doing this stuff, upwards of 80% of our time was being spent essentially just JSON.parse()'ing, and we were looking to move to protobufs to avoid that.
Ideally you want to be yielding back to the event loop at least every 1 ms. Anything that takes too long without yielding will show up as a latency delay before your code is able to start handling a new request (technically a background thread in Node.js will pick up the request, but your code won't start executing in response to it until you yield back to the event loop again).
To be honest the more difficult thing to diagnose sometimes is event loop overburdening. If each of your execution spans are taking 1ms, then you can only do a max of 1000 of them per second (assuming there was no delay between executions, but there is). So if you are trying to handle a large number of requests per second the event loop may end up with say 1005 execution spans per second that it needs to execute to handle that request volume. Because you can't do 1005ms of work in 1000ms the extra work will queue up.
So gradually you will end up with 5 backlogged execution spans stacking up per second. Each second you will get 5ms more latency. The overall request latency will just gradually increase and increase as work gets further and further delayed in the queue.
Overall I just think of Node.js as a fancy CPU scheduler. As long as you give it even, decently sized chunks of work to schedule, and you don't give it too many to schedule you will be fine. Anyway I'm a huge fan of Node.js but yeah its easy to fall into some gotcha's if you don't study how it works. The simplicity is a bit misleading
Disclosure and bias: I work on Node core and always hear ranting about incorrect usage in async_hooks in anyone but Elastic APM in core meetings. I used both products and have no affiliation to other companies.
Do you know what they are doing differently with async_hooks in Elastic APM?
To be honest that comes with a lot of Node.js experience though. You probably need flame graphs to start out with, but eventually I find New Relic traces to be all I need, as I already have a sense for the relative CPU weight of the various calls inside a trace span.
This is true, and that JavaScript is mostly a synchronous programming language with host environments that can provide asynchronisity.
A caveat though is that the most important part of I/O is network I/O (tcp/udp sockets) and Node uses real async operations there rather than a threadpool.
FS is just really hard to get right in a cross platform way and that's why it's on the threadpool. Some other stuff like dns is also famously on the threadpool but tcp sockets are not - it's a big part of why Node is fast.
I'm curious as to why. For large scale applications like this, you have other options that offer higher performance ceilings, have more safety and correctness features, and are likely more productive as well. What is the attraction to node?
A guy has to invent a scripting language for browsers in 9 days -> he decides on a lisp -> management says no it has to look like java -> he comes up with something -> its dynamically typed -> lets run a huge banking infrastructure on this
wat
You can achieve safety and correctness features for node via good lint rules and typescript/flow.
Why would you duct tape hacks on top of hacks to achieve the result you want instead of using a language that has already all of the functionality built-in?
However, if you haven't taken a look at Joi- https://hapi.dev/family/joi/?v=16.1.8 you probably should. To me this was a very happy medium- you can enforce your "types" at the point of ingress and egress at your api, and even do validations there, while still being able to enjoy the flexibility of type safety within your code.
Don't get me wrong, my favorite language is Rust, but pretending the JavaScript ecosystem is unusable doesn't make you cool. I can be extremely productive in TypeScript.
What I consider hacks is the hundreds of different tools and dialects of JS that are used to bend JS into doing something it wasn’t really designed for.
JS is fine in the browser, and TypeScript is also fine there because you don’t really have the choice to run anything different (although Web Assembly might change the game soon). But on the backend you have the privilege to pick between dozens of different languages that are better suited to the task and support the features you want out of the box without layering hacks on top. Why not just go with one of those?
This grants us a magnitude of benefits, like being able to do server-side rendering and share code between the client and server. It also means we only need to hire people who can write JS, instead of JS and something else.
TypeScript, Flow, and ESLint aren't hacks. These are mature tools used by some of the largest, most sophisticated engineering teams in the world.
Then, you take a look at node. You look at a getting started tutorial. Its javascript on the front, and on the back. The JSON in between is "native" and is convenient and easy to use, easy to read, lightweight, and just makes a lot of intuitive sense- especially when I had found myself neck deep in XML in previous jobs for the same tasks. I had a nice looking HTML5 web app running in a few minutes- my mind was blown. Then you take a look at the frameworks- express and hapi, and the vast module ecosystem- and how easy it was to build a simple CRUD website with leveldb, or mysql, or really an endless array of storage options. And people were using those options! It wasn't just the bog standard RDBMS being used every place, with your only real choice being mysql, postgres, or if you had money, Oracle. Building endpoints with routes in these frameworks made your code so easy to divide up along clear lines, and there just wasn't the endless miles of boilerplate/scaffold code, and ugly syntax and type systems to fight with and plan ahead of. Things Just Worked. Turning around a code change was a matter of seconds, not a minutes long build process- I had never felt so productive- and writing code was fun again! Deploys were easy, restarts were fast. Rollbacks, when necessary, were painless. There was a plugin/module for everything (too much in hindsight).
Now, this was 6 years ago. Go was around, but still kind of a blip on the radar, Ruby/Python were probably the closest real contenders. Ruby had lost steam, I honestly took some cursory looks at it, but it didn't seem to have traction. Python, suffered from its single threadedness and GIL, and its popularity with the ML crowd- Flask and such existed, but was pretty rudimentary compared to what Express/Hapi were offering, and no one seemed that interested in those projects. I like Go a lot, and for a pure backend service, it might be my go-to today, as one of the original arguments for Node was "its the same language on the front and the back, no more delineation between FE and BE developers, anyone can jump in and fix the bugs!" Which, along the lines of my original comment, don't really work out in reality, at least not on larger systems. People drawn to FE work usually have never done real systems development and don't understand how things work under the hood- which isn't a problem, until one day it is and then its a huge one.
The dynamic typing argument... is somewhat valid, but I found that enforcing api contracts with hapi/joi gave you the equivalence of type safety at your interface borders, while still giving you the flexibility of dynamic typing within your code. In fact, Joi went even farther than just type checking, it could check that your int was within range for the field, that your dates were formatted properly, etc... In mega large codebases, this will come back to bite you, but I found the plugin architecture of Hapi really discouraged that kind of crap from leaking in and it was easy to build truly modularized code.
The performance ceilings aren't that different, and not that impactful, at least not until you get to FANG scale, and I mean literally only FANG scale. We were running a billion dollar business with on 8 fairly small VMs for the API layer, which handled all of the ecommerce transaction handling. I remember at one point we encountered a memory leak of some sort in node, and the instances were falling over and dying about once an hour, but restarting and recovering- this was causing a few % error rates to our customers. I was insistent that we get all hands on deck to figure this out ASAP, and our head of Ops type person said "kevstev, we can throw hardware at this problem to meet SLOs until you get it under control. Your monthly server costs are less than my studio apartment cost me per month in Jersey City 15 years ago."
You just have to have a basic understanding of whats going on at an architectural level, something a few hours of doing the right reading and experimenting can get you if you have the proper background. The number of gotchas to avoid to get that performance were an order of magnitude, if not more, fewer than in a language like C++ (Which I feel has actually gotten so complicated and difficult to grok its become a parody of itself- and I say that as someone who used it and adored it for 15 years).
Yes, there are a number of footguns, but that's true of any language and platform. You can do stupid things in any number of platforms and languages. I don't even see anything particularly egregious in TFA for that matter.
There are fortune 100 companies with systems handling hundreds of thousands of requests per second on a couple dozen servers in Node.js ... is it absolute performance for CPU intensive operations, not really, does it handle more simultaneous requests than everything else, not even close. What it does offer is a really good mix of good enough performance with unmatched developer productivity.
I'd love to see any data or case studies that claim the opposite if you have any.
> Is a 15% reduction in bugs making it all the way through your development pipeline worth it to you?...
https://www.reddit.com/r/typescript/comments/aofcik/38_of_bu...
> 38% of bugs at Airbnb could have been prevented by TypeScript according to postmortem analysis
I've never seen a number far outside of the 15-30% range.
In my experience, most bugs are operator error. Developers didn't code for branching paths that should've been accounted for, etc.
Personally, I'm a fan of TypeScript. Just don't expect to remove the majority of your bugs via its usage. The old "no silver bullet" adage.
Being used to Node, I was flabbergasted when writing C for Linux* . The file system commands just leave my thread hanging while the result is being generated, if I use it on a network drive it might hang for a minute before timing out, so I have to make a tread for each file system command, solely so that it can stall without bringing down the whole application.
* I have no delusions that Windows is any better, Linux is just what I have first hand experience with.
This would be true if you hired someone to write a server in C about 15 years ago. It's not true today. And I hope you're not putting a naively coded server like that in production, or at least doing the hour of research once you notice it's awfully slow to solve the problem.
Like if you wrote your backend in Go, Rust, Java or any number of languages (even C/C++ with common dependencies!) and did a little reading while you designed it, this issue wouldn't exist.
It's only tangentially related to your question, but I can't help but ask this question: why people use JSON instead of protobufs at all?
I'm mostly a client-side developer, and most of my server-side experience is in hobby projects; still, I always used protobufs and loved it. They never damaged my feature velocity, apart from an hour to set up the build system in the beginning, and type safety helped me quite a few times when I forgot to sync changes in protocol on client and server side. Are there some secret advantages of going with json that I don't see because of limited experience?
There is some friction to them though, and I think a lot of it is that most tutorials and beginner books like to keep things as simple as possible, and people start their little project, it gets traction, and then they figure out they need protobufs but now its hard to introduce. In most projects, even today, it seems that its the version 2.0 that gets protobufs, v1.0 keeps JSON for simplicity, unless you have a bunch of seasoned devs involved.
Another one is: what happens when a node process completed execution?
// node ex.js
function foo() { // something async here }
foo()
console.log('bye...')
This is a fun question to discuss (I think some consider this a bug in node).The first thing that really bothered me about our use of nodejs was no one could say why stuff would fail in production. So many moving parts. One of my team members figured out some edge case interactions between nodejs and nginx (used for HTTPS), which I would have never figured out on my own. It wouldn't have even occurred to me to look there. But other crashes, caused by apparent leaks, were mystifying.
The second, and bigger, thing that really bothered me about nodejs, and expressjs in particular, was the notion of back pressure is completely missing. If it's in there, I couldn't find it. So our endpoints were still accepting new socket connections without processing responses from backend services (eg redis, other nodejs endpoints, auth services), which would either zombie or ABEND those backends. And no one could figure out why.
I only understood what was happening because I'd already been through all that "architecture" madness a decade earlier with Java services.
I guess what I'm saying is while I LOVE nodejs' closeness to the metal, I didn't like going back in time 10-15 years.
Also, npm is crap.
I don't even think it's a terribly bad thing to do assuming it favors feature velocity.... but at that point, I'd recommend moving away from Node towards something like Python. And if you wanted to dip your toes back into async plumbing land, explore Go or Elixir.
God knows they could be waiting for some reel to reel tape to spin up somewhere...
I don’t buy it.
Async is nice - if you can handle it. But this is not easy to do in complex systems and processes. It is certainly easier to work with an old-fashioned process that blocks when waiting for whatever you need to wait for, and just scale by letting the OS run lots of those in parallel. Sure, it's less efficient. But it's easier for the devs to handle.
I just read the hidden undertone of this article as "our devs aren't that smart after all".
Taking a wild guess: Some of their bank integrations probably require browser automation. If you're doing browser automation, the best tool for the job is (currently) Puppeteer, which runs on Node. There are other third-party language bindings for the Chrome dev tools protocol, but Puppeteer is developed by Google as a first-class citizen alongside Chrome.
It's really just bindings for the dev tools protocol.
Half the GitHub issues result in "well the protocol requires X and we can't change that".
Pupeteer is popular because it's web automation protocol bindings for a web language, not because it a sophisticated layer or does very much.
There are literally dozens of language bindings for the protocol. [1] Some are quite good and widely used, for example chromedp (Go bindings). [2]
[1] https://github.com/ChromeDevTools/awesome-chrome-devtools#pr...
FWIW, I reliably have 6 puppeteer/chrome instances (headful, even) going on a single box and it's not even at half capacity.
We do use Go for almost all of our other services, and there are an increasing number of integrations written in Python. But we're still using and investing in our Node integrations code for the foreseeable future, and this was an important step for simplifying our infrastructure.
We certainly hope the tooling and rollout process in the post were instructive for anyone using Node, even if their stacks were pristine from day 1 and never need this sort of complex migration :)
I have never seen a good argument for using golang for business logic. If you are writing the actual server then sure, use golang. If you are writing some high-speed network interconnect, use golang. Some crazy caching system, sure use golang. The public WS endpoint, use golang.
But if you need to access a DB with golang for anything more than, like, a session token, then you made the wrong choice and you need to go back and re-assess.
Elixir is in the "germination phase" and I predict massive adoption in the next 5 years. It is a truly excellent platform, every fintech company I know at least has their toe in the water. Everyone I show this video to [1] just says "well, shit."
> Each Node worker runs a gRPC server
Not going to lie, this kind of surprised me. When I think of a Node backend I think of ExpressJS. Not because I think Express is better, but because it's been pushed around in the past few years as the fastest, simplest way of running a backend.
Yet, if you're going to be running a gRPC server, why not use a more performant language with better multithreading support? I thought this article was about them optimizing a grandfathered-in solution (such as Express), but I can't tell why they built out a gRPC server in Node in the first place.
With perfect hindsight, it's a fair point that all the pros and cons could net out to another language being best for our integrations. Integrations are the largest and most quickly-changing codebase at Plaid, so such a migration would be a massive undertaking. We definitely didn't want to block scalability improvements on doing a language migration.
I can't be the only person who reads stories like this and wonders how they arrived at that solution in the first place?
Failing to scale because their previous approach to scaling was a worker per request, a model which was roundly moved away from, because that's how CGI and Apache modules worked and it didn't scale well.
I thought one of the key selling points with Node was an fully async standard library, enabling better scaling in process.
But then you read stories like this, and I find it hard to relate to the original problem.
We still have an event loop that is trivially blocked by very simple programmer errors, destroying the whole advantage that you describe here.
The fact that Node ships a fully asynchronous standard library doesn't in any way fix the fact that Node is a runtime for a language that itself is a mistake.
So they fixed the issue that some requests blocked... by making all requests blocking.
There is a massive deadlocking design mistake in the centre of the language - literally a huge red button with DO NOT PRESS printed on it. Thousands of programmers pass it by every single day, or hour, or minute, and the creators of the runtime insist that it is impossible to fix that button whatsoever; instead, all users need to work around it by ensuring that their code in no way presses that red button on purpose or even by accident.
These people insist that it is impossible to program normally and in a language that is actually sane and does not advertise obvious and gaping design mistakes as "features of the language". These people advertise the analogue of Python's Global Interpreter Lock as the core foundation of their language.
These people advertise Node and the language it implements as practical for implementing multithreaded applications. Posts such as this show what sort of bullshit it is; it is only practical to use Node for parallelism if each single Node instance is only ever run single-threaded. You don't parallelize by running multiple threads, you parallelize by running multiple Node runtimes.
This is no longer an act against productivity or usability. This is simply insane and shows one of the most basic things that are wrong about Node's language and approach. It is impossible to write a multithreaded program if your language of choice makes it trivial, and practically unavoidable, to globally lock your whole runtime with every single line of code you write and import as your dependencies.
The first Node.js service I wrote and maintained, processed thousands of requests in parallel and was successfully in production until the company it was developed for ran out of money.
However, running multiple containers for parallelism sounds a little bit crazy. In the worst case, each container may be running on its own server, but even assuming multiple containers per host, I'm guessing they were running an insignificant number of instances, which is probably why they were able to save $300k in server costs.
> We were running 4,000 Node containers (or "workers") for our bank integration service.
All I'm saying is that the choice of language is (usually) not the issue. Poor architecture design causes a lot more problems than whether you choose python/java/ruby/node for your webapp.
Some features would become unusable, like Promises and async/await, but those would be worthless in such a design anyway.
As much as I'd like to see it, your JS without an event loop is a purely theoretical construct so far.
I wouldn't implement it for the HTTP server/backend use case, because I find the blocking/share nothing architecture of PHP limiting.
The point is the language doesn't prevent this kind of architecture.
No different than having a dedicated threadpool for asynchronous programming on the JVM.
Yes blocking the event loop is easy. No it's not THAT easy. I've never done it because you think about it while writing code. It's part of the environment. I have had to fix lots of reports etc that try to load up the world and iterate through it in a loop that doesn't yield to the event loop. It's possible to never make that mistake by understanding your environment (just like how managing pointers in C is "hard").
To be pedantic, all JavaScript functions block the event loop. It's just that the vast majority of functions execute so quickly that the amount of time your function blocks is very short.
I once had a loop that processed tons of data and would block the event loop for 1-3 seconds. I ended up solving it by writing an asynchronous loop which used promises and process.nextTick() to make each iteration a separate execution block on the event loop. But I've only had to do something weird like that once in 10 years of node development.
At one company I built API middleware that calculated time spent in the event loop vs waiting on dependencies (other apis, Redis, etc) by managing a couple counters per request. It was very helpful.
As someone not well-versed in js, could you describe one such case? Concurrent access to a global from two threads? Mutexes? My background is more with systems languages and I have done very little js for the browser, so I do not see that big red button.
If your function blocks, for example, by performing a wait on something without yielding, then nothing else gets computed because of the wait, but the event queue is occupied since the function has not yielded. This breaks the cooperative part of the cooperative multithreading mechanism of Node.
If the event loop is part of the Node language/runtime, I can see the case for making it preemptive.
All this has happened before. All this will happen again.
In Javascript world I guess you could do the same by replacing time.sleep with some CPU bound code. eg a big calculation or an infinite for loop.
I'm working on a project now that's very CPU bound and using limited workers behind an MQ as a distributed RPC behind a fronting API interface so that I can handle scaling... though it's also a Windows-only library involved. There's definitely an art to scaling certain types of workloads and many different options. Sometimes the simplest solution you can come up with really is the best option.
I can't help but also feel that is also an issue and in another given language this issue might not happen ... but they'd hit another.
It's so easy to say "don't do X because problem Y won't happen" but hard to predict what happens when you move from (language, platform, or whatever) X to (language, platform, or whatever) Z.... and I suspect people often hit issues and realize that maybe Y wasn't the problem.
I see it all the time and I feel like "Wait guies I'm not sure we're fixing the right thing!?!?!"
This article raises a lot more questions than answers IMO.
Can you give an example please?
I think it's much easier to block a thread with C#'s async programming model than node's...
Node can't not be single-threaded. Because Javascript is. Node.js is single threaded. It has a single event loop in a single thread, and all the "concurrency" is simply queued on that loop. It offloads some tasks to libuv for some system-related tasks but that's it. And the thread pool that libuv creates is very limited.
Anything that doesn't end up in libuv (that is, probably vast majority of user code) will only ever run in one thread... because Javascript is single-threaded hence V8 is single-threaded hence Node.js you get the gist.
And of course, node.js even has a separate documentation section titled "Don't Block the Event Loop (or the Worker Pool)" [1] because it's trivial to block the event loop.
[1] https://nodejs.org/en/docs/guides/dont-block-the-event-loop/
https://stackoverflow.com/a/24684037/1026671
"Concurrency" doesn't warrant scare quotes even when describing a single threaded program.
It's true that Node.js allows one task to block many others, but that's an implementation detail of Node.js, not a guaranteed result which necessarily applies to any program using only one OS thread. Other programs suffer from analogous starvation/deadlock/livelock/priority inversion problems, but those would be implementation details too, not guaranteed results from using multiple OS threads.
AFAIK there’s nothing in the JS language that forces a single-threaded implementation. I think the real reasons that Node is single threaded are a) legacy and more importantly b) a lot of existing code would break if events were dispatched in parallel.
Yes, I can. Node.js doesn't do that.
The only thing I can think of where a programmer could "easily" block the event loop would be if they explicitly use sync filesystem or other blocking calls instead of the async API. But in any project with reasonable code review I don't think this would happen.
Parsing JSON. Or executing a regex. Or anything, really, that blocks the thread: https://nodejs.org/en/docs/guides/dont-block-the-event-loop/
There's no magic
In practice I've rarely seen production nodejs applications cause significant CPU blocking issues. Huge JSON parses sometimes, yes, but then it has to be one hell of a payload to cause any significant issue.
Regexes? Have you really seen regexes block CPU for significant time relative to the rest of the application? I'm sure it's possible with a crazy enough runaway pattern, but I've never seen it happen.
I really don't understand what people are getting at here. Nodejs is an async programming model, it's non-blocking by default. Are the people saying it's trivial to break nodejs developers?
> I think it's much easier to block a thread with C#'s async programming model than node's...
Which spawned a discussion in which people seemingly think that you can't block a thread in node, or that since node is async it means it's not single-threaded etc.
Since it is single-threaded, it's quite possible that the original post wouldn't need 4000 node instances otherwise (emphasis mine):
> We were running 4,000 Node containers (or "workers") for our bank integration service. The service was originally designed such that each worker would process only a single request at a time. This design lessened the impact of integrations that accidentally blocked the event loop
Creating a system in the 21st century that tries to follow ideals from the 90s gives us the kind of idiotism that we can witness here.
You are otherwise on the right track, though Node does technically have one advantage, which is that it is a cooperatively-scheduled island in a preemptively-scheluded overall OS. In the 1990s, when the cooperatively-scheduled program was not cooperative, you locked the machine, not the process [1]. There is a reason why Apple went very aggressive with the OSX rewrite; the previous Systems had basically written themselves into a corner where they had to use cooperative multitasking because so much code made use of the implicit promises it provides, yet they could no longer afford to compete with Microsoft if they didn't get off it it, because the complexity just kept going up, up, up and the problem was going to continue getting exponentially worse.
For a Node program, you only have to account for the Node program itself, not everything running on the computer. Still, you're in the same exponentially-growing-complexity trap (with a very initially-safe-seeming low exponent, but it still gets you in the end), you just reset yourself back to a point earlier on the curve.
[1] There are various details, caveats, interrupts, etc, the picture is more complicated than one sentence can convey, but the principle still held and it was still possible to wedge the machine fairly badly for varying periods of time with simple bad code.
It doesn't matter if it's in an event loop or thread per request. Architect things correctly.
I'd bet that the orders of magnitude of speed from Moore's law did in CMT by making PMT doable without a huge speed hit.
Nothing I've seen from async is cleaner, easier to maintain, or better from a cognitive load POV. It's just more efficient for certain types of loads because you're being a consenting adult and not breaking things.
Node's concurrency and parallelism probs don't require deep lang changes but runtime ones and some ergonomics to be closer to Go -- and not far IMO bc of last few years of work for Async, safe (fresh env) eval, and json/buffer msg's.
More exciting would be something like Apache Arrow / Berkeley's 'Plasma', but that stuff is still more exploratory.
https://dev.to/nestedsoftware/is-cooperative-concurrency-her...
This article does a good overview
Still interesting...
But also, though, you have to consider that most places aren't Plaid, and most places developer time is more expensive than throwing an extra machine at the problem.
In terms of what issues caused us to move away from parallelism in the first place, it was all the CPU-bound stuff that you might expect: ReDoS-style issues, post-processing arrays in very large edge cases, programmer error, etc.
But these are not parallelism problems. These are single threading problems, which the core problem with Node.js, not parallelism in general. Hence I think the question stands: why did you choose node for this?
Hire an architect costs what? Putting the genie back in the bottle was a problem Plaid baked into its early success, which is common with startups hiring engineers with zero architecture knowledge.
For example Discord reached out to Rust and built tiny Rust components that are called from Elixir for their server user list. Some servers have 200,000+ people online, and Elixir wasn't cutting it performance wise. Rust, boom now it works.
Here's how it probably worked: they liked Node, they liked containers, they put Node into containers and it worked, and they stuck with it as the user base grew.
> We hypothesized that increasing the Node maximum heap size from the default 1.7GB may help. To solve this problem, we started running Node with the max heap size set to 6GB [..], which was an arbitrary higher value that still fit within our EC2 instances.
Sounds like they were utilizing their EC2 instances very poorly. Why not run more workers per instance, or switch to an instance type with less RAM (or more CPUs)?
> In terms of what issues caused us to move away from parallelism in the first place, it was all the CPU-bound stuff that you might expect: ReDoS-style issues, post-processing arrays in very large edge cases, programmer error, etc.
But it's trivial (a single line) in Node to place breaks in CPU processing to allow the event loop to fire, and as for "programmer error"... many commenters below are also complaining async programming is too hard or finicky.
But that's like complaining about C because pointers are hard, or Java because OOP is hard, or databases because planning indexes is hard.
Once you "get" async, pointers, OOP, or indexes, it's easy. And it's part of your job as a professional programmer to get it. Async is no trickier than anything else.
The setup in the first place makes absolutely no sense to me, using a language exactly opposite of how it's meant to be.
What I don't get is how nobody treats it as an issue when developers coming from Python or Java to C don't get pointers.
The assumption is that you learn.
But for some reason, people think it's "OK" to not get async, that it's the language's fault rather than the programmer's. That's what I don't understand. It's like a different cultural standard gets applied.
I have no patience for persons who don't belong to the discipline.
I agree that it would be nice if all developers were infallible – I'm reminded of a friend describing their company, where "we don't write tests because we all write good code". At a certain point, you have to look for processes – linters, monitoring, testing, language choices [1] – where people can't shoot themselves in the foot. (Code reviews being only moderately less fallible than a single engineer.) It's not enough to just say "be better" whenever bad code is written.
I think when the decision was made (years ago) to handle a single request per container, they couldn't find such a process to prevent event loop blockages, other than migrating an already-large codebase away from Node. As others have pointed out, maybe such a migration is necessary – after all, event loop blockages are still an inherent risk because of how Node works. It's just a lower risk than it was a year or two ago, because we've significantly improved our usage of the event loop, and also have tooling in place to catch blockages before they become an issue.
Yeah, external libraries for Node ought to be designed so that any function that ever might take any length of time whatsoever should always be callable as async. But if they're badly designed or not intended for large inputs, they might not be. You'd definitely need to find another library or write your own there, so I get that.
No you are not. I wonder which CTO would allow this; like everyone here, the exact case is not really clear (or at least why this solution is a great solution for it), but this sounds like a weird solution (and expensive) to some issue. I really don't understand these 'solutions' and I am almost 100% sure I (with a team! but the point that this is not the best solution for the problem) can whip up something far simpler and more efficient for this problem. But ofcourse there are problems that might fit?
For instances where you actually know you need lots of CPU, there are now strategies for offloading that specific work, although they have taken a while to get nice and easy to use.
On a negative note: FOR THE LOVE OF ALL THAT IS HOLY, HOW DID THIS HAPPEN.
The insistence on using Javascript is just beyond lunacy at this point.
The true treasure is Erlang / Elixir's runtime though. The parallelism, the self-healing, the preemptive scheduling.
> Since V8 implements a stop-the-world GC, new tasks will inevitably receive less CPU time, reducing the worker’s throughput
But there is this Google blog post vom January 2019:
https://v8.dev/blog/trash-talk
> Over the past years the V8 garbage collector (GC) has changed a lot. The Orinoco project has taken a sequential, stop-the-world garbage collector and transformed it into a mostly parallel and concurrent collector with incremental fallback.
So I guess they used an older node.js version. The current LTS version is 12.x and it is from around the middle of this year.
---
PS: If the blog author reads this, there is an accessibility problem with the Google-hosted inline images. If I try - without ad blocker - in an anonymous window I see none of the inline images. Logged into Google with my own account I can see some but not all the images. Apparently which images I can see depends on being logged in to my Google account? I also tried IE Edge just to see if the browser makes a difference - no inline images visible there either.
Your client does not have permission to get URL /Iw-RdHoPjbwuSAqJHK3C0Sy8m29NqzeHPtmJ7CVFuYqwr4CbwpGjwn9O4bcDNtCf_hLD4FGc75nkQYnJBgyA-CT2ikBDWQD-nAtqxXa4Lw2yDuh_-ywcsDaer6m4LyVtljwfrajO from this server. (Client IP address: [redacted])
Rate-limit exceeded That’s all we know.
V8 JIT means that things like order of keys in an object or number of different calls to a function might affect whether your function gets optimized.
And there's no easy way to find out if a JS function is falling back to slow mode or to tell the buildsystem 'this is a hot path, don't let me write code that deopts this call'.
Since they provide an API, it seems like some of the calls where they think a user isn't present might actually have one present.
In fact, it sounds like they think "linking an account" is the only "user present" API call:
"Only 10% of Plaid's data pulls involve a user who is present and linking their account to an app"
Plaid's use case here is automating logins, responding to captchas, manipulating those on-screen virtual keypads to respond to security questions, chaining together multiple HTTP requests, and then parsing out frequently invalid, rapidly changing, and just plain broken content from a wide multitude of banking websites.
And things like data types should be strictly enforced, otherwise you can get unpredictable results, which is especially bad when you're dealing with money transfers.
You would have a really difficult life in web scraping. You do not have the guarantees of well-formed data. Instead you get HTML with mismatched tags, JSON with newlines in the middle of strings, content that claims it's UTF-8 but upon closer inspection is actually GB2312, pagination endpoints with off-by-one errors, etc. It's an absolute mess and taking the stance of "well, they didn't encode their JSON correctly, so we're not going to operate on their data" isn't a very effective strategy.
> Which is especially bad when you're dealing with money transfers
Afaik Plaid is read-only. They fetch information from financial institutions and make it available through an API.
That's why I'm so uncomfortable handling things like bank transfers over such inconsistent, buggy systems, which is what Plaid does. It's not read-only: https://plaid.com/use-cases/consumer-payments/
Not to say I don't trust Plaid, I'm sure they're aware of all this and very careful about how they do things.
I have no affiliation with plaid, I've honestly only heard negative things about those guys, I was only empathizing with the difficulties in maintaining thousands of different scrapers and why I felt a scripting language provided far more latitude to get things done.
That is, an mmap based kv store so that if you choose to run more than one node process on a single server, it has a fast kv cache?
I'm aware you can use redis or similar, but a simple mmap kv store is simpler and faster for a single server use case.
If you want a simple open source lib to do exactly that for you and provide an easy to use API, you can use something like https://www.npmjs.com/package/tmp-cache .
I'm aware of the runtime model differences between node and PHP.
It isn't published on NPM (you can use it as a git dependency) but if people are interested I can.
Why don't you publish releases?
> our system is more robust to increases in external request latencies or spikes in API traffic from our customers
Provisioning and deploying with ECS is usually just mouse clicks.
Honestly, the accounting for which would've been higher impact – investing in parallelism earlier, or adding infrastructure and having more resources to devote to other pressing needs – is difficult to do, even in retrospect. There was surprisingly little effort required to get to 4,000 node containers in an ECS cluster, other than deploy speed issues which we talked about in a previous post [1]. But it's possible this migration process would have been easier if we had done it sooner.
[1] https://blog.plaid.com/how-we-reduced-deployment-times-by-95...
What the f*? Of course it would have been easier if you had done it sooner. What you lacked was the willpower from decision-makers who had growth of dollar-signs in their eyes.
You've littered this thread with comments explaining how every move you made was based on ROI. That's the kiss of death for architecture concerns, and bizarrely it puts Node.js on the list of runtimes for data/stream processing backends.
No matter how many times you explain how you made these decisions, I can't help getting the feeling you were wearing horse blinders.
Edit: I find it impossible to imagine that nobody on the engineering team ever shouted, Hey look out! We are basically a Web farm for banking-related requests, this is insane! Surely you've heard from those people and they were let go.
Different companies make different decisions when weighing ROI against architecture concerns. We're heavy on pragmatism and impact at Plaid, so it's quite intentional that we don't fall all the way on the latter end of the spectrum. I appreciate the discussion in the comments as to how effectively we are balancing these two concerns – certainly this is an area where reasonable people can disagree.
This blog post gives me the impression that either Plaid is filled with either junior or incompetent engineers - to scale to 4k containers serving 1 request each for an API workload is absolute insanity.
These engineers are building stuff for banking. Banking!! There is literally no way I'm going near Plaid with a very long bargepole after reading this.
It I was someone senior at Plaid, I'd be pulling this blog post before it harms reputation any further.
Instead the rational thing to do is build something quick and dirty and optimize later, and that's exactly what they've done.
The difference here is that what they did wasn't even the simplest thing - it was a crazy, insanely wasteful thing that just happened to work for a while. Being honest, for me, it's an indefensible approach.
> Why should they worry about $100k or whatever when they're funded for > $350M? Their bottleneck is engineer hours, not dollars
Arg, but this rubs me up the wrong way! Any half-way competent engineer could have built something simpler and much more performant, and likely in many less hours too. Sometimes stopping, thinking and discussing for a few minutes or hours will save numerous hours. I mean, how many hours did they spend on this "diagnosis" alone?
Their bottleneck was software being able to scale past a hard stop. I guess having a known breaking point of scalability is a good thing? But building things in a way where you either have to overhaul your development runtime or not be able to scale past a certain point is pretty terrible.
It seems like the only reason they did this was because they really felt the pain of it from the business and dev side and they were lucky enough that they had traffic spikes to raise these issues. If they had more consistent day-to-day traffic then this would have just hit a breaking point one day and they would've been fucked until it was fixed.
OTOH, I do still feel this is so bad they need to be called out on it, and it really does scare me off using them. Given they're being transparent, it boggles the mind that they're tried to justify this, rather than just owning it, admitting it was the result of letting a junior do some resumed-driven-developlemt (or however it came about).
To put a positive spin on this (perhaps a first for me in this thread!), I plan on doing something of an internal post-mortem with my team, where we'll look at the deficiencies of this design, try to reason about how on earth it came to fruition, and critique our in-place review processes to make sure something like this never happens to us.
Did you consider the likely (and more charitable) explanation that they were aware their design was "bad", but had higher priorities until now?
If I were you, I'd be pulling your comment before it harms your reputation any further. :)
I think marketing and VC valuations grew them into a multi-billion dollar company; whether they remain so, to a large part relies on how fast they burn through VC cash - so, not looking too good on that front...
No even half-way competent engineer would come up with such a complex, unperformant solution to a simple problem - I think a higher priority should be hiring engineers who actually have a clue what they're doing.
As for meeting business requirements... while this might have worked for a while, it was plainly not a good way to meet them, and given Plaid are in the banking sector, really doesn't bode well for the future (I'm having flashforwards already to security breaches, plaintext passwords etc...).
WeWork is a "multi-billion dollar company" in the same way that Plaid is. Private funding valuations don't really mean anything anymore.
Worse is- they never really explain where that 30x improvement came from- or if they even understand it themselves? They talk a lot about getting their memory issues under control, but hardly at all about actual parallelism- and it seems that even then they confuse it with merely speeding up operations that are blocking.
I kind of expected this post to be "We did a whoops and had a blocking call to a DB/fs/compression call/whatever. This was all happening in the event loop and not being farmed out to the threadpool by libuv. We fixed it and now look like heroes to our CTO!"
Though I'm somewhat surprised they didn't use Worker patterns per node with self monitoring for health above and beyond what they already did.
I can guarantee you a VPE or CTO who can say they helped do that... but ran into a scaling issue from their success will have no issue with employment and no reason to be ashamed. All the more impressive if it was just a bunch of junior engineers.
But you'd go to a competitor who hasn't published a blog post, whose internal code you haven't audited and simply presume is just fine?
In plaid's defense, lack of performance tuning isn't necessarily a lack of security focus.
Come on, this is not about "performance tuning", where you're trying to eek out every last drop of performance - it's about a completely indefensible, complex, wasteful solution to a simple problem.
I'd say engineering insanity at this level is very worrisome for what they've done at the security side of things.
Some people think you can just write software, sell it to customers, and it's "tuning" to make it work properly.
You should be fired from whatever job you have.
My guess is that you have no job, you are fronting USD.
In which case you have absolutely no place in this conversation and you should be ashamed of yourself for speaking up.
A fool and his money are easily parted.
If you don't want to be banned, you're welcome to email hn@ycombinator.com and give us reason to believe that you'll follow the rules in the future.
The simplest solution is to scale to one worker per node initially if you're doing anything compute intensive... once you've done that, and/or you need better performance for any number of reasons including cost, then you can do more. Now, I'm not sure I would have gotten to 4k nodes before I started to re-evaluate parallelism or better scaling options, but the initial implementation is absolutely fine.
I get it, but come on - this was not a "performance optimisation" issue, but one of bad architecture; an architecture that certainly doesn't inspire confidence in the priorities you mention: accuracy, simplicity, safety.
Now, in addition to probably optimizing what they've done, converting to, for example another container system, like K8s where they can scale vertically a bit better may have been another approach.
The biggest issue that I see is gRPC doesn't work that great with Node. You can't use it with cluster which means you have to self-manage threads/processes and it adds complexity there. Yeah, there's definitely issues that come into scaling in terms of performance optimization.... but where they started from isn't unreasonable imho.
That doesn't prevent us from thinking about performance. GTFO with this nonsense.
I would have probably pushed for a shift in orchestration to kubernetes along with some tweaking as an initial uplift. Others would re-write the whole thing in another language. They chose to add a bit of complexity for multiple requests per node support. They all have their pluses and minuses, but in the end it doesn't mean the initial approach was bad, or their refactor wasn't pragmatic or practical.
Dramatic rewrites to a codebase lead to instability and in practice fail as much as succeed.
Otherwise known as re-architecting?
I don't think we've tried to assert that the old system is perfect. We went into some detail in the post about why it took us this far. Certainly, the single request per container approach wouldn't scale if our unit economics were different. We didn't get into this too much in the post, but the Node service sits behind a couple of layers of Go services, so the we had more control over scaling API traffic than it might appear.
Likewise, I hope we didn't give the impression that the new system is perfect. We've explored other languages for integrations in the past (even Haskell, at one point), and are continuing to do so. A migration away from our years-old Node integrations codebase would be a massive undertaking at this point. Absent that, it doesn't seem consistent to say "you're incompetent for handling 1 request per container" and also "you're incompetent for writing this post" – if you believe the former then it makes sense to be an advocate for this project, at least until a language migration can be done.
I think the set of hoops we had to jump through in order to add concurrent requests without adding latency is a good demonstration of why we didn't do this sooner. It wasn't a massive undertaking by any means, but it wasn't trivial. At any rate, we're not really looking for a gold star here – just putting this out there and hoping this will be useful for others who are, as other commenters have put it, building their own "Frankensteins" :)
1. We used a system which uses event loops to achieve great concurrency, but we turned that off because we don't trust it. 2. Instead, we spent $300k/yr rolling out one-process-per-API as though we were using Apache 1.3. 3. We used an arbitrary JSON library without knowing anything about its performance characteristics, which it turns out were inordinately bad
It's not that this wasn't a great exercise in engineering and problem-solving, or that it's not a great demonstration of how to solve scaling problems at scale, those are definitely true. It's more that "we spent $300k/yr more than we needed to so our engineers didn't need to learn how to use our technology stack properly."
I'm not meaning to be harsh, I've kludged enough garbage into production in my lifetime, but more that the fact that you got into that situation in the first place gives a poor impression of either your development team or your development processes.
FWIW, I don't believe Node was chosen specifically for its concurrency – it was just the language chosen for the entire stack by the company's founding engineers, and lives on in just this one service.
In a parallel universe, a barely-competent engineer would have designed something more far more obvious, simpler, performant - all while using less hours, and not borderline-fraudulently wasting substantial amounts of your VC's money.
If the company 10x's again, it won't be because of poor engineering, it'll be because of marketing and VC's who don't know how you're wasting their money. If the company 0.1x's, it'll likely be because of a security breach because of appalling design.
Is that overly negative or just the right amount?
I'd be very interested in reading about that.
In general, I think it'd be great if more companies communicated their engineering practices and war stories externally – I'd love to read such posts myself! Unfortunately, it takes quite a bit of time to write a post (for engineers who are already stretched thin), and it seems that being honest about shortcomings at an early-stage company is an invitation for people to be personally disrespectful. It is what it is, but I imagine that's one reason we see such posts from only a small handful of startups, and the subject matter is often cherrypicked and sugarcoated.
So anyway, thanks for engaging with our post all the same – and maybe we'll be back with a post on architecture reviews in 5 years when all the kinks are ironed out :)
https://monzo.com/blog/2016/09/19/building-a-modern-bank-bac...
Instead, you've peppered this thread with comments that kind-of, sort-of justify the approach taken.
I'm sorry, but this approach cannot be justified - it's overly complex, and far from the simplest or most obvious approach. I'm truely shocked that Plaid has produced an architecture like this, and doubly so that Plaid would try to justify it. My guess here (and given the attempts at justification, this is me being really charitable) is that a junior dev was given too much leeway, and did some resume-driven-development, just so they could say they'd worked with 4k containers.
So how would you have engineered it? I would just send the data uncompressed granted that the receiving server is probably in the same data-center with switches capable of handling Tbit's of data per second.
I liked the article, but would have wanted more details. I love optimizations, it's such a drug, the rush when you make something x times faster. This article doesn't give me a bad impression. Contrary I'm thinking about sending an application.
LOL
Seems like everything went right to me.
I would be worried if the blog post was "we randomly tweaked some stuff and we can't measure it but it's a little better" or "we rewrote it in go and in the rewrite introduced 87 new bugs while fixing 42 old bugs". They engineered a solution, built from good investment in infrastructure, rather than ninja-ing a hack. That, to me, is a very good thing.
A lot of people seem deeply upset that Node was involved, but I think that's a red herring. The problem they had -- allocate a large chunk of memory, keep a reference to it while it is slowly sent to another server, free memory -- is going to happen in any language. (I don't super agree with their solution of "make the server faster" because one day it's going to be slow for some other reason and this problem will crop up again. Instead they probably just need a fixed amount of memory to dedicate to this process and to drop the debug payload when the buffer is full. Or just put it in the request path if it's crucial that it be produced every time no matter what. At least that will apply backpressure to calling services, pop the circuit breaker, and redirect requests to a region where S3 isn't broken. But I don't think the debug information is THAT important ;)
So, yes, horizontal scaling is good, especially for stateless workloads - but that doesn't mean you run the most hopelessly under-performing code imaginable on each node, so you basically have to scale out like this! I mean, seriously, 4000 containers to serve 4000 concurrent requests? I mean, I can't even...
I honestly can't believe the attempts in this thread to justify such an utterly, horrendously bad architecture - there are 1001 better, simpler even, ways to approach this.
Yes, premature optimisation is bad, but optimisation here was nowhere near premature.
When you start a business, you have no idea what it's going to grow into, or if it's going to grow. So you start simple. The design was good enough for there to one day be too many customers. That is huge.
When this happened, they started a second copy of their app, and could now handle twice as many customers. Repeat 3998 more times. Now the toy app is making some real money, so you can afford to deep-dive into the system and fix the technical problems.
They avoided the real issue that kills startups, having a customer call you because they want to buy your service and you saying "sorry, we aren't accepting any new customers right now because Hacker News comments don't like our software architectures."