Provisioning and deploying with ECS is usually just mouse clicks.
Honestly, the accounting for which would've been higher impact – investing in parallelism earlier, or adding infrastructure and having more resources to devote to other pressing needs – is difficult to do, even in retrospect. There was surprisingly little effort required to get to 4,000 node containers in an ECS cluster, other than deploy speed issues which we talked about in a previous post [1]. But it's possible this migration process would have been easier if we had done it sooner.
[1] https://blog.plaid.com/how-we-reduced-deployment-times-by-95...
What the f*? Of course it would have been easier if you had done it sooner. What you lacked was the willpower from decision-makers who had growth of dollar-signs in their eyes.
You've littered this thread with comments explaining how every move you made was based on ROI. That's the kiss of death for architecture concerns, and bizarrely it puts Node.js on the list of runtimes for data/stream processing backends.
No matter how many times you explain how you made these decisions, I can't help getting the feeling you were wearing horse blinders.
Edit: I find it impossible to imagine that nobody on the engineering team ever shouted, Hey look out! We are basically a Web farm for banking-related requests, this is insane! Surely you've heard from those people and they were let go.
Different companies make different decisions when weighing ROI against architecture concerns. We're heavy on pragmatism and impact at Plaid, so it's quite intentional that we don't fall all the way on the latter end of the spectrum. I appreciate the discussion in the comments as to how effectively we are balancing these two concerns – certainly this is an area where reasonable people can disagree.