I say "possible" because our system observability was less mature even 12 months ago. Firefighting 10 different root causes of memory or event loop issues without the right tooling in place would be a nightmare. That's why we did a deep dive into the tooling that we considered to be a prerequisite for this project – hopefully it's helpful for others in our situation.
Different companies make different decisions when weighing ROI against architecture concerns. We're heavy on pragmatism and impact at Plaid, so it's quite intentional that we don't fall all the way on the latter end of the spectrum. I appreciate the discussion in the comments as to how effectively we are balancing these two concerns – certainly this is an area where reasonable people can disagree.