June 25th, 2023 Deno Deploy Postmortem
deno.com
deno.com
> The start-up payload will be cached across the globe
This is pretty interesting. Doesn't seem like the right scale for a small amount of metadata to topple over a DB. That's a pretty small load even if we're talking thousands of nodes.
Just a wild guess but the fact that connections are heavy in pg is a foot gun that tends to rear its head during large scale-up events
From the moment one runs a system in two nodes, one being the application server and the other the database server, that system is distributed. Two nodes is all it takes.
Perhaps I wasn’t clear enough, but in your comment you make it sound that having multiple nodes interacting with a database is an anti-pattern, which is not, the discipline of distributed system is well researched and from the top of my head I risk to tell you that such anti-pattern is never mentioned, and the concept of a database rarely appears in DS, instead it focuses in a more abstract concept, which are the network nodes.
To support my argument, from the moment one decides to replicate nodes, most likely those same nodes will be connected to the same data source, because data replication isn’t as trivial as application replication.
In a Microservice architecture it can be considered an anti-pattern to have different bounded contexts using the same database.
Keep throwing rocks if you or your org haven’t overlooked something in the system you work on.
Typically, it seems that the second major issue is when a company like deno needs together get serious about addressing known, but not profitable risks.
This kind of metastability _should_ be detectable through fairly simple load till failure testing and recovery monitoring. The 10x load cited is a red herring when pulling static config during recovery is their bottleneck
What you don’t know about are all the other attacks that did not impact the service, because the way load testing was performed already lead to some mesures being taken.
Was this a global outage for 90+% of deno deploy hosted projects? “A large chunk” doesn’t give me much faith that it was isolated to any one region
The framework provides a clean way to define expectations and to produce data that can then be visualized in grafana.
What are the best ways of stopping this without using cloud infrastructure these days? I’ve become too used to using the cloud offerings so I’m out of touch.
You protect with a combination of services and some cloud providers will offer cost compensation for resources you might have to create to stay responsive during the attack.
I’m wondering if this was open source rather than proprietary this sort of bug allowing a DDoS would have been caught before production. I know of multiple times in the past security issues were detected by the community and dealt in Node before they reached production.
It must be a challenge for the Deno team competing against the likes of Amazon or Google without the andvantage/burden of open source.
Under the hood Deno Deploy is just Deno with scale built in. It's obviously not a simple soundbite but I recently built the new Deno.serve API implementation and the same code effectively runs in the CLI and Deploy modulo what's necessary to get it working in the Deploy environment.
Load shedding is exactly what would prevent this in the future, scaling up capacity and adding a cache is only going to help buy some time.
Wish I could have a dollar for each time SREs duct tape rather than solve problems with systems thinking.
I believe you are getting downvotes because it's difficult to give a direct solution like you did without having more context on the matter.
> Look at all these super knowledgeable “engineers” down voting this.
Maybe people thought "look at this super knowledgeable 'engineer' shooting a one-off solution just by looking at an RCA"