Cellular Architecture was largely a reaction to the S3 outage [0]. I agree that one is still bound to fail due to unknown unknowns or unpatchable known unknowns, but reducing the blast radius [1] to not be globally unavailable [2] is a step in the right direction.
[0] https://www.youtube-nocookie.com/embed/swQbA4zub20
[1] https://blog.acolyer.org/2016/09/12/on-designing-and-deployi...
[2] https://blog.acolyer.org/2015/05/07/large-scale-cluster-mana...
‘Cellular architecture’ is how anyone not going down during their prior outages was doing it for over a decade, just not cleverly branded.
Good links, showing base ideas getting published half a decade ago. I’ve seen use for at least 15 - 20 years, pre-dating ec2 and AWS.
I don't advocate for ever-more-complicated solutions as a rule. e.g. I think multi-cloud setups are probably way more trouble than they're worth for most companies.
I certainly agree that graceful degradation where possible and not too expensive is ideal. For example, if S3 is having problems in one region, being able to fall back (gracefully degrade) into read-only mode might be a nice thing to have.
(In this particular case having a secondary region also probably helps with disaster recovery, which is pretty much mandatory in B2B, for better or worse.)
We also run our underlying Content Delivery APIs in two AWS regions so this was a logical extension.
If the added complexity is worth for your use-case can only be decided by you and I hope the article provided some guidance around that vs. just being a copy & paste gist.
Source: I work at Contentful.