436 karma · joined November 13, 2011
Previously: SRE at Google Cloud (taking care of GCE and GCP) and Forward Deployed Engineer at Palantir.
https://www.palcu.net/
[ my public key: https://keybase.io/palcu; my proof: https://keybase.io/palcu/sigs/425bZ1Ip-RwE52Uws06rLF54qq21RU2pQQyi8cXICv8 ]
> On March 26–27, 2026, customers experienced elevated error rates when using Claude Opus 4.6 and Claude Sonnet 4.6. The issue was caused by a networking performance degradation within our cloud infrastructure that disrupted communication between components of our serving stack. We resolved the incident by migrating the affected workloads to healthy infrastructure, restoring normal service by 9:30 AM PT on March 27.
> Between 14:17 and 17:11 UTC, our primary application database experienced severely degraded I/O performance following a routine maintenance operation, causing slow or failed requests on Claude.ai and preventing new or refreshed sign-ins for Claude Code and the Console. API traffic via Claude Developer Platform was unaffected.
https://status.claude.com/incidents/jm3b4jjy2jrt
Sorry again, thanks for bearing with us as we're dealing with the influx of new users and scaling up all of our systems.
I was also fortunate to be using Claude at that exact moment (for personal reasons), which meant I could immediately see the severity of the outage.
https://status.cloud.google.com/incidents/8cY8jdUpEGGbsSMSQk...
I hope the pendulum swings the other way around now in the discussion.
[disclaimer that I worked as a GCP SRE for a long time, but not left recently]
The more important question is that they've removed the Activity feed, where you could see likes from other people. Which was like a realtime feed to what your friends were doing on the website. The website is way more boring now.
Personally, I wasn't part this time for the actual mitigation of the overall Paris DC recovery, as I was busy with an unfortunate[0] side effect of the outage. These generate more anxiety, as being woken up at 6am and being told that nobody understands exactly why the system is acting this way is not great. But then again, we're trained for this situation and there are always at least several ways of fixing the issue.
Finally, it's worth repeating that incident management is just a part of the SRE job and after several years I've understood that it is not the most important one. The best SREs I know are not great when it comes to a huge incident. But, they're work has avoided the other 99 outages that could have appeared on the front page of Hacker News.
Hey Dang, thanks for cleaning up the thread. One thing to note is that the title is not correct. The entire region is not currently down, as the regional impact was mitigated as of 06:39 PDT, per the support dashboard (though I think it was earlier). The impact is currently zonal (europe-west9-a), so having zone in the title as opposed to region would reflect reality closer.
Finally, there's lots of good feedback on this thread and on the previous one (https://news.ycombinator.com/item?id=35711349), so we obviously have a lot of lessons to learn.
The Wolfram Alpha custom keyboard on Android is absolutely the best keyboard for coding.
Needless to say I’m pretty happy they didn’t take my money in the end.
https://www.lesswrong.com/posts/niQ3heWwF6SydhS7R/making-vac...