HNHacker News
TopNewBestAskShowJobs

eastdakota

11,162 karma · joined December 6, 2010

A little bit geek, wonk, and nerd. Repeat entrepreneur, recovering lawyer and former ski instructor. CEO & co-founder of CloudFlare. [ my public key: https://keybase.io/eastdakota; my proof: https://keybase.io/eastdakota/sigs/_uDY0ZsLTEWaNu5daRtuwzZtJDJJrtGS4uXoYxwI634 ]
submissionscomments
eastdakota··on Clef: Open-weight decision models, and new RL fine-tuning platform
I think you just called me old.
eastdakota··on Precursor
You’re confusing us with Google. We don’t have a visual CAPTCHA.
eastdakota··on Cloudflare outage on November 18, 2025 post mortem
That’s not accurate. As with any incident response there were a number of theories of the cause we were working in parallel. The feature file failure was one identified as potential in the first 30 minutes. However, the theory that seemed the most plausible based on what we were seeing (intermittent, initially concentrated in the UK, spike in errors for certain API endpoints) as well as what else we’d been dealing with (a bot net that had escalated DDoS attacks from 3Tbps to 30Tbps against us and others like Microsoft over the last 3 months). We worked multiple theories in parallel. After an hour we ruled out the DDoS theory. We had other theories also running in parallel, but at that point the dominant theory was that the feature file was somehow corrupt. One thing that made us initially question the theory was nothing in our changelogs seemed like it would have caused the feature file to grow in size. It was only after the incident that we realized the database permissions change had caused it, but that was far from obvious. Even after we identified the problem with the feature file, we did not have an automated process to role the feature file back to a known-safe previous version. So we had to shut down the reissuance and manually insert a file into the queue. Figuring out how to do that took time and waking people up as there are lots of security safeguards in place to prevent an individual from easily doing that. We also needed to double check we wouldn’t make things worse. The propagation then takes some time especially because there are tiers of caching of the file that we had to clear. Finally we chose to restart the FL2 processes on all the machines that make up our fleet to ensure they all loaded the corrected file as quickly as possible. That’s a lot of processes on a lot of machines. So I think best description was it took us an hour for the team to coalesce on the feature file being the cause and then another two to get the fix rolled out.
eastdakota··on Cloudflare outage on November 18, 2025 post mortem
There’s lots of things we did while we were trying to track down and debug the root cause that didn’t make it into the post. Sorry the WARP takedown impacted you. As I said in a comment above, it was the result of us (wrongly) believing that this was an attack targeting WARP endpoints in our UK data centers. That turned out to be wrong but based on where errors initially spiked it was a reasonable hypothesis we wanted to rule out.
eastdakota··on Cloudflare outage on November 18, 2025 post mortem
Team has a good sense, typically. In this case, the names of the columns in the Bot Management feature table seemed sensitive. The person who included that in the master document we were working from added a comment: “Should redact column names.” John and I usually catch anything the rest of the team may have missed. For me, pays to have gone to law school, but also pays to have studied Computer Science in college and be technical enough to still understand both the SQL and Rust code here.
eastdakota··on Cloudflare outage on November 18, 2025 post mortem
* published less than 12 hours from when the incident began. Proud of the team for pulling together everything so quickly and clearly.
eastdakota··on Cloudflare outage on November 18, 2025 post mortem
Next time open your dev console in your window and look at how much is going on in the background.
eastdakota··on Cloudflare outage on November 18, 2025 post mortem
We incorrectly thought at the time it was attack traffic coming in via WARP into LHR. In reality it was just that the failures started showing up there first because of how the bad file propagated and where it was working hours in the world.
eastdakota··on Cloudflare outage on November 18, 2025 post mortem
Chicken burrito from Coyo Taco in Lisbon. I am not proud of this. It’s worse than ordering from Chipotle. But there are no Chipotle’s in Lisbon… yet.
eastdakota··on Cloudflare outage on November 18, 2025 post mortem
Well… we have a culture of transparency we take seriously. I spent 3 years in law school that many times over my career have seemed like wastes but days like today prove useful. I was in the triage video bridge call nearly the whole time. Spent some time after we got things under control talking to customers. Then went home. I’m currently in Lisbon at our EUHQ. I texted John Graham-Cumming, our former CTO and current Board member whose clarity of writing I’ve always admired. He came over. Brought his son (“to show that work isn’t always fun”). Our Chief Legal Officer (Doug) happened to be in town. He came over too. The team had put together a technical doc with all the details. A tick-tock of what had happened and when. I locked myself on a balcony and started writing the intro and conclusion in my trusty BBEdit text editor. John started working on the technical middle. Doug provided edits here and there on places we weren’t clear. At some point John ordered sushi but from a place with limited delivery selection options, and I’m allergic to shellfish, so I ordered a burrito. The team continued to flesh out what happened. As we’d write we’d discover questions: how could a database permission change impact query results? Why were we making a permission change in the first place? We asked in the Google Doc. Answers came back. A few hours ago we declared it done. I read it top-to-bottom out loud for Doug, John, and John’s son. None of us were happy — we were embarrassed by what had happened — but we declared it true and accurate. I sent a draft to Michelle, who’s in SF. The technical teams gave it a once over. Our social media team staged it to our blog. I texted John to see if he wanted to post it to HN. He didn’t reply after a few minutes so I did. That was the process.
eastdakota··on Cloudflare outage on November 18, 2025 post mortem
Can attest: not a single LLM used. Couldn’t if I tried. Old school. And not entirely proud of that.
eastdakota··on Cloudflare outage on November 18, 2025 post mortem
That’s correct.
eastdakota··on Cloudflare outage on November 18, 2025 post mortem
Because we initially thought it was an attack. And then when we figured it out we didn’t have a way to insert a good file into the queue. And then we needed to reboot processes on (a lot) of machines worldwide to get them to flush their bad files.
eastdakota··on Cloudflare outage on November 18, 2025 post mortem
We don’t know. Suspect it may just have been a big uptick in load and a failure of its underlying infrastructure to scale up.
eastdakota··on Unpacking Cloudflare Workers CPU Performance Benchmarks
That definitely is one not-wrong conclusion.
eastdakota··on Unpacking Cloudflare Workers CPU Performance Benchmarks
Blog run by the engineering team. I wouldn’t even know how to veto a post if I wanted to. Not in our DNA.
eastdakota··on Cloudflare Email Service: private beta
I’m not going anywhere anytime soon.
eastdakota··on Ask HN: Why is Cloudflare sending my US traffic to London?
Free customers are served from the nearest location where we have capacity. If we’re capacity constrained then free customers will be the first to be rerouted to another facility with capacity. That typically only happens for a very narrow window during any day. It has nothing to do with load to your particular site. It has to do with a region’s capacity and a group of customers (e.g., FREE PRO BIZ ENTERPRISE).
eastdakota··on A look at Cloudflare's AI-coded OAuth library
I agree with Kenton’s aside.
eastdakota··on Chaos in the Cloudflare Lisbon Office
We’re in Munich, pretty close, and Lisbon, which is lovely. Unlikely Austria any time soon.
eastdakota··on Chaos in the Cloudflare Lisbon Office
Had no idea there was a patent. Even if we had, think we’d have risked it.
eastdakota··on Chaos in the Cloudflare Lisbon Office
If it were merely marketing spend for customer acquisition, I bet the ROI on the lava lamp wall in SF has been 100,000x. This isn’t hard to figure out.
eastdakota··on Chaos in the Cloudflare Lisbon Office
Actually, it’s because we believe innovation is easier the further you are from HQ. And Portugal is wonderful. I’m spending about 4 months of the year there with the team who hail from around the world.
eastdakota··on Chaos in the Cloudflare Lisbon Office
No we’re not. We opened an India office because it’s a big market for us. But nothing has gotten “offshored” there.
eastdakota··on Chaos in the Cloudflare Lisbon Office
If you figure out how to model this fluid dynamics accurately over any reasonable period of time, call me. Lots and lots of more valuable things you could do with that, e.g., accurately predicting the weather.
eastdakota··on Chaos in the Cloudflare Lisbon Office
They should never all be off at the same time. We do cycle through each of them turning off for a period of the day. But, even if they were all off, there are lots of other sources of entropy we use, most of which are for more traditional if far less visually interesting.
eastdakota··on Chaos in the Cloudflare Lisbon Office
Mistakes inevitably happen at scale. Sometimes they’re not caught by traditional channels. What I try and encourage is our leaders to take responsibility and fix them wherever they see them. Which is what John did above.
eastdakota··on Chaos in the Cloudflare Lisbon Office
This is correct.
eastdakota··on Chaos in the Cloudflare Lisbon Office
Great thing about entropy is that adding more never hurts. This is one of many sources — both more conventional as well as unconventional — that we use. If it were to go offline, or somehow be corrupted, it wouldn’t hurt our ability to generate entropy across the Cloudflare network.

What I love about this, the lava lamp wall in San Francisco, and the double pendulums in London, is that it takes something very abstract and makes it tangible for our team and our customers.

eastdakota··on When the Dotcom Bubble Burst
I have plenty of friends who went down that path. They’ve done very well as lawyers. But suffice it to say that if offered they’d readily trade places.
Page 1 of 21Next →