Quick takes on the recent OpenAI public incident write-up
surfingcomplexity.blog
surfingcomplexity.blog
Something that I wish all databases and API servers would do, and that few actually do in practice, is to allocate a certain amount of headroom (memory and CPU) to "break glass in case of emergency" sessions. Have an interrupt fired periodically that listens exclusively on a port that will only be used for emergency instructions (but uses equal security measures to production, and is only visible internally). Ensure that it can allocate against a preallocated block of memory; allow it to schedule higher-priority threads. A small concession to make in the usual course of business, but when it's useful it's vital.
Like, why would operating systems allow themselves to run out of headroom entirely in this day and age?
I think consumer systems do it. macos shows the "quit some app, you're out of ram" but the system itself works.
but if you are asking that OS knows this is containers process and this is control plane process and treat them differently, I think no one does that.
Imagine how frustrating would it be if your os does it wrongly and you end up paying 10x because the os nerfed your main process because it thought something more important was running
cgroups on Linux does exactly this, and are a standard part of ensuring that containers don't exceed their allocated resources.
I am not sure where it is about "one process"?
Unless you are talking about VMs in cloud computing, but even in those cases the VM is usually abstracted away and the end client only sees a VPS (e.g. with EC2 or GCE).
Imagine how frustrating it is when the Linux on your desktop suddenly decides to start randomly killing processes, or worse, attempts to swap some memory out, causing a feedback loop of delays that completely freezes the system until you power-cycle the machine.
That's one of the two things Windows always did better (the other thing is disabling write buffering for removable storage, on the account of that storage being, well, removable).
Resource limits are not something you want to discover when you've exceeded them; they need to be managed and alerted about in advance. "Makes computer slower" and alerting the "people who pay for compute" is preferable to crashing - especially in distributed systems, where failures cascade (particularly when the whole system has been penny-pinched / overoptimized to the same degree as any single computer within it).
Someone replied about cgroups on Linux and how it is bog standard stuff (that ClosedAI just didn't know how to use apparently?)
> why would operating systems allow themselves to run out of headroom entirely in this day and age?
Following the analogy, perhaps for the same reason the richest cities in the richest, most developed Western nations, are rapidly beginning to look very much like third-world slums?
Over-optimization is the name of the game. All systems need some amount of slack in them to stay flexible, robust (and livable). Unfortunately, cutting into the slack is always profitable on the margin, so without top-down intervention, the slack will be cut into until it's gone entirely. "Look, these machines are utilized only to 75% of capacity; adding this new system will only increase it by 5%"; "look, we can save XX$/month by cutting on compute, the machines will still be maxing out at 90%, so we'll still have a buffer". "Oh, this new telemetry service will bump that only by 1%".
"Look, there's so much free space here in between these blocks of flats; adding another block won't hurt."
And so on. Until your machines are running at 99% capacity and you risk global outage every time someone sneezes near the server room. Until your city starts to look like London, and if you're from Central Europe like me, you may start to realize that being a few months or years behind on the most recent gadgets is small price to pay in exchange for cities that are affordable and clean.
There are tons of tools built into modern OSes to manage prioritization of processes, etc. Whether it’s worth the tradeoff for you to deal with the operational overhead of harnessing those features is up to you.
This would require per-cluster testing (and is complex to test since you need to induce representative load) so I suppose hardly anyone does it.
[0] - https://kubernetes.io/docs/concepts/cluster-administration/f...
I personally tend to leave a root shell open (and both my feet remain largely hole-free to this day), but it's pretty common advice, to avoid accidentally typing something unfortunate like rm ./* into the wrong shell
And you're basing this assumption off...?
The entire purpose of the /etc/sudoers file is to configure which users have access to sudo and which commands they can use.
Your top comment’s parent didn’t say the ssh login user had all sudo permissions. For best security, there should be many users which each have different limited permissions. Navigating the multiple `sudo su` is frustrating but has a purpose.
In all of my career I had seen that at one company. Everyone else just leaves is unrestricted. I would be impressed to see sudo used the way it was intended in more places. Some places even use passwordless sudo and ssh multiplexing which together with simple phishing give unfettered and unlogged access to production.
You don't even need to break everything, take 1 apiserver out for admin access (provided etcd is not overwhelmed too).
I suppose one could set up a "fast pass lane" kubeconfig that adds a header that haproxy would understand, and route to a priority class in its queue with e.g. https://www.haproxy.com/documentation/haproxy-configuration-... . But there's no easy `kubectl --with-priority` (or, to my knowledge, good guidelines for the various gitops solutions) that follows this pattern out of the box.
In short, the root cause was a new telemetry service configuration that unexpectedly generated massive Kubernetes API load across large clusters, overwhelming the control plane and breaking DNS-based service discovery.
The DNS song seems appropriate.https://soundcloud.com/ryan-flowers-916961339/dns-to-the-tun...
I think the idea of just serving cached responses indefinitely when api server is unreachable is what you're describing but not sure if this is default. (and probably has other tradeoffs that I'm not sure about too)
My guess is they were running CoreDNS on control plane nodes since that's the kubeadm default.
Dude.
(I'll see myself out.)
Job well done.
“I HAVE NO TOOLS BECAUSE I’VE DESTROYED MY TOOLS WITH MY TOOLS”
The text columns look like the side of that hallway rubber mat that my dog keeps chewing on.
spend a lot of time trying
edge. However, as someone
lieve that true progress is
mes, and for the chickens
y zombies, and the polite
to eat your brain to acquire
be prepared; thus, in the
e scientific breakthroughs,
ast inevitably becomes
he main thing that I ponder is
post-apocalyptic survival
ag-tag group of associates.
cruit: a locksmith (to open
ith has run out of ideas);
row snakes at my enemies
g is a reasonable way to
ble in my ultimate successChatGPT Down - https://news.ycombinator.com/item?id=42394391 - Dec 2024 (30 comments)
maybe instead of relying on kubernetes DNS for discovery it can be closer to something like envoy.the control plane updates configs that are stored locally (and are eventually consistent) so even if the control plane dies the data plane has access to location information of other peer clusters.
That is generally how it is done... there's this never ending conflict between "push" architectures and "pull" architectures, and this scenario sure makes "push" seem better, and it is... until you're in one of those scenarios where "pull" is better. ;-)
That'd be a pretty big architectural change, though
Also, stale-if-error is a far safer pattern for service discovery than ttl’d dns.
How can such a large change not be staged in some manner or the other? Feedback loops have a way of catching up later which is why it’s important to roll out gradually.
"To make error is human. To propagate error to all server in automatic way is #devops."