We reduced 502 errors by caring about PID 1 in Kubernetes
about.gitlab.com
about.gitlab.com
Some less common knowledge: Do your pre-startup config with a shell script and then `exec` your process. If you need some sort of process manager, use supervisord and configure stdout/stderr capture and forwarding. Supervisord will do anything you need from init in a container.
If you're running a jvm in the container, set the dns cache ttl in the java security config.
Sometimes it's very useful to include a preStop hook that will end your process and then sleep for N seconds so that your process unbinds its listening socket and then for a period of time the pod will respond to incoming SYNs with a RST. So your upstream app doesn't sit there waiting on a timeout because for various reasons it still made new connection attempts despite the pod being removed from the service endpoints or w/e.
Fix your resolver configs. ndots. timeout. retry. Remove all the hosts search domains by setting `--resolv-conf=/dev/null` in your kubelet options. Then set dnsConfig.options in your pod spec. While you're at it, set `--allowed-unsafe-sysctls 'net.ipv4.tcp.keepalive*'` so you can configure your pods with sane values. Update pod spec with securityContext.sysctls to set the per pod values. Update the kubelet before applying that pod spec because lol if you don't... (your pods do not inherit these values from the host! So if you think you've set them, you're probably wrong.)
For people new to Dockerfiles, it's too easy to be oblivious to the fact that the two forms do something differently.
For people who've used Dockerfiles a lot, it's too easy to miss the formatting difference, especially when the two different CMD forms can successfully run what you want it to run.
CMD should have been split into two separate statement types, either completely separate types (e.g. SHELLCMD vs RAWCMD) or two explicit subtypes of CMD (e.g. CMD SHELL vs CMD RAW). Then you couldn't read each of the CMD variants without seeing the difference.
New users would see the difference and think "Why is there a difference? I should look it up." Old users would more easily notice when the shell CMD is used when the raw CMD should be used, and vice versa.
It also, though, makes me think about whether technology has progressed by adopting Kubernetes, or if we've just made things more complicated.
Reading through the discussion, it doesn’t look like it was quashed; the project owners had a reasoned debate about the pros/cons of what would happen if such a support was included at the time and its a totally reasonable argument to me.
Replicating the capabilities it provides in another system that offers such capabilities wouldn’t involve less complexity. In fact, I would argue that the way the authors reasoned through the issues was made possible because they were able to operate with these abstractions. Imagine doing this but with eg AWS NLBs. This team might have discovered similar quirks in that setup. But now their knowledge and what they could share would be restricted to AWS NLBs. Since they’re working with K8s, the lessons they learned are a lot more widely applicable.
Another thing to consider (as pointed out by a sibling) is that an average user of k8s may not likely have to worry about minimizing 5xxs during application termination. But if they did want to, now they can.
It's always amazing to me how complicated web stuff has gotten these days. I started out in the LAMP days over 20 years ago. I'm not this complexity is necessarily bad. I don't know the details and I don't know how to make a web app scalable today. But damn if this doesn't sound like a pile of hot garbage waiting to fail to my ears.
GitLab (the software) is a child of the microservices culture. That said, that architecture is probably a very good fit for GitLab (the company). It's just not very good for other people to self-host it.
One problem is it's still too hard. Hosting your own Gitea/gogs shouldn't be any more difficult or less secure than installing an app on your phone. This is where my work is currently.
Another problem is things like repo stars. You can federate, but then every instance has to deal with spam. Not sure what a great solution to this is.
I'm not saying the technology isn't good or solves a certain problem, but the reality is kubernetes is being heavily marketed and in many cases I've encountered it wasn't a good fit for what the team was trying to do. A server and copying files would have been easier but they've already made an investment and going back would look like failure.
Running dumb-init (or tini or whatever, there's lots of these) when your application can handle SIGTERM itself is unnecessary.
These are tools for when you don't control an executable in the stack and it doesn't do the right thing itself.
(I do have a lot of experience with helpful corefiles, it's just been totally orthogonal to my experience with "at-scale" backend deployment. Different languages, different dev/ops relations, different attitudes towards abends in general.)
Thats a big “when”. Many applications don’t. e.g. they spawn child processes which they forget about
And "many" is kind of a cop-out word; yes, I'm sure by absolute number it's a lot but:
- If you're in e.g. Go, you spawn thousands of child green threads but it's trivial to route a SIGTERMable context into all of them and wait for only the relevant ones.
- If you're in e.g. Python, you'll be running in an application server that handles it and you're probably not spawning many threads yourself because it's usually pointless.
- If you're any piece of software written to run outside of containers, that's also what non-container environments have used for decades to handle smooth upgrades.
- etc. for other common cases in the end, "spawning threads" is no excuse for not handling clean listener shutdown and seems to be more about how "Unix-like" the application is and not threading - e.g. old Java programs tend to be worst in my experience.
I've used tini et al mostly when putting entire legacy environments (e.g. nginx+PHP+memcache+mysqld) into a single container with a hacky shell script around them because we needed them off bare metal yesterday but there's no dev resources to properly adapt the application. There's definitely "many" of those in the world, but it doesn't speak to picking tooling to support that by default.
Do shells not forward signals correctly to their child processes? This seems surprising to me.
https://www.gnu.org/software/bash/manual/html_node/Signals.h...
I'm confused about this myself because I thought bash ignores SIGTERM only when it's interactive:
https://www.gnu.org/software/bash/manual/html_node/Interacti...
In this case bash is not interactive, so I'm not sure why SIGTERM is ignored, but FWIW, the docker documentation explicitly says that it will:
> The shell form prevents any CMD or run command line arguments from being used, but has the disadvantage that your ENTRYPOINT will be started as a subcommand of /bin/sh -c, which does not pass signals. This means that the executable will not be the container’s PID 1 - and will not receive Unix signals - so your executable will not receive a SIGTERM from docker stop <container>.
https://docs.docker.com/engine/reference/builder/#cmd
The bash documentation for -c says nothing about different signal handling. POSIX mode doesn't address this behavior either.
They list this in the Takeaways but I don't see that they took any action to work around this problem. SIGTERM to the pod and updating the service/endpoint race and you will absolutely see traffic getting sent to terminated processes because of it.
The only "solution" I know is to put a sleep in the preStop hook to delay the SIGTERM so that the pod can be removed from the service. Cloud native!
Once the graceful shutdown was properly executed, it closed any open connections to that pod and stopped the 502s they were seeing. Sounds like either the race wasn't happening or they didn't see/care about it.
I would caution one against "javascript framework syndrome:" bah, this framework is too complicated, I just want something simpler ... ok, it doesn't work right in all circumstances, I'll just add this one feature(, ... ok, maybe just this one other feature)+ ... bah, this framework is too complicated!
Yeah I wouldn’t touch that with a 10ft pole. No way you are going to do something complex without introducing hideous errors because bash is a horrible language for doing anything complicated.
I probably made the mistake of making it too complicated already. I'll go back at some point and strip most of it out.
etcd (the core of k8s) uses raft as its consensus algorithm. I’m not even sure what you are conflating, or what this magical single packet you are referring to is. Kubernetes will very much not continue to work if etcd goes down which is why people tend to host it outside of k8s itself for production use cases.
Kubernetes will continue to run in a degraded state if etcd loses quorum. Your services will still be reachable.
Also,
> We had a bit of tunnel vision and kept focusing/blaming that we aren't removing the Pod from the Endpoint quickly enough.
The underlying root cause should be identified so that this same issue can be corrected when retries aren't a solution, but if this was causing Gitlab Pages to burn through it's SLO (which is stated in the article), the Pages team could have (and should have) stopped the leak with a retry which the root cause was being determined.
Yes, "do some backoff", etc., but then we're in much murkier waters where Pages probably also has a latency SLI and if you're talking to a connection that might not let you know it's going to reset for up to 30 seconds, that's probably not going to fly either.
If you're in some other environments, e.g. a browser, this might not even be possible.
And again: If the request is non-idempotent I'm SOL no matter what. (I don't see anything in the post that says it was or wasn't.)
Reliable systems are built at the client, not at the backend. You get those extra 9s from having a smarter client, not from having a more reliable backend. Trying to make your lowest layers flawless is a fool's errand.
This is kind of a big assumption you're making. If you're going to do retries when a backend is unavailable, explicitly doing it against the same backend is obviously a decision that defeats the entire purpose.
> probably also has a latency SLI
Then take this into account w.r.t timeouts on internal requests/retries.
My point is that this was a problem that was (at least partially) addressable by the team with the impacted SLO, and did not have to wait for someone to fix the root cause in order to improve their situation.
This doesn't sound good to me and translates to "our staff may not know what they're doing". Lack of or outdated documentation, lack of time to read it?
My job is currently mostly detective work, and it’s neither a function of incompetence nor laziness, it’s just a function of needing to prioritize business objectives over sacrifice-able things like documentation.
And that's perfectly fine. This post was published so GL folks can say "we figured out something tricky/interesting and our scale is what made it more tricky". I love figuring out "Frankenstein" type of bugs, errors and problems myself. What I dislike figuring out is something that shouldn't require any figuring out at all.
I read GitLab or Cloudflare's posts with pleasure. These doesn't seem to be written solely as technical recruitment "we do cool stuff" posts.
I left every single company which wanted me to spend a lot of time (weeks, months) to figure out their custom stuff due to no documentation and lack of people who worked on this and knew what makes the thing tick or what to do when it stops. Having noone or a single person kind-of knowing how to do something in a company is a managerial knee shot.
I value my learning/working time maybe more than an average person due to focusing issues. Spending it on figuring something which could have been documented but wasn't absolutely isn't my favorite thing to do. Weeks of detective work are weeks when I'm deprived of ability to learn something useful for my next job.
So while this GitLab post doesn't specify what kind of things are there to figure out, I consider stating that there's always something to figure out as unprofessional as it points out possibly unprepared staff.
My own sense here is that detective work is a deep part of being a software developer, because fundamentally we can’t escape the particulars of the problems we work on, but I guess that really depends on the complexity of the systems we’re working on.
Figuring out things you neither wrote nor fully studied beforehand, without breaking everything, is not some waste of valuable time, that IS the valuable expertise itself.
If you don't like doing that then that just means you're not a good fit for that role, not that there's some failing somewhere else that it's even needed.
It's not possible to document everything, and even if it were, no one would ever be qualified to do the job by this standard, because 3 things changed in production just while the on-call was making coffee.
This story was utterly normal and the only bad thing in it is just generally how tall and shaky and complex "normal" has become by now. But that's not anything any single company can do much about. Large web services need load balancing and orchestration and a lot of different moving parts, and every one of those parts are changing every minute because somewhere in the world a productive developer just committed a new line of code in something you use.
Well, one company can certainly make it worse, by e.g. ignoring years of API ergonomics around process spawning and give you a vaguely-specified "command" string, then pumping VC money into their marketing budget to make themselves an "indispensable" daily tool.
(Mostly I agree with you - just want to point out that in addition to the necessary complexity, the tooling around it also produces a lot of unnecessarily wrong-by-default behavior.)
It does.
What I meant by one company doing something about it was, in realistic terms, no one can do much about the fact that the rest of the world uses a lot of complex systems that aren't as robust as earlier simpler systems. You have to use the current tools and that's just how they are now. You can try to be less stupid as possible, but you can't change the ecosystem you need to live in and interoperate with.
Amen to that! It's basically how you start understanding how complex systems work in real-life and the lessons you learn while there can and will help you in every job you will have afterwards. Personally I also think it is what makes remarkable developers vs developers that just stay in their (usually tiny) domain.