Upgrading Executable on the Fly
nginx.org
nginx.org
[1]: https://github.com/caddyserver/caddy/blob/v1/upgrade.go
That seems relevant if the process is using a non-privileged port that's >= 1024. If we're talking about privileged ports (<= 1023), though, only another root process could hijack that, and those can already hijack you many other ways.
1. there's a control process and worker processes
2. on upgrade, control process launches new worker processes from the new binary
3. requests are drained from old worker processes
4. most of the time nginx request handlers allocate from a per-request allocation pool, so requests mostly don't share memory
5. for the cases where there are global states, there's a separate shared memory pool that you need to allocate from (which is kind of hard to work if you are not using built-in nginx primitives)
In the one hand, an application should never be able to replace itself with "random code" to be executed. I want my systems to be immutable. I want my services to be run with the smallest set of privileges required.
On the other hand, it encourages "consumer level" users to keep their software up-to-date, even when it wasn't installed from a distribution's repository etc.
So I think in general it's a good feature to have, as advanced users/distributions will restrict what a service/process is able to to anyways and won't have any downsides of not using this feature.
It should be optional, that's all!
To clarify: it doesn't, nor has it ever worked that way. You have to be the one to do that (or someone with privileges to write to that file on disk). Most production setups don't give Caddy that permission. And you have to trigger the upgrade too.
I always assumed it would be the Caddy process itself taking care of downloading the update & replacing the binary, before restarting it.
But I suspect that most Caddy deployments are done via docker, and that requires a whole container restart anyways.
In general I am personally not a fan of Docker due to added complexities (often unnecessary for static binaries like Caddy) and technical limitations such as this. All my Caddy deployments use systemd (which I don't love either, sigh).
I don't like all parts of systemd but I think the socket activation is pretty elegant. The decoupling also allows to bind on port 80 and 443 (because systemd runs as root), and then still have Caddy run as a user process. I think it would unlock nice things for Caddy.
Inetd is listening to a port and then for each new connection, spawning a new process, binding stdin/stout to the socket pair. The main issue was that it could lead to system resource exhaustion pretty easily if too many connections were being opened and there were no good ways to control that.
With systemd, the listening socket is only bound by systemd and passed to the service. The service itself is responsible for accepting and handling the new connections. So it has the control on the rate of new connections, and can also more easily share memory. The main advantage of that approach is that as soon as systemd binds the socket, new connections won't be rejected by the system and will be on hold until the service accept() them. So no connection gets dropped, even during a restart. The service itself is still responsible for gracefully shutting down existing connections on SIGTERM.
I'm not sure it is super critical in the age of containerized workloads with rolling deploys but at the very least the connection draining is a good pattern to implement to prevent deploy/scaling related error spikes.
Not sure how the cloud providers do it though, maybe combination of low DNS TTL and rolling restart since they often have huge fleets of servers which handle ingress?
With ingress you'd have a load balancer in front or have it routed in the network layer using BGP.
But I agree it isn't as easy as a in place upgrade.
Multihomed IP <-> Loadbalancer <-> Application
By having the same setup running on multiple locations you can replace the load banacers by taking one location offline (stop announcing the corresponding route). Application instances can be replaced by taking the application instance out of the load balancer.
At least some I am familiar with operate at the packet level and can hand off live "connections" to a peer or hot standby, along with full session state.
Remember that with TCP or anything else, the abstract session is an illusion.
0. https://www.envoyproxy.io/docs/envoy/latest/intro/arch_overv...
https://www.haproxy.com/blog/truly-seamless-reloads-with-hap...
(but I don't think nginx supports h3 out of the box yet)
It's only a problem if your state is tangled and impossible to serialize or bundle up to hand off.
UDP is perhaps the easiest because there's nothing to do in the basic case, for example with DNS.
Basically the parent executes the new binary after it receives a USR1 signal. Once the child is healthy it kills the parent via SIGTERM. The listener socket file descriptor is passed over an environment variable.
https://github.com/monroeclinton/- (this is the proper url, it's called dash)
In my proposals, there would be a simple application-aware http proxy process that we'd maintain and install on all environments. It would handle relaying public traffic to the appropriate final process on an alternate port. There would be a special pause command we could invoke on the proxy that would buy us time to swap the processes out from under the TCP requests. A second resume command would be issued once the process is running and stable. Ideally, the whole deal completes in ~5 seconds. Rapid test rollbacks would be double that. You can do most of the work ahead of time by toggling between an A and B install path for the binaries, with a third common data path maintained in the middle (databases, config, etc)
With the above proposal, the user experience would be a brief delay at time of interaction, but we already have some UX contexts where delays of up to 30 seconds are anticipated. Absolutely no user request would be expected to drop with this approach, even in a rollback scenario. Our product is broad enough that entire sections of it can be a flaming wasteland while other pockets of users are perfectly happy, so keeping the happy users unbroken is key.
DNS not required. You can use a load balancer to do the same thing. If you don't want a full second setup, do a rolling restart of application servers instead.
Edit: I forgot... you can do this with containers too.
You can follow the trail by searching for ngx_exec_new_binary in the nginx repo.
https://www.nginx.com/blog/socket-sharding-nginx-release-1-9...
Page also has an example of how SO_REUSEPORT effects flow.
https://www.freedesktop.org/software/systemd/man/systemd.soc...
So, simply replacing the binary on disk will cause all new connections going forward to use the new binary, while existing held connections (with in-memory references to the old binary's inode) will finish the operations. Once they are done and all references to that inode are gone, the blocks referencing the binary will be removed.
inetd supports both process-per-connection and single process/multiple connections using the "nowait" and "wait" declarations, respectively. The former passes an accept'd socket, the latter passes the listening socket.
Upgrading an Nginx executable on the fly - https://news.ycombinator.com/item?id=8677077 - Nov 2014 (1 comment)
cp new/nginx /path/to/nginx kill -SIGUSR2 <processid>
That does sound pretty neat if you're not running nginx in a container. I wonder if they've built a Windows equivalent for that.
I don't know why they haven't moved on from it; it only really made sense when uni-core processors were the norm.