The New HAProxy Data Plane API: Two Examples of Programmatic Configuration
haproxy.com
haproxy.com
I wonder if something could sit in between and translate. It would be great for all the existing control planes to work with multiple data plane proxies.
HAProxy's API looks like more of a transactional RPC API and HAProxy manages the configuration internally.
Getting an entire snapshot of configuration data is quite compatible with this, since then the explicit use of transactions is not necessary and the dataplane can take care of it on its own.
But think about a system which is dynamically adding domains with TLS and SNI, or constantly adjusting backend workers. Or trying to automate a clean rolling upgrade.
You want to quickly shed load, or add new TLS certificates, etc. If you have a process which is accepting these change config requests and controlling a file, it potentially has to worry about atomicity if several different things are happening at the same time — could one change blow away another if you script a “read config, update config, link new config, trigger reload, poll till complete” function without locks, one change could blow away another?
Presumably the API would protect against this and allow a simpler function to push a given state change?
But I agree that it’s likely over designed and not strictly necessary assuming the hot reload doesn’t have unanticipated side effects. E.g. What happens to stick tables, session state, etc.?
It’s not like the REST API eliminates all these concerns too. For example, if you make a change via the API, presumably it is persisted across reloads, or there is at least a way to make it be?
HTTP offers none of this. No atomicity and no consistency of the full configuration. Want to edit a service or maybe alter some hosts? Better hope it's already setup as expected before editing and you're doing all the right calls in the right order and they all succeed. You will have to write the hell of a state machine to ensure to get to the expected state.
The only sane use case for live-changes is enabling/disabling a server, which is necessary to perform rolling maintenance or blue-green deployment.
If you're balancing traffic by session count, the new process doesn't know about the sessions in the old process (not sure if you can workaround this, it's not a problem I have, just something I'm aware of)
If you use source port ranges, the new process doesn't know what the old process is still using, so you get extra logs about ports in use, and extra CPU spent on syscalls to find free ports. For me, on a busy system, this makes it even busier for quite a few minutes before it slows down to normal again. I haven't looked at this API yet though to know if it will help for me.
One issue we have with our internal instances is that while the inbound sockets are passed across reloads, the connections/sessions state have to be initialized on every reload. We also use mTLS, terminated between services by haproxy itself on both sides.
So we have some services using haproxy as a client side load balancer where there are hundreds of backends for the service and that list changes somewhat frequently triggering reloads. This leads to heavy CPU usage on all the backends who have to deal with clients reloading and re-negotiating mTLS sessions.
I haven't been able to tell from the various marketing messages, but does the new data plane API avoid restarting the process, and hence this re-negotiation?
Because this stuff is only available in the nginx paid version.
This pretty much makes making k8s ingresses very easy.
P.S. in typical haproxy vein, its a bit hard to get started. for example the Docker page doesnt have a trivially runnable version for haproxy 2.0 (https://hub.docker.com/_/haproxy/).
As for why people use APIs for configuring things, it's because infrastructure changes frequently enough that encoding it in a static file isn't practical. Imagine you are rolling out a new software release. There are 100 replicas. You could start up a new replica, see if it passes health checks, edit the haproxy config to configure that as an endpoint, wait for haproxy to start sending it traffic, check that it's working, edit the config file again to remove the old version's replica, wait for haproxy to acknowledge that change, shut down the old replica... and finally do that again 99 more times. Nobody wants to do that, so there's an API to inform your frontend load balancer of where your workers are. When it needs some endpoints, you tell it.
I guess an even better example is TLS certificates. They expire every 3 months. What do you do when a new one is released? Drain traffic from one load balancer, copy the certificate and key over, wait for the load balancer to restart, and then undrain traffic? No, that's a huge pain. You just have your load balancer connect to your secret discovery service and have the same program the renews the certificate send your load balancers the new certificate. Then you don't have to drain; when a new connection comes in, you present the new certificate.
All these config APIs make it possible to operate software at scale. You probably don't need it for your personal website. But you could probably eschew servers entirely and just nc -l on port 80 and type the response in your terminal when a request comes in and still have a mostly-working website.