Stickiness, rate limiting, gzip compression, configurable logging, caching, Lua scripting, load balancing to domain names, more options for health lifecycle. It even has a control socket to instantly manage the server without restart. HAProxy is actually the inventor the PROXY protocol which a number of other systems have now copied.
The convenience of cloud proxies is no small matter, but on a feature-by-feature basis HAProxy easily dominates.
[1] https://www.haproxy.com/blog/announcing-haproxy-data-plane-a...
[2] https://www.haproxy.com/blog/announcing-haproxy-data-plane-a...
More control/flexibility (i.e. cloud provider usually gives you a limited number of "checkboxes" that they want you to be able to modify), generally better observability (HAProxy timing metrics [1] and session state at disconnection codes [2] are very helpful for debugging), and granular control over timeouts [3]. You also get builtin rate limiting [4], which from what I have seen with most cloud load balancers usually requires an up-charge through their "WAF" product.
Better health checking, circuit breaking, and layer4/layer7 retries [5].
Another consideration is whether you have spiky traffic patterns. For AWS load balancers at least, they need time to "warm up" to larger traffic levels. If you manage your own load balancing tier you can scale on demand.
[1] https://www.haproxy.com/documentation/hapee/latest/onepage/#...
[2] https://www.haproxy.com/documentation/hapee/latest/onepage/#...
[3] https://www.haproxy.com/documentation/hapee/latest/onepage/#...
[4] https://www.haproxy.com/blog/four-examples-of-haproxy-rate-l...
[5] https://www.haproxy.com/blog/haproxy-layer-7-retries-and-cha...
Question about the suitability for a dynamic backend case. I have a system where clients are assigned to backends exclusively for the duration of a session. At the moment I have a pool of backends and each client is explicitly told which backend to use so they can route to it. Ideally I’d make this transparent so it was easier to reassign on the fly. Is this case reasonably supported with HAProxy and which features should I be looking at to get a sense of it?
I’m working under the assumption that I’ll need to build the machinery to maintain the pool and registry of where clients should be routed to.
[1] https://www.haproxy.com/blog/introduction-to-haproxy-maps/
[2] https://www.haproxy.com/blog/introduction-to-haproxy-stick-t...
I haven't heard many problems with AWS or Google load balancing, but sometimes there's issues where people talk about 'pre-warming'... more control would probably help there.
If you're running HAProxy on AWS, be sure your AWS firewall rules are stateless, otherwise you may well run into the unpublished connection tracking limits.
This.
I got bit by this the other day, a bunch of servers were having very low CPU usage but new connections to them were failing. Turns out it was because they were hitting the conntrack limit. Changing the security group rules to allow all from all "fixed" it.
Although after a quick test, I think there's another limit, also unpublished, as I still can't establish as many connections as I want on a smaller instance – same number of requests on fewer connections works, and also a bigger instance works with many connections. In every case, the CPU was largely unused.
AWS doc regarding the issue: https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/security...
Any chance this is an OS limit in your instance? Either a connection tracking firewall there, or FD/socket count limits? I don't have experience beyond the AWS connection tracking limits because I ran into them while testing something, but only found out how to avoid them much later (the test didn't go well, because of the limits, so we didn't go forward, and our assigned rep didn't tell us about the workaround either; I only found out by complaining about the limit to enough people that someone knew how to fix it)
If it's OS related, it may be some dynamic limit based on RAM size. But the limit was something around 100-200 connections, so nothing particularly crazy.
The way I found out about the limit was that downstream servers couldn't establish connections, and a manual test showed significant packet loss, while the servers were under 10% CPU and accepted connections locally. I opened a support ticket and the guy apparently saw that they were hitting the conntrack limit.
Once I replaced the security group with a blanket allow all as he suggested, connections could be established again. But two days later, when the number of connections increased further (while the security group was still allow all), I started seeing again connections not being established, although the servers were still not particularly loaded. As these are UDP (DNS) servers, I can't usefully quantify the number of different clients.
I set up a quick lab with two isolated EC2 instances and figured there must be another limit to the number of connections, although it seems somewhat higher – or maybe the conntrack limit as shown by ethtool doesn't update immediately. This happened yesterday, Saturday, while I was on call, so I didn't push the tests far enough to reach any clear-cut conclusion. I just doubled the servers and everything went back to normal.
However, the limit does seem to be related to the instance size and per interface. Meaning that adding another interface doubled the connections I could establish. Again, I didn't push the tests any further, though I intend to this sometime next week.
First to me is the ability to retain the IP address. When using AWS Load Balancer, the LB scale automatically and replace the IP.
Second is sophisiciated routing. Example route by cookie. What it buy is that by setting a cookie you can make sure the request hit a specific pool. You can also have concept of backup
Third is lots of useful metrics with their stats page.
Fourth is the ability to alter the responses such as removing header. You have to pair with lambda edge for these on AWS. What it helps is imagine you can write back the user_id to HAProxy, HAProxy log it together with the request, and delete it(so client won't see it). It helps to easily tail the log for a specific user.
And a lot of rate limiting built-in features that just take a few lines of config.
With an application load balancer, you can set up multiple backends, do basic redirections, do OIDC authentication, etc.
With a network load balancer, you can only spread TCP and / or UDP connections.
We deploy auto-scaling groups for each major app version (because customers can choose when to upgrade, so we can have anywhere from 1-3 versions live at a time). There's a database that has an entry for each customer system, with one or more domains, the app version and the db name (usually auto-generated but for historical reasons can be set manually). There's a UI to manage all this.
A script takes this data and builds haproxy config, creating backends for each version group and routing rules for domain to the proper group. This part could maybe be done with ALB now, but I am not certain of that.
We also automatically configure SSL for all domains: anything that doesn't have a static .pem file gets LetsEncrypt cert. Most of these are done via HTTP-01 because they're customer-owned domains that just CNANE to us. None of this is doable via AWS built-in stuff.
There's also a bunch of other hacks that haproxy does:
doesn't redirect to SSL for a couple specific (non sensitive) URLs+user agent that doesn't follow redirects;
returns a fake "success" page for a long-gone service called by an obsolete client some customers of (ex-)customers are still running, which effective causes a DDoS attack due to retries if we return 404 or 5xx;
Has some awareness of backend state and shows better error pages than just a generic 5xx depending on situation.
Haproxy instances are behind NLB, but otherwise there's a single (layer 7) hop to the app server.
The end result is you can configure a new system via our management UI, and so long as DNS is setup (using a wildcard subdomain we own, and/or customer's CNAME entry) within a few minutes it will be live (database deployed, app servers aware of connection/domain mapping, proxy configured with SSL).
For early stage projects, and particularly for personal projects it's more cost efficient to use haproxy.
As an added bonus it's got all the bells and whistles to match demands as you scale, and no (cloud) vendor lock-in