HAProxy 2.4
haproxy.com
haproxy.com
It looks like it's being adopted by a few larger app monitoring products, wondering if datadog will follow suit or if they will stick with open tracing and their custom implementations for logging and metrics. I assume they will support ingestion from ot at some point though. It sounds neat tho, will definitely read up on it more.
[1] https://github.com/haproxy/haproxy/blob/master/addons/ot/REA...
[2] https://github.com/haproxy/haproxy/blob/master/addons/ot/tes...
[3] https://github.com/haproxy/haproxy/blob/master/addons/ot/tes...
[4] https://github.com/haproxy/haproxy/blob/master/addons/ot/tes...
[5] https://github.com/haproxy/haproxy/blob/master/addons/ot/tes...
[1] https://www.haproxy.com/blog/haproxy-enterprise-2-3-and-hapr...
use_backend fix_servers_a if { var(txn.sendercompid) -m str firmA }
default_backend fix_servers_b
Then "fix_servers_a" could be defined as a round-robin backend.- the site does not do a good job of describing the products and what they do. https://www.haproxy.com/products/haproxy-enterprise-edition/ has a download button as the start and tells nothing about the product.
- it is hard to find requirements from the first page. (https://www.haproxy.com/documentation/hapee/2-2r1/getting-st...)
- is there any a page with basic architecture of the product? Some diagrams with typical configurations.
- is there a simpler HA setup than the 3 layer one here: https://www.haproxy.com/documentation/hapee/2-2r1/high-avail... ? What is the simplest HA configuration for 2 web servers?
> - the site does not do a good job of describing the products and what they do.
The page does lead off by telling you it is a software load balancer, but I agree, we can do a better job of giving more information on the product itself. Most people who pursue HAProxy Enterprise are already familiar with HAProxy. With that said, we're actually working on a complete overhaul to this page and I believe the new version will address this.
> - it is hard to find requirements from the first page.
This is just a reference point and typically organizations are working closely with a sales engineer.
> - is there a simpler HA setup than the 3 layer one here:
A simpler version is found under active/standby here [1]
[1] https://www.haproxy.com/documentation/hapee/2-2r1/high-avail...
I haven't heard many problems with AWS or Google load balancing, but sometimes there's issues where people talk about 'pre-warming'... more control would probably help there.
If you're running HAProxy on AWS, be sure your AWS firewall rules are stateless, otherwise you may well run into the unpublished connection tracking limits.
This.
I got bit by this the other day, a bunch of servers were having very low CPU usage but new connections to them were failing. Turns out it was because they were hitting the conntrack limit. Changing the security group rules to allow all from all "fixed" it.
Although after a quick test, I think there's another limit, also unpublished, as I still can't establish as many connections as I want on a smaller instance – same number of requests on fewer connections works, and also a bigger instance works with many connections. In every case, the CPU was largely unused.
AWS doc regarding the issue: https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/security...
Any chance this is an OS limit in your instance? Either a connection tracking firewall there, or FD/socket count limits? I don't have experience beyond the AWS connection tracking limits because I ran into them while testing something, but only found out how to avoid them much later (the test didn't go well, because of the limits, so we didn't go forward, and our assigned rep didn't tell us about the workaround either; I only found out by complaining about the limit to enough people that someone knew how to fix it)
If it's OS related, it may be some dynamic limit based on RAM size. But the limit was something around 100-200 connections, so nothing particularly crazy.
The way I found out about the limit was that downstream servers couldn't establish connections, and a manual test showed significant packet loss, while the servers were under 10% CPU and accepted connections locally. I opened a support ticket and the guy apparently saw that they were hitting the conntrack limit.
Once I replaced the security group with a blanket allow all as he suggested, connections could be established again. But two days later, when the number of connections increased further (while the security group was still allow all), I started seeing again connections not being established, although the servers were still not particularly loaded. As these are UDP (DNS) servers, I can't usefully quantify the number of different clients.
I set up a quick lab with two isolated EC2 instances and figured there must be another limit to the number of connections, although it seems somewhat higher – or maybe the conntrack limit as shown by ethtool doesn't update immediately. This happened yesterday, Saturday, while I was on call, so I didn't push the tests far enough to reach any clear-cut conclusion. I just doubled the servers and everything went back to normal.
However, the limit does seem to be related to the instance size and per interface. Meaning that adding another interface doubled the connections I could establish. Again, I didn't push the tests any further, though I intend to this sometime next week.
Stickiness, rate limiting, gzip compression, configurable logging, caching, Lua scripting, load balancing to domain names, more options for health lifecycle. It even has a control socket to instantly manage the server without restart. HAProxy is actually the inventor the PROXY protocol which a number of other systems have now copied.
The convenience of cloud proxies is no small matter, but on a feature-by-feature basis HAProxy easily dominates.
[1] https://www.haproxy.com/blog/announcing-haproxy-data-plane-a...
[2] https://www.haproxy.com/blog/announcing-haproxy-data-plane-a...
More control/flexibility (i.e. cloud provider usually gives you a limited number of "checkboxes" that they want you to be able to modify), generally better observability (HAProxy timing metrics [1] and session state at disconnection codes [2] are very helpful for debugging), and granular control over timeouts [3]. You also get builtin rate limiting [4], which from what I have seen with most cloud load balancers usually requires an up-charge through their "WAF" product.
Better health checking, circuit breaking, and layer4/layer7 retries [5].
Another consideration is whether you have spiky traffic patterns. For AWS load balancers at least, they need time to "warm up" to larger traffic levels. If you manage your own load balancing tier you can scale on demand.
[1] https://www.haproxy.com/documentation/hapee/latest/onepage/#...
[2] https://www.haproxy.com/documentation/hapee/latest/onepage/#...
[3] https://www.haproxy.com/documentation/hapee/latest/onepage/#...
[4] https://www.haproxy.com/blog/four-examples-of-haproxy-rate-l...
[5] https://www.haproxy.com/blog/haproxy-layer-7-retries-and-cha...
Question about the suitability for a dynamic backend case. I have a system where clients are assigned to backends exclusively for the duration of a session. At the moment I have a pool of backends and each client is explicitly told which backend to use so they can route to it. Ideally I’d make this transparent so it was easier to reassign on the fly. Is this case reasonably supported with HAProxy and which features should I be looking at to get a sense of it?
I’m working under the assumption that I’ll need to build the machinery to maintain the pool and registry of where clients should be routed to.
[1] https://www.haproxy.com/blog/introduction-to-haproxy-maps/
[2] https://www.haproxy.com/blog/introduction-to-haproxy-stick-t...
For early stage projects, and particularly for personal projects it's more cost efficient to use haproxy.
As an added bonus it's got all the bells and whistles to match demands as you scale, and no (cloud) vendor lock-in
First to me is the ability to retain the IP address. When using AWS Load Balancer, the LB scale automatically and replace the IP.
Second is sophisiciated routing. Example route by cookie. What it buy is that by setting a cookie you can make sure the request hit a specific pool. You can also have concept of backup
Third is lots of useful metrics with their stats page.
Fourth is the ability to alter the responses such as removing header. You have to pair with lambda edge for these on AWS. What it helps is imagine you can write back the user_id to HAProxy, HAProxy log it together with the request, and delete it(so client won't see it). It helps to easily tail the log for a specific user.
And a lot of rate limiting built-in features that just take a few lines of config.
With an application load balancer, you can set up multiple backends, do basic redirections, do OIDC authentication, etc.
With a network load balancer, you can only spread TCP and / or UDP connections.
We deploy auto-scaling groups for each major app version (because customers can choose when to upgrade, so we can have anywhere from 1-3 versions live at a time). There's a database that has an entry for each customer system, with one or more domains, the app version and the db name (usually auto-generated but for historical reasons can be set manually). There's a UI to manage all this.
A script takes this data and builds haproxy config, creating backends for each version group and routing rules for domain to the proper group. This part could maybe be done with ALB now, but I am not certain of that.
We also automatically configure SSL for all domains: anything that doesn't have a static .pem file gets LetsEncrypt cert. Most of these are done via HTTP-01 because they're customer-owned domains that just CNANE to us. None of this is doable via AWS built-in stuff.
There's also a bunch of other hacks that haproxy does:
doesn't redirect to SSL for a couple specific (non sensitive) URLs+user agent that doesn't follow redirects;
returns a fake "success" page for a long-gone service called by an obsolete client some customers of (ex-)customers are still running, which effective causes a DDoS attack due to retries if we return 404 or 5xx;
Has some awareness of backend state and shows better error pages than just a generic 5xx depending on situation.
Haproxy instances are behind NLB, but otherwise there's a single (layer 7) hop to the app server.
The end result is you can configure a new system via our management UI, and so long as DNS is setup (using a wildcard subdomain we own, and/or customer's CNAME entry) within a few minutes it will be live (database deployed, app servers aware of connection/domain mapping, proxy configured with SSL).
This is analogous to the move we're all witnessing from HTTP to HTTP/2 (standardized from SPDY) and HTTP/3 (QUIC).
The desire to reduce the amount of bandwidth spent pushing uncompressed text around and reduce the amount of CPU needed to parse strings is what spurred market participants to explore alternatives to FIX, namely in the form of binary encoded message that costs less to parse (from a CPU and bandwidth perspective).
As libraries release support for alternative protocols and the industry compiles their projects against these newer versions, the adoption of non-FIX will keep increasing.
Similar to how you get HTTP/2 for "free" when using the Go stdlib so its downstream consumers such as Caddy and Traefik can provide that as a default to end users.
https://en.wikipedia.org/wiki/List_of_electronic_trading_pro...
EDIT: Added year for context on timeframe of first implementations
HAProxy is virtually unusable as a file server, but really shines as a reverse proxy. One of the main reasons I almost always deploy HAProxy in tandem with (stock) nginx is the superior HTTP rewriting / configuration. See also my comment back on HAProxy 2.0: https://news.ycombinator.com/item?id=20198232
HAProxy's configuration is a procedural style configuration. The HTTP rules (e.g. adding or deleting headers) are processed in order if their condition matches. (Stock) nginx has a declarative configuration which makes it hard for me to understand which options apply when and the inheritance by nesting different blocks is sometimes very confusing. One example would be an `add_header` in a `location` completely overriding (instead of extending) an `add_header` in a `server`. So if I want to put e.g. HSTS into the server to apply to the full host and also want to add headers for different paths, then I need to duplicate the HSTS stuff.
And if you don't know of the product, just Google it.
I sometimes feel the collective of spirit of HN can never be satisfied. If you do put that sentence in, someone will probably question why it's there :p
But I feel sorry for him now, with HTTP/2 and HTTP/3 the pain will never stop. And IPv6 makes most of this obsolete if it happens!
I stick with HTTP/1.1 since I'm building a native 3D MMO; only Microsoft can block port 80 which I doubt they will atleast in the coming 5 years!
Linux wont, they are not that stupid.
Edit: allready the bots have downvoted, this is beautiful.
If all your machines have their own unique public address you can loadbalance with your own software by pointing people directly to different machines.
What?! Load balancers tend to be in place for availability and scalability reasons, not to save IP addresses. Sure, there are many use cases where you might not need to have public IPs on the backends a load balancer is sending traffic to, but to say that's the only "real" reason for a load balancer seems to be missing the point.
It can also be used for canary deployments where you roll out a feature to 1% of your user base.
There are several other use cases. I'm pretty sure 'saving IP addresses' is not a use case you'd use a load balancer for.
NAT. You're thinking of NAT.
It's common to do internal services that way but you can't control how or who or the internet might try to connect to your service so the load balancer offers a layer of protection so servers can't be picked off individually.
They can also offer an abstraction/routing layer so you can do things like canaries and controlled rollouts which would be much more difficult just advertising IPs directly to the client.
You cannot use different ports from the client, it has to be 80 for HTTP, I meant; do you own loadbalancing from the client!
So f.ex. the browser would load parts of the site from another server via javascript or redirect to another dns with another IPv6 address.
To have a bottleneck between all your clients and all your servers is not good architecture.
Maybe not, but everyone else is doing it. If your clients are web browsers, it's hard to convince them to do useful things like notice servers have changed; much easier to focus their attention on a handful of load balancers on static IPs and figure it out from there.