HAProxy 1.8
haproxy.com
haproxy.com
If this is on a server with a few GB RAM, is there any reason not to use it as the only cache? Is there any reason not to set the bufsize to a few MB, for example (assuming memory is not a constraint)?
I have a somewhat complex setup with several micro services each served by a dynamic set of backends. Caching for some of it would be nice, but I couldn't figure out a way to use varnish (or other external cache) without ending up with 1 varnish instance per backend (a pain to maintain), introducing a single point of failure (one per service - and still a pain to maintain), or losing some of the benefit and power of haproxy as a frontend (with varnish in front of it).
Disk access can never be deterministic, which is terrible if you care about latency.
Awesome! Thank you!
I can't find it now, but I'm sure they had pretty compelling benchmarks that showed that with keepalive & low-latency networks, performance wise, there's little, if any, performance difference.
http2 backends support is on the roadmap https://trac.nginx.org/nginx/roadmap - so you can expect it to be implemented.
Is HTTP/1.0 (vs 1.1) actually used anywhere? Both versions are also compatible so what's the issue? Also seems rather disingenuous to say "even" when Envoy has plenty of other advanced features.
Disclaimer: We develop https://github.com/appscode/voyager which is a HAProxy based ingress controller for Kubernetes.
Is this true for A records too?
If so, neither haproxy nor nginx expire cached A records.
Nginx Plus does, and a few nginx plugins do, however.
https://github.com/airbnb/synapse is a process that polls DNS, and updates haproxy config accordingly and SIGHUPs haproxy I've used synapse to solve this issue, but it's a moving piece I'd rather not have involved.
HAProxy won't follow-up the TTL returned by the server. It's up to the administrator to decide how HAProxy should behave with DNS responses.
From my point of view, you don't need synapse any more if your usage of synapse is limited to this single feature.
About persistence: Does HAProxy have IP-based session persistence, i.e. routing to always the same backend server for the same client IP, as long as that server is available?
Look up "stick" functions in HAProxy docs. You can even persist TCP connections.
Please note that you can already do dynamic routing using ACLs or MAPS and updating your ACLs or MAPs content at runtime using the Runtime API (stats socket). there are "set map" and "set map" commands for this purpose.
If so, I have a suggestion: Have a look at what traefik (and before that Hipache) are doing - they use a KVS for the config and sync at runtime, which is quite convenient for dynamic usecases IMO. For the dynamic mapping it would be a better fit than just a file that needs to be manually managed.
After 1.8 releases we plan on contributing DNS SRV capability to the controller as well.
As you're the author and you're promoting your own work, may I ask what sort of testing this claim is supported on?
Seriously though if you have the time to benchmark it and compare to your favorite alternative please post the results !
hits ^hits hits/s ^h/s bytes kB/s last errs tout htime sdht ptime
32871 32871 31515 31515 11997915 11503 11503 0 0 19.1 92.2 16.5
41915 9044 20957 9450 15298975 7649 3449 0 0 23.4 128.0 30.8
43968 2053 14651 2050 16048320 5347 748 0 0 8.3 59.1 9.8
65994 22026 16494 22026 24087810 6020 8039 0 0 60.4 393.2 66.5
69176 3182 13832 3182 25249240 5048 1161 0 0 70.1 393.2 109.8
69176 0 11527 0 25249240 4207 0 0 0 0.0 0.0 0.0
69176 0 9880 0 25249240 3606 0 0 0 0.0 0.0 0.0
69746 570 8717 570 25457290 3181 208 0 0 941.4 1847.7 835.2
69746 0 7748 0 25457290 2828 0 0 0 0.0 0.0 0.0
69746 0 6973 0 25457290 2545 0 0 0 0.0 0.0 0.0
And after that it totally stops responding until I restart it. On 404 it's more around 37000. This is using 2 threads.I looked at the code and saw a select() in use so that limits the number of concurrent connections to ~512 on recent libcs (1 fd for the socket, 1 fd for the file, 1024 total). Reducing the number of concurrent connections seemed to help a bit (it delayed the occurrence a bit).
Also it doesn't seem to support keep-alive so we can't get more performance on the client side by reducing the kernel-side work in the TCP stack.
I'm getting the same level of performance out of a single process on thttpd.
Doing the same stuff with haproxy gives me 72000 requests/s in close mode as well delivering a small object with the errorfile trick, at only 80% CPU (my load generator reached its limit), so I guess there's still room for improvement since it's possible to achieve twice the performance using 80% of a single thread on the same machine, and it reaches 88k using the cache. I don't know how other servers compare though, but I definitely think that some might get better results.
The number of concurrent connections is limited by your resource limits and the fact that file descriptors are cached. You can tune the cache, but you should update your resource limits if so, this is documented here in the man page available here: http://filed.rkeene.org/fossil/home?name=Manual
It does support keep-alive and I have no idea why you think it doesn't ( http://filed.rkeene.org/fossil/artifact?ln=1130-1132%201240-... )
Using only 2 threads (way less than the default) is also very sub-optimal since most of those threads will be waiting for the kernel to deliver the file. On average filed does about 2 system calls before asking the kernel to send the file to the client -- One read() of the HTTP request, one write() of the HTTP response header, and then the sendfile() of the contents. There's no reason not to use more threads than 2 other than to limit I/O (since they do not increase CPU utilization significantly).
And it's really hard to scale using a model requiring one thread per connection. You'll hardly stand one million concurrent connections this way, and this can definitely happen for large objects.
Out of memory: Kill process 7134 (filed) score 779 or sacrifice child Killed process 7134 (filed) total-vm:7680976kB, anon-rss:7140908kB, file-rss:1988kB, shmem-rss:100kB oom_reaper: reaped process 7134 (filed), now anon-rss:7143588kB, file-rss:1980kB, shmem-rss:100kB
7 GB for 100 connections is a bit excessive in my opinion :-)
There are definitely still a number of issues to be addressed before it can be used in production, you definitely need to have a more robust architecture and request parser first.
As always, appreciate your open source work! Thanks!
I run haproxy in docker - create a shellscript that allows your jenkins (or whatever you use) to send a SIGHUP to the haproxy container, and it will reload the config.
In fact this is my biggest issue with docker - that properties like bound ports or volumes cant be changed on the fly and without deleting the container.
It would be nice if I could just do this via the HAProxy admin socket .. or not even reloading the whole config but just tell it to rescan the SSL certs.
Don't do this, please. It is a security risk as access to the docker socket can be used to take over the entire system as root.
Letsencrypt support? I realize that's a big undertaking though. There is a Lua plugin, but it does depend on certbot being installed and running, which is a bit difficult if you use the docker image (there are some docker containers on github that achieve this by running supervisord in the docker container for cron/certbot and haproxy).
This is probably outside the scope of HAProxy really though, and could probably be implemented entirely as a lua plugin that handles ACME, maybe?
Kudos to the HAProxy devs and contributors. I've always been a fan and great to see them pushing forward with these big improvements.
For now, there is one in HAProxy Enterprise, and it manages HAProxy's configuration file and triggers reload.
We're working on opening it, as soon as we have improved it: make it use both HAProxy Runtime API (stats socket) and configuration file to trigger reloads only when required.
Stay tuned as we say :)
Here I thought the one big feature is behind a paywall. Competing with https://traefik.io/ in that regard.
I already have a nice anycasted cluster of HAProxy servers that already have the battle tested BGP session teardown logic/etc. integrated - I really don't want to have to use a different load balancing method to HA my UDP services and re-do large portions of that tooling.
It's certainly an annoyance, and being able to define traffic in the same configuration file/syntax/etc. makes it much more maintainable for an ops vs. engineering group.
Chrome detects that email field, and assumes it might be part of a login form. (Because it is <input id="email" name="email" […] type="text">).
That triggers that.
All around good stuff, but I can't wait for multithreading.
As a rule of thumb, I'd say that if you have a cache, keep it. If you only have haproxy, you can try with a few rules to enable very basic caching and see if that improves the situation for you. Then you'll get a better idea about what would be required in your application to deploy an advanced cache and what benefits it could bring.
The promise of http2 is to lower latency on the client side and having H2 on the frontend will help with that.
That said, we still want to address server-side H2 in 1.9 to address the CDN type of workloads where the origin servers might be far away. But that's less important.
The server at www.haproxy.com is taking too long to respond.
Not the best advert for a high availability proxy...