NGINX open sources TCP load balancing
hg.nginx.org
hg.nginx.org
In the same tense, haproxy is adding Lua support[1], which has been available in nginx - using openresty[2] - since 2011, and nginx core is doing the same with Javascript[2].
Interesting times aroung haproxy and nginx.
[1] http://blog.haproxy.com/2015/03/12/haproxy-1-6-dev1-and-lua/
[3] http://www.infoworld.com/article/2838008/javascript/nginx-ha...
As for supporting a single product, I don't see the point of that. Using Nginx for load balancing will probably be much different than using Nginx as a web server, so the learning curve is similar.
Not that I don't welcome competition, I just don't see a real need in this space.
EDIT: btw, the Lua thing was an April Fool's joke...
EDIT 2: no it wasn't, my mistake. I was surprised by this so I checked the page and jumped a gun when I saw "April 1st" on http://www.haproxy.org/news.html. Sorry about that...
Are you sure ? The 1.6 dev repo contains Lua related code [1]
[1] http://git.haproxy.org/?p=haproxy.git;a=blob_plain;f=src/hlu...
http://nginx.org/en/docs/http/ngx_http_limit_conn_module.htm...
It's not about configuration; it's about security. Fewer products in your stack means fewer things to patch. Rather than updating nginx some times and haproxy other times, you just update nginx across all your machines (both web servers and load balancers), and you're done. This also gives you more time with which to vet any given nginx update.
Kind of the reverse of the defense-in-depth principle eh? ;-)
A shared-library vulnerability means both Nginx and HAProxy get broken in their own ways, which is worse, I think, than just having your whole stack rely on one or the other, and having that one break—it's more similar to having two independent vulnerabilities arise simultaneously.
However, the scenario you describe is one where you would likely NOT be doing defense in depth, because you'd be using the same library to handle a vital piece of your security infrastructure.
Regardless, when a shared library is updated for security, you don't need to apply updates to packages using the shared library. That's kind of the point. The only exception is when the flaw is in the interface to the library.
The win entirely derives from the case of having the two independent vulnerabilities. Since they are broken in their own ways it isn't sufficient to find a way to exploit one system (which would work great for attacking a system with both). You have to find a way to exploit each, and you have to find a way to connect the two so you can get all the way through.
// Query Times
varnishncsa -F '%t %{VCL_Log:Backend}x %Dμs %bB %s %{Varnish:hitmiss}x "%r"'
// Slow Queries
varnishncsa -F '%t %{VCL_Log:Backend}x %Dμs %bB %s %{Varnish:hitmiss}x "%r"' -m "VCL_Log:SlowQuery"
// Top URLs
varnishtop -i RxURL
// Top Referer, User-Agent, etc.
varnishtop -i RxHeader -I Referer
varnishtop -i RxHeader -I User-Agent
// Cache Misses
varnishtop -i TxURL
// awesome dashboard
varnishstatopenresty well it is standalone product, it is build on top nginx with many 3rd party modules and many improvements accepted/not accepted by upstream, and of course big Lua support
Correct me if I am wrong but I think this is actually incorrect, because there is no concept of "request" at the tcp level. If I understand correctly it will rather load balance "connections".
Something to read while we wait for the announcement page to come back up :-)
[0]: http://nginx.org/en/docs/http/ngx_http_realip_module.html
i thought haproxy was used basically to proxy http stuff, not general tcp.
[0] http://nginx.com/resources/admin-guide/tcp-load-balancing/
I've used HAProxy for a long time and been very happy with it. But, everything else being equal, a stack with n-1 components is better than a stack of n components.
First, doing anything intelligent with HTTP slaughters HAproxy's performance by an order of magnitude because of the way you must configure it. Second, sticking requests to a backend is easy if you have a header that you want to decide upon. If you want to elect a different backend based upon a path component, this is much harder and yields an unwieldy configuration.
HAproxy is not designed to operate extensively on HTTP. It is designed to balance quickly and efficiently, and grew HTTP intelligence because people started wanting the convenience of making HAproxy do far more than its core focus. Rather than HAproxy getting smarter about HTTP, I'd much rather have the protocols that service my applications handle themselves and use HAproxy for its bread and butter, TCP availability and balancing. I can then focus on optimizing that using HAproxy's really clever mechanisms, like keeping the entire TCP conversation in kernel memory without reading it out (which you must do to "be powerful," as you say). This also means if I want to support SPDY or HTTP/2 or Websockets, I'm not waiting for HAproxy to support them because I painted myself in a corner.
The stack I've deployed at the frontend of every startup I've ever consulted for or operated looks like this:
/- [haproxy AZ A] -- [openresty AZ A]
[ELB] -- [haproxy AZ B] -- [openresty AZ B]
\- [haproxy AZ C] -- [openresty AZ C]
This is my Standard Frontend Deployment A. My other Standard Frontend Deployment, B, is if I have the budget and comprises Netscalers because of my experience with them from Google and other companies. Startup/low budget, ELB/haproxy/openresty. High budget, Netscaler and done.We are speaking to years of my own operational experience. I apologize if it sounded like dismissal; I actually think Tarreau would agree with my observation and opinion, if I'm perfectly honest.
1) HAproxy does support SPDY, now via NPN/ALPN. And you don't need http mode for this.
2) HTTP performance (and performance in general) is now much better. Particularly on newer kernels (IIRC somewhere in the 3.12-3.15 range, when splice was fixed for small objects).
3. I'm not sure why you found it fragile. I've run HAProxy in HTTP mode at much higher QPS volumes than you were seeing - this did require a bunch of tuning, and I'm actually hoping to have time sometime in the next few weeks to write some articles on doing this. But it worked, and worked well.
What's worked well for me is the following:
/- [haproxy]
[router] -- [haproxy]
\- [haproxy]
Using anycast routing to distribute the load over the haproxy servers, who run BIRD[1]. It's simple and pretty effective - although you want to chose your router hashing method carefully, and if you use persistance (stick tables) you have to use a recent version of haproxy, with support for peering.It is difficult to make it work, though - just how difficult I only recently appreciated when trying to help a friend over IRC make haproxy scale to very high volumes. I'm hoping that I'll find the time to write the articles I mentioned above, which will hopefully be useful to others with this sort of problem to solve.
Some users reported more than 400k requests per second in HTTP mode, or 50 times more than what you experience. Sure, here any form of HTTP processing adds a few nanoseconds to the processing time and will slightly lower the numbers. But 8k is the level of performance you should expect from tens of thousands of HTTP rules which probably is not what you're doing.
So I'm interested in knowing what trouble you're experiencing. Feel free to bring that to the mailing list, a design is always better when more people are involved.
Note that Nginx already has a simple proxy built in that does very basic HTTP load balancing. HAProxy's is vastly superior to Nginx's in that it supports a sophisticated set of filters ("ACLs"), transformations (eg., header rewriting), queue behaviours (eg., queue limits, backup backends, health checks, retries) and proxy-specific request logging.
A big difference is that HAProxy's main balancing algorithm is "fair", in that traffic is distributed evenly among target backends, whereas Nginx's load balancing is purely round-robin (there is a third-party fair balancing module [1], but it's not maintained).
I don't think that's really true - haproxy has lots of load balancing modes, none of which are called "fair". It's also very unclear what you'd mean by "fair" in this context. Least connection? Maybe, but with a traffic pattern with large volumes of very short-lived requests, least connection actually won't be very fair at all, and will end up loading some servers more than others. Roundrobin isn't really "fair" either, if you have a mix of short and long requests.
Now use HAproxy.
Now you see one example where nginx is kind of nice.
HAProxy has been able to do SSL termination OR pass-thru. It's ability to do TCP load balancing allows it to do ssl pass-thru, where SSL connections are "passed through" to other servers (so the web nodes would de-encrypt the SSL connection, rather than the load balancer). This is a good use case for some where they prefer or require data to be encrypted up to the last minute (although it's not the only way to do it).
TCP load balancing is neat for doing things like load balancing MySQL connections, which aren't HTTP (although that's not necessarily recommended according to some things I've read).
I believe, but can't find the sources, that Nginx can be as efficient a load balancer as HAProxy. I know I for one would prefer to use Nginx over HAProxy to keep my stack simpler (same technologies throughout), although HAProxy may have more advanced balancing algorithms and some more power around it's tcp socket "API" for adding/removing nodes dynamically. (I think Nginx Plus can already do some of that).
Would love to hear the opinions of those with more experience/knowledge on the differences between the two!
I want to note big difference between haproxy and nginx, first of all "nginx is webserver" (nginx HTTP server) and then proxy/loadbalancer(may be used) whereas haproxy is pure loadbalancer. Enterprise bare metal and hardware appliances for proxy/loadbalance built on top of haproxy
nginx is wide spread because of usage as good minimalistic web server
http://nginx.com/resources/admin-guide/tcp-load-balancing/#u...
Is there a use-case where the TCP load balancer being with the web-server made a lot of sense ?
Integrating too many features into a single software can be risky as it may compromise simplicity and the UNIX way, 1 tool for 1 job..
As to my original comment, you can probably get by without having full TCP load balancing.
for me, in the first place haproxy and then nginx
I good welcome opensourcing parts of nginx, take that way
http://blog.haproxy.com/2014/01/02/haproxy-advanced-redis-he...
Though I agree it'd be nice to have the functionality integrated.
Or you're referring to something else?
Just a bad title or am I missing something?
New features should be developed and tested in an open version, so the feedback, testing, patches and even unexpected new improvements from high skilled enthusiasts would be incorporated much more quickly than any closed team with QA (look at the Linux kernel).
We have seen too many examples of "acquiring" open source projects to monetize on its user base (how I hate that idiotic MBA slang) which then became stagnant - from MySQL to Xen, you name it.
I wonder what mr. Sysoev is writing these days?)
The "open core" model is horrible IMO. It pits the open source version and the commercial version againts each other. What happens when someone would like to contribute features already planned for the commercial version.. ?
Fedora became a test-bed for a new technologies, to amortize the too rapid changes (systemd and other crap, you know), so they could provide stable and compatible RHEL versions for existing customers.
And Redhat is a service, not the code company.
Or let's say that one's knowledge and skills are profitable, not the code itself.
The code should be open, so it could grow on the same pace as everything else grows, according to the real demand for features, like early FreeBSD or Linux or even MySQL grew, or even nginx before 1.0.
However, I am looking specifically at nginx. The open version of it is too good and too complete, and by its nature it's not something you are going to run externally to your application. It's what Alan Cox (I think) described as "legacy software", meaning it's a fundamental part of your infrastructure that you expect to just sit there and work.
What I am trying to suggest is the simple idea that as long as the code become closed it cease to grow and become stagnant, but while it is open to everyone, like Linux kernel, and just grows and grows, it is not possible to monetize it, and the one possible working model is to develop new features sponsored by someone, like it is the case for the Linux kernel, or to sell your own services with open code, like so many do.
It seems like the model acquire, close and sell does not work with open source projects, because when it does not grow it dies (stagnates), just according to the laws of big numbers. Or rather it changes its status to being a commercial product, which is completely different story (paid developers, support, QA, etc) to which very few people would contribute - no one wants to grow other people's wealth.