Running Rails on Heroku Update
blog.heroku.com
blog.heroku.com
I've played around with various routing methods before, and intelligent routing does seem to be incredibly difficult at scale. I don't think anybody gets it right? Any examples of intelligent routing at scale that works properly?
If not, might be a good place to start an OSS project into and benefit everyone. Intelligent routing should be far from impossible, all you need is very basic communication between servers and routers, possibly using some form of minimal broadcasting or multicast?
A simple solution could be having each server send 'dont send me more requests' and 'im ready for more requests'
One good notion is to embed server statistics into the response, like in an HTTP header. That way the LB can be aware of queue size, queue wait, utilization, etc. But that only works if the knowledge of server health is global to all load balancers, or if some LBs only talk to some servers.
If your requests are more or less the same "size", eg they take roughly the same amount of CPU / wall / etc, all of this gets vastly easier.
A system like this could have all sorts of knobs to turn. Requests could be partitioned into two groups: "probably fast" and "probably slow". Or three, four, n groups etc. Then the way in which the LBs distribute the requests could be tweaked. For example a ratio of 1/5 slow/fast requests per server.
This does require some feedback from the servers to the LBs. However it doesn't have to be fast. The servers could push some (request, time) pairs to some aggregating system at their leisure. Then the response time prediction algorithms used by the LBs are updated at some point. Probably doesn't have to be immediate.
1) In very large systems there are many many load balancers
2) Server "spikes" (eg swapping) can happen very suddenly, so distributing health info is a hard problem
3) Determining a priori which requests are heavy and light is also hard. Think about "SEO urls": they blow up the cardinality.
4) Even assuming perfect global knowledge, you are proposing to solve a variant of the knapsack problem.
Heroku likely has their own custom routing software that communicates with their databases to route incoming requests to the correct dynamos, etc.
EDIT: Key problem with the 'at scale' part is that Heroku does not have a single router, they have many many routers. I don't work there, so I have no idea if/why a single HAProxy couldn't handle it all, but obviously it can't or they'd just do that and save money!
Using a HAProxy with 'maxconn' flag works only for that single router instance. If a request arrives at a different router, it has no knowledge that server is wants to route the request to is already processing a request. That knowledge is internal state in a different router.
The key here would be to communicate this state between all routers. As I thought in the OP, the correct way to handle this at scale is some form of multicast - across a 1gbps network, broadcasted status updates should account for a tiny percent of traffic/processing, yet allow intelligent routing that would decrease latency immensely.
So I guess the original question still stands: is there any router project that communicates state between multiple load balancers/routers to allow intelligent routing at scale ?
EDIT2: Here we go - looks like research is being done in this area. Maybe Heroku should read this paper:
http://digitalcommons.unl.edu/cgi/viewcontent.cgi?article=10...
Yes. Consider that Heroku publishes DNS records for all their customers to the same set of nodes; those nodes must be responsible for routing all Heroku traffic. Sharding customers onto individual Haproxy LBs would require either updating an upstream load balancer (which just pushes the problem around) or updating DNS (which comes with long convergence times).
Consider the following scenario:
- Two types of requests, A and B
- Request A takes 200ms to process
- Request B takes 10 seconds to process
- Each dynamo can take 4 concurrent requests
If you are receiving 1000s of requests per minute, it is very likely that you will eventually allocate more than 4 request B to a single dynamo. One this has happened, that dynamo is now locked for 10 seconds. All of the request As that we route to that dynamo will take over 10 seconds to return.If request A is a credit card transaction, and request B is just some data lookup with a loading UI spinner, then every time this occurs our poor app has lost money as the user has navigated away from our app before the transaction (request A) can complete! Ouch!
Intelligent routing solves this by ensuring that request A will only go to a dynamo that is open, thereby ensuring no user will have to wait 10 seconds for their quick credit card transaction.
The take away here is that the important measure is the slowest user facing request, not the average request time across all requests.
Oh, wait, Heroku already tells you to do this: the B-type requests are why they introduced "workers." Heroku's recommendation has always been that you're supposed to write your "app servers" to serve A-type requests directly, while asynchronously queuing your B-type requests to be consumed by workers. You then either poll the app server for the completion state of the long requests, or have the worker queue a return-value paired to the request-ID, that the app server can dequeue and return along with the next request. (Erlang's process-inbox-based message-passing semantics, basically.)
To put it another way, it's the old adage of "don't do long calculations on the UI thread." In this case we have a Service-Oriented Architecture, so we've got a UI service--but we still don't want it to block. By default, Heroku basically exposes Unix platform semantics; unless you wrap those with Node or Erlang, you have to deal with how Unix does SOA: multiple daemons, and passing messages over sockets. Heroku could "intelligent-route" all they want, but there's no level of magic that can overcome applications designed in ignorance of how concurrent SOA architecture works on the platform you're designing for.
Note that this overlooks the fact that a node that has requests queued only drains them at a certain rate, as long requests tend to pile up. However, if your framework is truly concurrent and your nodes are not CPU bound, but rather waiting for other services, it's possible to get much higher concurrency than just 20.
It's possible to do several hundred requests per second on a single dyno; just because Rails doesn't allow you do to that due to not being concurrent doesn't mean that other stacks can't.
Still, there's a resource limit to how much concurrency you can achieve within a dyno as opposed to by buying more dynos. You shouldn't have to scale horizontally just because the routing scheme is inefficient with the width that it has.
If Heroku really wanted to abandon Rails because other web stacks are easier for them to scale with, they should have let us know rather than turning around and shitting on the platform that made them as a business.
Concurrency is the ability to do work on multiple things at once, while parallelism is the ability to execute multiple things at once. For example, an operating system running on a uniprocessor is concurrent, but not parallel. Another example is HAProxy, which is highly concurrent but not parallel(it can handle several thousands connections, but is single-threaded).
The distinction is important when you're trying to scale. Adding more threads/processes does not help, as you quickly reach OS limits(eg: C10k problem). Having a concurrent web stack(nodejs, EventMachine, Play, Xitrum, Twisted, Yaws, etc.) does, as it allows extremely large concurrency with a limited resource impact, whereas adding more processes quickly hits memory limits(at least on Heroku).
And even at that point, you're still using more dynos than you would with intelligent routing.
Leastconns routing falls flat in many cases, such as long polling/streaming/websockets, which count as a connection but barely take any resources. In the case where you would have a load balancer that did leastconns, the servers with many long poll requests would end up underloaded(ie. doing nothing) while the ones with few requests would end up serving most of the application.
Maybe you need entirely different load balancing strategies for different designs of web application, which means Heroku's promise of a single infrastructure stack for everything is bogus. But I'm skeptical that they deliberately chose to favor evented frameworks rather than choosing what was easiest for them to implement.
If your web app can only process one thing at a time, everything is expensive. Someone made a database query in the admin interface that runs for 20 seconds? Oops, hope you have other servers. Ten 100ms requests queued? The last guy has to wait 1000ms for his reply.
If it can process things concurrently, it doesn't matter, as long as you're not doing something that uses all the CPU(which shouldn't happen, assuming a sane design). Someone made a database query in the admin interface that runs for 20 seconds? Doesn't matter, you still have n-1 database connections to process the other queries. Ten 100ms requests hit your server? That's fine, the last guy will have to wait maybe 120ms for his reply.
If your app cannot process things concurrently, it should not accept connections for further work if it is working on something, period. This shifts the load balancing back onto the load balancer, because it has to find a worker that can process a request. And guess what, random assignment works fine in that case.
I don't see how forking is insufficient concurrency for the cases you mention, either. Even the Ruby GIL will allow threads to switch while blocking on I/O, so your ten second DB query is covered.
But there's no free lunch--if 2000ms of CPU ends up on one of your servers, and your median request is closer to 100ms, the unlucky server that gets the 2000ms request is still a little fucked without any mechanism to counterbalance. Even letting the server stop accepting new requests from the LB, as you suggested, would be sufficient.
It does and it's in the configuration for your app, not on the Heroku side of things. For example, if you're using Unicorn, check the documentation for :backlog, which controls the connection queue depth of the web server.[1]
> I don't see how forking is insufficient concurrency for the cases you mention, either. Even the Ruby GIL will allow threads to switch while blocking on I/O, so your ten second DB query is covered.
Can it handle 10, 20 or 100 long database connections/web service queries? I do admit I don't know much about Ruby threading and per-thread overheads and in that particular case, green threads are perfectly fine.
> But there's no free lunch--if 2000ms of CPU ends up on one of your servers, and your median request is closer to 100ms, the unlucky server that gets the 2000ms request is still a little fucked without any mechanism to counterbalance. Even letting the server stop accepting new requests from the LB, as you suggested, would be sufficient.
There shouldn't be any requests that take that much CPU time, unless you're doing something ridiculous like sorting millions of integers or whatnot. If your application requires sorting millions of integers or any computationally expensive requests, for some reason, those should be handled in a backend queue and sent to the client using long polling/comet/websockets/meta refresh so that your front end can worry about delivering pages quickly and shoveling data between the backend and the client instead of crunching numbers.
[1] http://unicorn.bogomips.org/Unicorn/Configurator.html#method...
do you have any proof/calculations?
Sure. http://aphyr.com/posts/278-timelike-2-everything-fails-all-t...
There has to be some happy medium between "embarrassingly parallel but shitty performance" and "intelligent routing that can't scale".
If you would extend the whole thing to have a negotiation protocol that lets the load balancer know how loaded the individual backend processes are then you could start talking about intelligent routing. That however currently does not exist in the Ruby world.
I'm inclined to give them the benefit of the doubt. They made a reasonable engineering and product tradeoff to increase their system's scalability, decrease it's latency, and allow new kinds of applications that are able to serve more than one request at once (eg. node.js servers holding onto tons of idle websockets).
Given an increasingly large infrastructure (which makes stronger queueing much harder if you don't want to add a lot of latency) and strong requests from customers to allow other languages that don't fit into the "one fast request per process" paradigm, I can absolutely see making that tradeoff. It makes the infrastructure simpler (and thus more robust), cuts some latency, and makes the many loud customers who wanted to run things like node.js happy.
Doing so, I would have been aware that in some cases, this would lead to suboptimal routing that increased latency, but I never really would have thought it would have as devastating effects on an application's overall performance as Rap Genius saw.
So yes, they do care a lot more about this issue now that there has been a big stir about it, but I think it's reasonable to assume that this may be more because they've realized it can have a huge impact on overall latency rather than a moderate one, and that it's come to their attention that some of their messaging and docs hadn't been updated reflecting the routing change.
But given how great of a company Heroku usually is, the fact that it is not in their interest to decrease their quality of service, and that there are very legitimate reasons to move away from intelligent routing (namely, the facts that it doesn't really make sense for applications that can do multiple requests concurrently and that it is extremely difficult to scale), I'm inclined to believe that this is more a mistake than a malicious lack of concern for their customer's well-being that they are only addressing due to vocal outside pressure.
(Something like my app name is heroku.ruby.free.cheapcustomers.myapp )
There will be lots of header handling but it seems simpler
The only difference is the URL here has a trailing question mark.
Why does this keep happening?
Or one that allows only k concurrent requests with a routing mesh that schedules m>k per server (dyno) or at least does that often enough.
Even with unlimited concurrent requests per server, similar problems may arise if the performance drops sharply with the number of concurrent requests being served (e.g. due to swapping.)
Fulfilling these commitments is Heroku’s number one
priority
I think this is the part that really breaks the camel's back for me.I don't doubt that dealing with the biggest PR disaster in the company's history is their number-one priority, don't get me wrong.
EDIT: It's probably better to make your opinion known on the blog post than here, though.
"Hey all our customers, We're terribly sorry for all the black magic stuff
we did (the shitty Newrelic reporting, etc.) and we've refunded your account
with $100*x for all the losses that we knowingly caused you and your business."
Unfortunately, that doesn't seem to be the case with this blog post.I'm not going to make it sound like this is something that can be fixed and mended by hiring some copywriters and PR consultants, but they're doing a very poor job of making me personally give them the benefit of the doubt.
It's like the asshole who says "I'm sorry you feel offended", or "I'm sorry I got caught". There isn't even a semblance of an apology and understanding of the consequences of this for their customers.