One good notion is to embed server statistics into the response, like in an HTTP header. That way the LB can be aware of queue size, queue wait, utilization, etc. But that only works if the knowledge of server health is global to all load balancers, or if some LBs only talk to some servers.
If your requests are more or less the same "size", eg they take roughly the same amount of CPU / wall / etc, all of this gets vastly easier.