I am not a webdev, but isn't that a task for the loadbalancer in the first place?
* limit the number of concurrent requests, and drop the others;
* limit the number of concurrent requests, but queue the others (with a timeout);
* distribute all requests uniformly (randomly or in round-robin fashion) to all backends;
* (any combination of the above);
However if the "customer" asks you to not drop or queue requests, then there is nothing the load-balancer can actually do...