A system like this could have all sorts of knobs to turn. Requests could be partitioned into two groups: "probably fast" and "probably slow". Or three, four, n groups etc. Then the way in which the LBs distribute the requests could be tweaked. For example a ratio of 1/5 slow/fast requests per server.
This does require some feedback from the servers to the LBs. However it doesn't have to be fast. The servers could push some (request, time) pairs to some aggregating system at their leisure. Then the response time prediction algorithms used by the LBs are updated at some point. Probably doesn't have to be immediate.