The "we need to keep load on providers nearly even for business/political reasons" constraint is fairly unique.
From a purely technical perspective, you would just distribute requests inversely proportional to response time. Probably under low load, one provider would get all the requests, and only in an outage or overload scenario would the other provider take the rest.