Every scheduler node has cached view of whole cluster and optimistically makes a scheduling decision, retrying on conflict?
Any tricks you did to reduce conflict rate? Is there a certain cluster saturation threshold (little free capacity) where conflict rates would get too high?