So it turns out that these metrics don't exactly follow a Gaussian distribution, so it's hard to get these algos to work right out the box. Additionally, speed was an important component for us (for training and evaluation), so we had to toss out a lot of the fancy, but slower, techniques.