What we have here is a population (interviewees) that we want to sample according to some metric. We want to apply some series of filters that limit the sample to only candidates that exceed some threshold of suitability for a task.
After forming a model for that filter, we apply it and begin to suspect that we're not extracting the optimal subset of interviewees.
In any other scenario, the answer to this lack of fit would be to change the sampling method/filters. But for some reason, unique to the interview process, interviewers decide the solution is for the population of interest to "study" at passing our filters. In other words, we find we have a bad measuring instrument and instead of designing a new instrument, the reaction is to insist that the samples work harder to fit the bad measurements.
I'd like to propose if qualified candidates have to study for your "test", you're test isn't very good at finding qualified candidates.
By the time a candidate makes it to an interview, initial screeners should have eliminated those that haven't proven their ability to study for a test when given the parameters of that test in advance. Try coming up with better questions to determine if they're a good fit for your company.
...
Now excuse me while I go re memorize all the big-O worst/best case performance for algorithms from my first year of undergraduate instead of just deriving it out or looking it up in a table like we all know we do in reality. (unless we use it every day)
edit
Oh, and the tool looks really fun and will probably help in said interviews. ;)