In science, when there's a two orders of magnitude difference, it's more indicative of distribution mismatch between the test sets. The logarithmic property of the learning curve suggests that, unless one approach is a completely random walk, two learning algorithms after a period of training must converge to their maximum capacity. As a scientist, I'm refused to be clouded by my prior judgement of the companies behind the approaches and must question the nature of the metrics and the various definitions that were used.