This paper is comparing AWS Rekognition, Google Cloud Vision, and Azure Computer Vision. I tried Google and Azure. I also tried clarifai.
To get 100% accuracy I ended up building my own service, and doing some really unconventional things. I used tensorflow, but I may rewrite the whole thing in pytorch.
I tried using neural architecture search but it was a dead end.
The key for me was training data distribution search.
The ML hype train has led to a frustrating amount of throwing the baby out with the bathwater, where people who should know better decide to use ML for more things than is reasonable (something something "end-to-end").
Ideally, ML should be used for a few very specific tasks, and then classical machine vision / geometric analysis / plain old logic should be doing the rest. If you don't do it this way, you eventually end up with the problems described in this paper, where performance is inconsistent and impossible to debug, and nobody can tell you what's going on or why.
Because all of the factors are known you could almost do pixel counts and compare to "true" circles (no need for AI models), any washers that didn't meet the critera were pneumatically puffed off of the line from an air hose.
Any tech that you think shows the most promise?
Like mentioned in another comment, targeted application of AI/ML can provide excellent results when the expectation is carefully defined and the objective is clearly stated.
Many AI/ML proponents jump directly to a 'black box' attitude where a product enters one side and leaves the other side and a smiley or frown face tells you if the product is good or bad. There are no definitions of dimensional tolerances nor geometric features. It's just good or bad based on the training model. But unfortunately, GIGO applies very strongly to the training model.