Interclass variation amongst rubik cubes is far lesser than intra-class variation amongst different objects. I would put that in the same category as dialogue systems generating responses using RL ( I've actually worked on this ), incidentally also on top of the page today. These things are mostly dog and pony shows. RL's application in practice is limited, even for Covariant .ai I'm not convinced that most of the paper they have posted is not just marketing. Academics are quite good at playing this game, see this paper on their site[1]. It has nothing to do w/ grippers. In practice once you touch hardware, all that policy gradient goodiness goes out the window, hardware considerations, domain specification, etc will dominate how good your solution is. So "domain spec and design" means you need to work closely w/ your customers, and have a say in how their warehouse is designed. Amazon doesn't run into this issue because everything is done in house. But if you try to deploy the same system at Walmart w/o strong institutional support, the system will fail.
Thus companies such as this is a pump and flip play. There's a direct comp that's Amazon's internal division, everyone is trying to copy Amazon's supply chain efficiency these days, so the best outcome is strategic investment from Walmart and such, and then followed by an acquisition.
You saw similar companies come out in the nascent days of the deep learning hype, socher from Stanford comes into mind. His company MetaMind was sold to Salesforce for xx million amount, all the engineers got a fine pay day, they didn't really release a product. But they certainly published some nice papers along the way.
[1] https://openreview.net/attachment?id=ByeWogStDS&name=origina...