Deep reinforcement learning will transform manufacturing as we know it
techcrunch.com
techcrunch.com
Most production planning problems are inherently high-dimensional and come with complex constraints. When you add noise to the problem, like random demand, all generic RL algorithms ran into trouble once you added complexity (like producing multiple products on the same line). So if you’re looking for good examples to break RL codes, throw them a nice stochastic economic lot sizing problem with sequence-dependent setup times.
By constrast what works well is policy search over parametric policies that consider domain knowledge. Sounds boring, but is quite hard to beat.
Our team at Pathmind has applied deep RL to multiple use cases in industrial control and supply chain management, notably various forms of scheduling, and MEIO, respectively.
You can see some demo simulations here:
https://pathmind.com/examples/
https://cloud.anylogic.com/model/15ee2fe9-c45c-425d-8b98-920...
https://cloud.anylogic.com/model/9c0725cc-7eb5-44ed-a2de-ff8...
https://cloud.anylogic.com/model/8769d942-43c4-4530-bc68-243...
https://cloud.anylogic.com/model/b7da42a0-734c-461f-9f68-382...
Yes, these show RL applied to simulations, and yes, real data and real physical plants are more complex than that.
But those physical systems are already controlled by optimizers (usually mathematical solvers like IBM Cplex or Gurobi), and deep RL happens to be an optimizer that can handle a few things better:
* data variability
* multiple objectives in complex scenarios
* multiple agents making simultaneous decisions in coordination
Solving the third problem gives rise to very interesting emergent behavior among teams of machines, which learn to behave in ways that are almost impossible for a hard-coded set of rules to specify. This is the source of many of the gains that RL produces.
I see a lot of preconceptions about RL in this thread that are partially true, but also solvable.
Yes, RL is data hungry. I acknowledged that in the TC piece. The problems with data are pretty much the same problems you get with ML in general. It's messy, partial, and usually not gathered with ML in mind. RL is no more impossible than other ML approaches, based on that argument.
A valid simulation of a physical system is a source of synthetic data that helps solve that problem.
Deep RL is already controlling machines on factory floors, and it will slowly help optimize larger and larger systems, from groups of assembly lines, to networks of plants.
In some cases it will serve as a "decisioning layer" (not my word) with HITL, and in others, it will drive more autonomous systems with low latency requirements.
Believe me or not, it's happening.
Without getting into too many details, the underlying algorithms are a flavor of policy gradient implemented with the open source projects Ray/RLlib (supported by AnyScale) and Gym (from OpenAI). They're great tools and I highly recommend them to anyone who wants to explore deep RL.
Fwiw, DeepMind says RL may be the path to AGI, and smarter people than me have placed their bets on that team, which has consistently surpassed expectations about what RL can do.
https://www.sciencedirect.com/science/article/pii/S000437022...
Microsoft is making a major push into this area using their Bonsai acquisition to power autonomous systems.
https://www.microsoft.com/en-us/ai/autonomous-systems
AWS, while not particularly focused on RL, is making a big push into operational technology (OT) as well.
The incentives driving their cloud business will spur them to transform the software that controls the factory floor, or at least prompt existing OT software makers like Siemens to move faster.
If you have any questions about how deep RL is being applied, my email is in my profile. Pls hit me up!
I suspect DRL will become extremely useful once we figure out how to use learned priors of the physical world to build more and more complex systems.
It is great to see that you are using Anylogic which I think is an excellent workbench for creating discrete-event models of production systems and entire supply chains even. Anylogic also comes with an integrated tool for parameter tuning using global optimization.
Now, the model that I was looking into was a stochastic economic lot-sizing problem with random orders. At the end of a production run or upon arrival of a new order, the policy decided which product to produce next on a single machine. Each change however incurs switching cost and any orders that met an empty inventory were lost (the backlog case was easier).
The difficulty with this problem was noisy rewards after each state transition due to order randomness, as well as for the training algorithm to communicate the future value of an action in a particular state far enough backwards in time. As said, it worked for easy problems but not for problems with 5 products or more (which I think is a ridiculously small and simple problem).
Now, I do not know what you guys are using, whether your policy is a DNN and you are tuning parameters using REINFORCE / evolution strategies or whether you are using value function approximation with fitted Q-iteration (FQI) or similar (like DeepMind's Atari players). While I think that policy search may actually work (despite rather slow with Anylogic), I'd expect that FQI will run into trouble on more difficult problems.
That said, deep RL is a great piece of tech, but I am a bit skeptical of the claim that RL will transform the future of manufacturing, as I have made the experience that it can easily struggle even with simple problems as soon as the action space becomes high-dimensional and reward information must be communicated across hundreds or even thousands of state transitions.
Btw: are there any scientific papers on this?
Deep RL is one of the most data hungry methods. And in physical systems you don’t get enough samples.
There are loads of data. But no manager of a shopfloor lets you produce scrap just to potentially learn.
There is an opportunity to combine ML with classical engineering models for manufacturing. Think differential equations for chemical processes. Then you end up with something like ML augmented control theory. That does work.
In some cases, you could be computationally bound and never achieve a sim representing your desired environment. In many, you lack theory to correctly build a sim. You need RL building the sims that you plan to use RL to explore optimal processes you seek within. We have a few disciplines that are well formulated enough where sims can be useful in but the vast majority are simply going to give you garbage or are just computationally bound. At best, many sims provide guidance that experts need to interpret.
What's good for AWS is good for America.
It's true that the process is data hungry. But this isn't such a big problem in some cases. Having a few dozens or more of robot arms just figuring out stuff by trial and error isn't such a big deal (you don't have to do this at the customer's shop floor...). For self driving cars and drones this isn't such a great idea (though it has been tried to some extent). That's where simulation and more clever algorithms come in. Including, of course some prior knowledge, like you are implying. I don't think anybody expects to "transform manufacturing" by just throwing REINFORCE into a random KUKA arm.
And even with the non-deep networks from back in the day, people felt compelled to include prior knowledge and even used Bayesian priors to deal with a lack of data.
[1] http://papers.neurips.cc/paper/447-neural-control-for-rollin...
On the other hand, if it hasn't been automated yet, who is going to invest hundreds of millions into a Deep RL system (never mind an army of robots) when they can just pay a slightly higher variable fee for humans. For example, sure a robot could theoretically sew jeans for $0.20 vs a human at $15.00 (living wage in US), but does any one firm have the scale to justify $500M to train Deep RL?
I see what you did there :)
A parallel - facial recognition today is a matter of `npm install face-api.js` (or some other library). Advances in hardware and ML architecture design can bring the currently-challenging world of RL to similar levels.
Another way to put my argument is through a question: do you think that even 20 years from now, the cost of a useful RL system will be around $500M (inflation adjusted)?
And that's not even accounting for the cost of robots (which will need frequent repairs when the untrained policies crash into physical objects), salaries for your engineers and all the raw materials that you're going to have to scrap because the robots didn't make saleable product.
In the world of supply chain and industrial control, people typically fight hard to get an additional 1% gain in optimization. Deep RL can often surface decision paths that lead to double-digit gains, and those gains, for many companies, will be worth a lot of money.
Fwiw, it does not cost anywhere near $500M to train deep RL.
1. https://pytorch.org/tutorials/intermediate/mario_rl_tutorial...
Personally I expect all real-world applications of RL to be trained in simulation, with tricks to make sure they also learn to adapt to reality at startup (meta-learning). For example by simulating each episode with parameters that are slightly off.
The bigger blocker right now is the simulation part itself (and there is also a lot work on that!) Training a DRL for real world physical processes ob real world processes has the Problem, that a failure is extremely expensive. You don’t want to see a manufacturing process moving too fast (because it’s still learning / optimizing) and then breaking a 500k router spindle.
Simulation is certainly one of the trajectories to deal with this though.
The third big block is then really the cost: for many manufacturing processes there are pretty good parameters found over multiple years. The trade off of the costs of a new DRL system - training it, potential failures, deploying it and training the users - to the gains need to be big enough to justify the use of DRL financially. Again standardization will help with this, but it requires significant R&D costs upfront only few are willing to pay.
In the end, in my opinion we will see the application more often, but it takes time and effort to improve and make it cost effective.
Large existing manufacturing firms with the kind of money to do this can't get out of their own way, and will never be a source of innovation. Small companies don't have the funds to throw at maintaining a staff of genius engineers and researchers.
Most outsiders with access to capital looking at innovating in the space probably don't even appreciate the difference between a job shop and an OEM, and are likely to lose lots of cash on projects headed by former OEM enterprise middle managers who exclusively have experience getting in the way of production and wasting money.
I'd love to be part of a highly funded team of smart people whose job was to reinvent manufacturing in a practical environment.