> It would be nice to be able to define a prior that puts very low to no probability on actions that are not physically possible. This would constrain the search space considerably. But it's not really clear how to do this with neural nets.
I dunno, I can think of several ways off the top of my head: 1. put large negative rewards on impossible actions; 2. only compute Q-values for feasible actions (just because DQN computes Q-values for a hardcoded set of actions doesn't mean you have to); 3. use Achiam's "Constrained Policy Optimization" https://arxiv.org/abs/1705.10528 ; 4. reparameterize the output to make illegal actions unrepresentable (domain-specific); 5. add illegal or not as an additional data input (ie in addition to the usual image or robot state, for the _n_ possible actions, include a _n_-long bitvector of possible/impossible), or include that as a regression target to make it predict whether each action is possible.