OpenAI Baselines: ACKTR and A2C
blog.openai.com
blog.openai.com
Is it really just that you can't easily test your program?
2) Many hyperparameters - it may be critical to get them right
3) No gradual implementation - it doesn't work until you get all the (often gimmicky) parts right. Take A3C for example - its paper version is parallel, may be hard to implement and debug in your language (and hardware) of choice.
4) Task and reward function choice may be hard - take a look on my question https://stackoverflow.com/questions/44781401/is-openai-gym-c...
Generally, RL algos performance is underwhelming in my experience, without heavy tuning (and so-called "domain expertise" in reward function), but I'm not an expert and OpenAI guys show that you can make the working thing like Dota2 bot, so they give me hope.