Based on a quick skim, what they're describing already exists in the form of a feature extractor + traditional RL algorithm. State -> features/symbols -> actions as opposed to state -> actions.
So, sure if you split up your decision making into more pieces, then you can specialize those pieces better to your specific task, then it will probably perform better at that task. But the whole point (and success) of deep learning is the feature extractor and decision model should are one network, where the early layers are acting as feature extractors (symbolic component) and the later layers are acting as the decision makers or value learners.
In my mind, this partitioning strategy runs counter to the greater goals of machine learning, which to me is a generic learner that can 'take raw data in' and 'get smart stuff out'.