It's clear that the next frontier is to have 3D-space instead of image space transitions. Language itself is very static and action verbs are not enough to specify scene dynamics. I suppose we would need:
A. an enriched version of natural language that refines the dynamic processes that occur in a scene
B. a data set of isolated processes labeled in the language described in A.
I've had a hard time finding ongoing work on A. and B, perhaps it isn't much of a priority for research groups.