This, over-dividing problem, along with the under-dividing problem you mentioned, are both huge hurdles for machine textual understanding and machine translation systems.
On the information-extraction system I'm working on now, roughly 80% of the entities we're trying to extract are multi-word expressions. Very difficult.
[1] https://en.wikipedia.org/wiki/Multiword_expression
[2] http://aclweb.org/aclwiki/index.php?title=Multiword_Expressi...