Those are all shaping (and bidi) issues though, independent from encoding. UCS-4 is a single Unicode code point and the simplest encoding. Shaping may combine code points many to one, one to many, have ordering issues, etc. I think it is important for people to understand how shaping works and the code point/glyph distinction. Trying to duplicate the effort of the shaper when processing strings UCS-4 is going to lead to a world of hurt. Even English text with strings like "1/4" and ligatures in the font may display as a single glyph (that's font specific shaping, not encoding).
I've heard more than one person tell me they don't need to worry about text shaping since they are using UTF-8. (That statement doesn't make any sense) There is a lot of confusion with Unicode text rendering stack.