Wow, so characters like U+022F (ȯ) and U+042F (Cyrillic letter Я) U+062F (Arabic letter د) are not allowed but nearly everything else is? Some of those are letters used in actual languages. That's sure to make people scratch their heads.
codepoint -> encoding in UTF-8
U+022F -> C8 AF
U+042F -> D0 AF
U+062F -> D8 AF
Remember, UTF-8 is self-synchronising: when you pick up at a random point within a stream, there is no ambiguity as to whether you are in the middle of a sequence or not. Valid lower codepoints appearing in the encoding of higher codepoints would violate this property.