Each unicode character has certain properties, one of which is whether that character indicates a break before / after itself.
I've done extensive research on this for my job, but unfortunately don't have time to do the whole writeup here. Here are several resources for those who are interested
Info on break opportunities:
https://unicode.org/reports/tr14/#BreakOpportunities
The entire Unicode Character Database (~80MB XML file last I checked)
https://unicode.org/reports/tr44/
The properties within the UCD are hard to parse, here's a reference if you're interested:
https://unicode.org/reports/tr14/#Table1
https://www.unicode.org/Public/5.2.0/ucd/PropertyAliases.txt
https://www.unicode.org/Public/5.2.0/ucd/PropertyValueAliase...
Overall, word / line breaking in Unicode in no-space languages is a very difficult problem. Where the UCD says there can be a line break isn't where a native speaker would put one. In order to do it correctly you have to bring in Natural Language Processing, but that has its own set of complexities.
In summary: I18N is hard!