I’ve made the mistake of simply splitting on spaces or punctuation, and eventually finding that will do the wrong thing.
Eventually I learned about Unicode text segmentation[0], which solves this pretty well for many languages. I believe the default Lucene tokenizer uses it. I implemented it for Go[1].
[0] https://unicode.org/reports/tr29/ [1] https://github.com/clipperhouse/uax29