Thanks for sharing! What sets Manticore apart from Meilisearch and Elasticsearch is that it lets you configure tokenization at a low level by:
- choosing which characters should be treated as token characters, and using the rest as token separators
- defining "blend chars" — for example, the hyphen (-) could make sense as both a separator and a non-separator in your case
- or optionally adding it to the ignore_chars list
- there's also regexp_filter to process tokens when indexing and searching
That said, setting things like this up perfectly is always tricky with any search engine, because the words and punctuation in real data often don't follow regular patterns. It's especially difficult when you want to find "abc def" by "ab cd ef" which may be a common situation in your case.