Would you mind giving an example of when this would be useful?
Would you mind giving an example of when this would be useful?
If you're looking for specific examples of "large alphabets", there's a very easy one in my mind, which is UTF-32. If you're looking for data that isn't a plain sequence of integers in any shape or form... I have to think back, as I don't remember what the use case I ran into was, sadly. I'll post here if I remember.
For most DFA engines I've seen, they take advantage of "strings" themselves and define the alphabet to be the set of bytes. If you're searching UTF-8, then this also implies that your DFA must include UTF-8 decoding (which is generally the right choice for performance reasons anyway).
Regex engines with modern Unicode support do generally need to support efficient set operations on large sets though, in order to support the various character class syntaxes (which include negation, intersection, symmetric difference and, of course, union). Usually you want some kind of interval set structure, and its implementation complexity, in my experience, generally correlates with the amount of additional memory you're willing to use. :-)