Show HN: Subsequential Transducer for Efficient Text Rewriting
github.com
github.com
The reason being that in a naive approach, a vocabulary of size M and a document of token size N is an O(m*n) operation in the best case - while this is claimed to be an operation of O(2n).
Not knowing the specifics, I would lean toward a template library, as it would be more precise, and could account for things like changing gender, plurality, a vs an, etc, in a structured way.
Wouldn't the vocabulary size fit into the order complexity? Vocabs that would be considered useful in this context tend to be quite large. Are you achieving worst case logn of search in the vocab?