Seems like a good place for a bloom filter.
If the overwhelming majority of words seen don't match, you would win by rejecting non-matches quickly. If instead it almost always matches, you will need to look at all the bytes anyway to be certain of a match.