Remember the "data items" in this case are bytes. Lots of them in a cache line.
That is, an unrolled loop on bytes can be N conditional increments on an index, with a check at the end to see if still in the bounds. Assuming N is small, will be hard to compete with that, honestly. The hash would be spreading the data wider than the linear search would be. Though, I agree I'd expect it to still be on a cache line for bytes.