Usually, we destructively compress (mean-pooling) both the query and the document, and then compare the two compressed forms.
With ColBERT, we compare first - at the full token level - for more detailed comparison. Then reduce the full set of comparisons to a single vector. Naturally this takes more memory and compute to do the more comparisons. The idea is it’s worth it because the more detailed comparisons lead to better results
tokens —> reduced vector —> comparison
Vs
tokens —> comparisons —> reduced vector