First, we don't see the spam that was outright blocked, so we have no idea of what the false negative rate is. Second, I bet that this is not a binary block/allow decision, but there are all kinds of ways of reducing the engagement that probable spam gets without outright blocking. The latter is operationally preferable since it reduces the cost of false positives and since it makes the iteration loop for the spammers a lot slower.
(But also, I can't remember when I last saw spam in my Twitter feed.)