I suspect that's the real fix. Now all those (16k) bound addresses aren't creating hash table entries, so other connections that happen to use a port that hashes to 21 (or 53 after enlarging the table) aren't being shoved into a hash bucket that starts with 16k entries already in it.
The enlarging of the hash table I think is less a fix for this problem (although it would halve the number of later connections being put in the bucket), and more just a good fix they happened to do at the same time.
Yes. It just reduces the risk that they run into this problem again with a different port constellation.
It does, however, reduce the overall impact when all connections are considered.
It works out the same for the application: 1 fd or 16k fds doesn't really matter if you're using epoll, and that single fd can accept connections to any of those 16k IP addresses.
Now, from a server point of view there are two types of connections: inbound and outbound. Our servers accept connections but they also establish connections, for example to your http origin hosts.
So from the point of view of our server the "colliding" packets will fit two categories: A) incoming packets to port 53 B) incoming packets to outbound connections which source port % 32 == 21.
For A) this is not that a big deal. DNS usually works over UDP, there are not _that_ many DNS queries done using TCP.
For B), since Linux choses source port incrementally, that means every 32'nd connection will possibly have some packets hitting the unhappy bucket.
Therefore increasing the hash size twice, reduces the chance of collision twice: now every 64th outbound connection will have some packets hitting the unhappy bucket.
The full answer is: depending on which RFC you read :)
Initially the RFC's specified that you could only use TCP if you got UDP truncation _first_.
Nowadays that's relaxed but it's very vague when you should use TCP except for after UDP TR. For example Bind will try to connect over TCP if UDP fails.
Generally speaking most of the traffic goes over UDP, and sometimes, in undefined circumstances, some stuff may be requested over TCP. No hard rule.
But yes, it seems like an unnecessary change.
Increasing the hash size doesn't fix the unhappy bucket, but it does reduce chance that packets will ever hit it. So yes, traversal of this bucket will be slow, but it will be hit less often.
In particular, a rewrite would have to make sure not to make the general case worse in an attempt to avoid this pathological situation.
For the 2nd level you could size the hash table appropriately since you always know the maximum number of IP addresses a host has.
Would be nice if this 2level array/hash was tunable from /proc or /sys.
Using a binary tree at each bucket could also work well but you would have to rebalance the tree periodically if listeners were inserted in sorted order. A self balancing tree could be used instead but then again this adds complexity.
Destination IP is a reasonable addition to the hash function, however.