[1]: https://www.google.com/search?q=(160+bits+%2B+32+bits)+*+16....
Include in the extension a hash table construction as follows:
foreach ID of an HN story submission
URL = the URL of the submitted story
URL = normalize(URL)
insert_into_hash_table(URL, ID)
insert_into_hash_table(key, val) is a function that inserts val into a hash table with key key. The hashing function does not need to be cryptographically secure.normalize(URL) is a function that takes a URL and normalizes it. What normalize means in this context is a little fuzzy, but the basic idea is that if URL_1 and URL_2 are different URLs to the same article, normalize(URL_1) == normalize(URL_2).
NOTE: what you include with the extension is the hash table itself. Conceptually it is probably just a sparse array containing HN IDs, with maybe a little more depending on how collisions are handled.
In the extension, do this:
URL = normalize(URL_of_current_page)
ID_list = lookup_URL_in_hash(URL)
foreach ID in ID_list
story = get_HN_story(ID)
if (normalize(URL_of_story(story)) == URL
show_story_comments(story)
The hash is only used for data retrieval from a local hash table, so does not need to be cryptographically secure.After the hash lookup we have a list of candidate stories on HN that might match the browser story. It's a list because due to collisions there might be more than one HN story with matching hash.
Note that all that is ever fetched from the server during operation of this are HN stories, so there is minimal information leakage.
For instance, if we had a hash table whose keys were names, and whose values were telephone numbers, we'd probably have to store keys with the phone numbers so that in the case of a collision we could figure out which phone number matches the search key.
If, on the other hand, we had a hash table that stored record numbers of employee records from our employee database, keyed by employee name, then we probably would not need to store keys with the hash table values. If there is a collision, we can just retrieve all of the colliding records from the database. Those records will contain the employee name, and we can use that to figure out which is the right one.
For the HN comment extension we are closer to the second case. The HN story contains the URL, so in the case of a collision we can fetch all the colliding HN stories and see which one is the right one.
Edit: Estimating quickly from the new submissions for the last couple hours, it looks like there are on the order of 100 submissions per hour. Of these, it seems like less than 10% have any comments. So if we assume that HN has been up for 12 years (100,000 hours) we get about 1 million commented URL's with comments, and that's probably an overestimate (dups, increasing usage, etc).