Hi, appreciate your comment. The sampling is from all posts / comments over the past 35 days, accessed via the API (https://github.com/philippdubach/hn-archiver). There might be a skew to sample higher voted posts first (i.e. if there is high volume posts and comments with zero upvotes don't make it into the database) so that would explain the high ration. I will definitely look into it before publishing the paper - this is exactly the feedback I was hoping for publishing the preprint. Thanks for pointing this out! Would love to see the mentioned classifier. If you find the time please reach out to the email on the page or on bluesky.