HNHacker News
TopNewBestAskShowJobs

costco

2,441 karma · joined December 30, 2019

chrisjtarry \at\ gmail.com

https://github.com/chris124567

submissionscomments
costco··on Show HN: I scraped 3B Goodreads reviews to train a better recommendation model
I would agree the results are generally OK but do not feel magical in most cases (I think in some specific cases they do though). The results can be not great if you add books across many disciplines. For instance if you add "The Elements of Typographic Style" and "The Design of Everyday Things" (https://book.sv/#671857,18518), you do get "Grid Systems in Graphic Design" but under its German name "Rastersysteme für die visuelle Gestaltung."
costco··on Show HN: I scraped 3B Goodreads reviews to train a better recommendation model
> Note 1: If you only provide one or two books, the model doesn't have a lot to work with and may include a handful of somewhat unrelated popular books in the results. If you want recommendations based on just one book, click the "Similar" button next to the book after adding it to the input book list on the recommendations page.
costco··on Show HN: I scraped 3B Goodreads reviews to train a better recommendation model
I use `fetch` on relative endpoints so that's odd. There shouldn't be any external API calls on my website other than whatever the Cloudflare captcha uses. I also use HTTPS-only in Chrome and did not experience any issues. I just tested Firefox with HTTPS-only on/off and Safari on my phone and I was able to import shelves for multiple users. Are you sure that you do not have any privacy settings on (can you access your shelf in Incognito mode)?
costco··on Show HN: I scraped 3B Goodreads reviews to train a better recommendation model
I use a Hetzner server with Ryzen 7 3700X and an SSD.

I think I could get the model to work with ONNX web but it'd be a 2GB download so the user experience wouldn't be too great. My Meilisearch index is ~40GB but I don't know how much that could be compressed down.

Here's how the similar page for books is generated, which I forgot to mention on the "how it works" page: https://gist.github.com/chris124567/8d06d64bfe827cb7f6121f93...

costco··on Show HN: I scraped 3B Goodreads reviews to train a better recommendation model
It's recommended that you put at least 3 books in. If you would like recommendations just based on one book, click the similar button on the book, it should take you to this page: https://book.sv/similar?id=297914
costco··on Show HN: I scraped 3B Goodreads reviews to train a better recommendation model
You can import the first or last 64 books of your read, to-read, or currently-reading shelves if you press the "Import Goodreads" button and provide your Goodreads ID.
costco··on Show HN: I scraped 3B Goodreads reviews to train a better recommendation model
Thank you for the recommendations. I didn't try BERT4Rec because I assumed it would perform the same or worse as what I already had after having read https://dl.acm.org/doi/pdf/10.1145/3699521. The TIGER paper seems interesting - I definitely want to explore semantic IDs in general and also because I think it could allow including more long-tail items.
costco··on Show HN: I scraped 3B Goodreads reviews to train a better recommendation model
Worked for me, could be due to server being overwhelmed

Here is the URL with your books: https://book.sv/#52752877,46049530,18437030,52480873,3260654...

costco··on Show HN: I scraped 3B Goodreads reviews to train a better recommendation model
Yes I would say the handling of series is probably the biggest problem. Once my test metrics got to a point I was happy with and my quality spot checks passed (can I follow the models recommendations from one generic history book to Steven Runciman, also making sure popular books don't always dominate the results), I was ready to release because I had been working on this project for so long. The solution is probably using the transformer model to generate 100-200 candidates and then having a reranker on top.
costco··on Show HN: I scraped 3B Goodreads reviews to train a better recommendation model
I didn't want an agent to get stuck on an infinite loop invoking endpoints that cost GPU resources. Those fears are probably unfounded, so if people really cared I could remove those. /similar is blocked by default because I don't want 500000 "similar books for" pages to pollute the search results for my website but I do not mind if people scrape those pages.
costco··on Show HN: I scraped 3B Goodreads reviews to train a better recommendation model
If you want recommendations solely based on one book, please try the similar page: https://book.sv/similar?id=13566692

These seem to fit the description you are going for better. The model is trained to predict the next book in the sequence. Those other books you listed happen to be very popular, so in the absence of information about you (only having 1 book), the model will tend to recommend those.

costco··on Show HN: I scraped 3B Goodreads reviews to train a better recommendation model
Everything (namely Meilisearch, Postgres and the web server in Go) besides the model inference is running on a Hetzner server with a large SSD and an "AMD Ryzen 7 3700X 8-Core Processor." The data.ms directory is about 40GB. Once the HN traffic dies down I will probably move the model back to the Hetzner server so I don't have to pay $0.15/hour for an A4000.
costco··on Show HN: I scraped 3B Goodreads reviews to train a better recommendation model
Not sure if I can. At the very least book descriptions most likely could not be distributed. There is an academic dataset with around 200M reviews though: https://cseweb.ucsd.edu/~jmcauley/datasets/goodreads.html
costco··on Show HN: I scraped 3B Goodreads reviews to train a better recommendation model
I think I will expand the input books limit (sadly requires retraining) and or the output books limit of 30.
costco··on Show HN: I scraped 3B Goodreads reviews to train a better recommendation model
I'm not an expert by any means but as far as sequential recommendations go, aren't SASRec and its derivatives pretty much the name of the game? I probably should have looked into HSTUs more. Also this / sparse transformers in general: https://arxiv.org/pdf/2212.04120
costco··on Show HN: I scraped 3B Goodreads reviews to train a better recommendation model
It's explicitly trained to predict the next book read in a sequence, which is why you get that behavior. There's probably a better way for me to handle it rather than having 5 books from the same series tend towards the top though.
costco··on Show HN: I scraped 3B Goodreads reviews to train a better recommendation model
Thank you for the compliments :) I used 50-100 datacenter proxies. I just logged requests made by the iOS app with Charles and then recreated the headers to the best of my ability though the server did not seem to be very strict at all. Worth noting though that static residential proxies are not too expensive these days anyways.

Re the API: The model does actually run fairly well on CPU so it probably wouldn't be too expensive to serve. I guess if there is demand for it I could do it. I think most social book sites would probably like to own their recommendation system though.

costco··on 'SIM Farms' Are a Spam Plague. A Giant One in NY Threatened US Infrastructure
Finally someone in this thread who gets it. In fact many of the temporary phone number registration services used to make bot accounts have pages on their website where they say "Do you live in country X and have access to large amounts of SIM cards? Partner with us - we will send you the equipment." Calling the US through standard means is so cheap that I suspect the only reason anyone would use a grey market provider is if they were making spam calls, but maybe I'm wrong.
costco··on NYC Telecom Raid: What's Up with Those Weird SIM Banks?
I've read that at least for for GSM VoIP gateway setups, people typically rotate SIMs because it looks suspicious for one customer to be making calls all day. In fact there is a whole industry most people have no idea about dedicated to detecting these so-called grey routes. Having perused some of these offerings, it seems fairly typical to offer 16-32 radios and 256 SIM card slots. For residential proxy setups, I've seen a bunch of 16 port USB hubs with Huawei LTE modems.
costco··on How did .agakhan, .ismaili and .imamat get their own TLDs?
Related: It’s only about $5k a year to run your own registrar (which is different than a gTLD). You have to sign agreements with Verisign and others to register domains on TLDs that people actually want but it’s not unduly hard and I think they even offer good payment terms.
costco··on Show HN: Building a web search engine from scratch with 3B neural embeddings
This is awesome, and the low cost is especially impressive. I rarely have the motivation after working on a side project to actually document all the decisions made along the way, much less in such a thorough way. Regarding your CoreNN library, Clearview has a blog post [1] on how they index 30 billion face embeddings that you may find interesting. They combine RocksDB with faiss.

[1] https://www.clearview.ai/post/how-we-store-and-search-30-bil...

costco··on Tor: How a military project became a lifeline for privacy
If you already have an understanding of how Tor works and want to know how attacks on it work, read these!

- https://github.com/mikeperry-tor/vanguards/blob/master/READM...

- https://github.com/mikeperry-tor/vanguards/blob/master/READM...

- https://spec.torproject.org/proposals/344-protocol-info-leak...

costco··on Tor: How a military project became a lifeline for privacy
This page on the mailing list has links to cases of people who were caught because of an unknown flaw in Tor: https://archive.torproject.org/websites/lists.torproject.org...

I can't find a link, but I think people have done simulations and the privacy benefits of more hops are not as great as one might think. If you control the guard and exit, then traffic confirmation is relatively easy by just looking at timing and volume of traffic no matter how many hops are in between.

costco··on Tor: How a military project became a lifeline for privacy
Were you running specifically a bridge or just a non exit relay? Bridges are generally unlisted and are somewhat expensive to mass scrape (the bridge distributors will require captcha or email or Telegram etc) so they are less likely to show up in those lists. Whereas all relays are listed in the consensus and can be trivially enumerated.
costco··on Bloom Filters by Example
I had used bloom filters in the past without really understanding how they worked. Then one day I decided to implement them just going off the Wikipedia article with the 32-bit MurmurHash function and was surprised at how simple it was. If you're using C++ you can use std::vector<bool> (or as of C++23, std::bitset) to make it even easier to store the bits in a space efficient way.
costco··on Unauthorized experiment on r/changemyview involving AI-generated comments
Look at the accounts linked at the bottom of the post. They actually sound real like people whereas you can usually you can spot bots from a mile away.
costco··on I'm done with coding
Clicked on this post because the domain name sounded familiar from the Tor mailing list. I knew you ran a large set of relays but didn't know you were also a pretty extensive contributor over many years! You're definitely smart and if you were able to get a job at Microsoft you're capable of getting a job at most other places, so this doesn't really have to be a permanent decision if you don't want it to be. You can work at Let's Encrypt, Tor, Signal, etc while making an impact and still doing pretty well for yourself. Anyways, in the spirit of this forum, I wish you luck with your startup.
costco··on Ask HN: Opinion on efforts to find prior art on outrageous priced drugs
Going through the IPR process will cost half a million and a year of time at least. The filing fees alone are in the tens of thousands. It's unlikely anyone would pursue a case unless thought they had a good chance of winning. Kyle Bass's strategy was just to hope that filing the patent dispute would cause the share price of the manufacturer to decline, which it did not in many cases. He ultimately lost most of the time in court, so I don't know how much he actually profited.

Have you heard of Cloudflare's Project Jengo [0] [1]? They were sued by Sable, a patent troll. So they made a website where anyone could submit prior art for any of Sable's patents and they would pay I think around $1000-2000 if it helped their case. Imagine if you had a website where you listed drugs alongside their patents, and a bounty in dollars if you found prior art. The bounty could be funded by hedge funds, generics companies, or other competitors. If the submissions were solid enough, they would take the case to lawyers and hopefully win.

    KEYTRUDA® (pembrolizumab)
    - Bounty: $100,000 
    - Patents:
        - U.S. Patent No. 8,354,509
        - U.S. Patent No. 8,900,587
        - U.S. Patent No. 9,834,605
        - U.S. Patent No. 11,117,961
        - U.S. Patent No. 9,220,776

    ...
You would probably need a lot of connections to make this work. You would also basically create a side hustle for bored patent lawyers or people with a lot of time on their hands. Though the people who are really good at this sort of work probably already make a lot of money, so maybe it wouldn't work.

This is basically your original idea, but there's a monetary incentive. I don't think people with the level of expertise needed to do this would do it for free.

[0] https://www.cloudflare.com/jengo/sable-prior-art-search/

[1] https://blog.cloudflare.com/the-project-jengo-saga-how-cloud...

Edit: but maybe I'm wrong. The CEO of Cloudflare says most of the people who submitted probably would have done it for free: https://news.ycombinator.com/item?id=41732580. But then again, Cloudflare was able to publicize their cause easily among technical people who can understand software patents on places like HN, and there was a moral righteousness element to it because patent trolls are parasites. It might be difficult to inspire the same level of enthusiasm about orphan drugs, and there is also likely a smaller number of people who have the skill to review drug patents.

costco··on Ask HN: Opinion on efforts to find prior art on outrageous priced drugs
If you can successfully challenge pharma patents you can get rich by shorting the stock. A hedge fund manager named Kyle Bass tried this with a couple dozen drugs to varying degrees of success. But also, a patent being invalidated doesn't mean prices come down immediately. ANDAs still take 3-4 years to get approved on average.
costco··on Treasury agrees to block DOGE's access to personal taxpayer data at IRS
Yes, but it'd be nice if there were insights in the comments not present in the original text on these sorts of political articles. I have not seen that lately, and frankly "anything that good hackers would find interesting" has become so tortured as to become meaningless. People are missing the "If they'd cover it on TV news, it's probably off-topic" line that follows.

The real issue is that there are too many articles on the front page that everyone can participate in (news, 150 word anecdotes on AI, language/editor wars, ...). If there is too much pent up demand for those topics, it should just be moved to a certain day of the week. I think you could more or less violently suppress it while having very limited collateral damage on actual technical discussions. And by allowing politics to remain on the front page for many days, you basically slowly change the composition of the community to people who want to debate politics all the time, which is explicitly not the intent of HN. I'm probably violating the guidelines by complaining instead of silently flagging the article, but hoping this inspires other people to start flagging as well.

← PreviousPage 2 of 13Next →