Not sure what you mean by scraping in parallel.
Not sure what you mean by scraping in parallel.
Maybe I'm just an old fart, and blowing 10GB of data is cheap nowadays?
I'm surprised some users think it's a lot on an individual scale. Consider an ideal persona for this submission: query some movies, then go stream several GBs for one. The 100 MB payload isn't much in comparison. I admit it is kind of bad form for mobile users who might be on metered data plans, and a warning and trigger for manual action for those devices would be kinder.
One thing that might help reset your notion of what a lot of bandwidth is would be to browse around with your network tab in the developer console open for a day. nytimes home page is 14 MB and they get a ton of traffic. Even a corporate blog on the HN front page right now that could just be the tiny compressed text is 2 MB. Single image loads on many pages can be 1 MB or more. Glancing at the submission, the response headers indicate it seems they're serving from S3 Cloudfront, which is free for the first TB per month, though after that it gets back to absurd pricing. AWS is not price competitive.
That’s an irresponsible waste of both bandwidth and storage. You should have really made that clear in the website.
Exclude movies with very low number of rating or potentially very low scores too.
The long tail reduction would be significant
It is on hosted on http://www.imdb-sql.com/imdb01-11-2024.parquet
In fact the main reason this project exists is
a) I wrote a jupyter notebook ages ago that'd join the raw data into a queryable form.
b) I committed it and forgot about it after the initial viewing of top movies/series.
c) For the Halloween I wanted to find a good-rated horror movie, a genre I don't watch much.
d) I found my notebook but it was a drag to get it running again, first pandas would keep throwing OOM errors so I had to migrate to polars. Secondly I had to find a spare working laptop since iPad is my primary off-work computing device as of late. Finally the schema is not so intuitive - I need to look up things like region/language codes.
I want something I can just leave unattended but still be able to use a few months later without going through a dev-env set-up. A static website with all the data is one potential way of avoiding maintaining it as it should keep running, just like an exe would keep running on Windows in the past.
I avoid maintaining personal projects as my software job requires me to maintain software and I don't find it enjoyable to do the same in my spare time.