How F5Bot Slurps All of Reddit
intoli.com
intoli.com
You can save a lot of bandwith by requesting compressed responses:
$ curl -s --user-agent moo/1 -H 'Accept-Encoding: gzip' "$pretty_long_url" > test.gz
$ wc -c < test.gz
63507
$ gzip -d < test.gz | wc -c
426941
(OK, that's 85% saved, not 95%, but hey.)If you do it right you can even keep the content stored compressed without re-compressing by saving the compressed byte stream directly.
https://github.com/madler/pigz/issues/36#issuecomment-249041...
Decompression can’t be parallelized, at least not without specially prepared deflate streams for that purpose. As a result, pigz uses a single thread (the main thread) for decompression, but will create three other threads for reading, writing, and check calculation, which can speed up decompression under some circumstances.
This is (unfortunately) not quite true. Since Reddit introduced "profile posts," there can be a post where the subreddit name is something like "u_Shitty_Watercolour" but the subreddit_name_prefixed is actually "u/Shitty_Watercolour", rather than "r/u_Shitty_Watercolour".
Example: https://www.reddit.com/user/Shitty_Watercolour/comments/84nh...
Maybe one is just an alias though? I wonder if you can make a r/u_$unused_username and then later register $unused_username
edit: nope, you can't make a sub that starts with "u_"
However the point of subreddit_name_prefixed (I assume) is to display something in a user-facing way. For this purpose, r/u_something is correct but not proper.
https://www.reddit.com/r/changelog/comments/7tus5f/update_to...
https://www.reddit.com/r/redditdev/comments/7qpn0h/how_to_re...
https://www.reddit.com/r/help/comments/1u0scj/get_full_post_...
https://www.reddit.com/r/help/comments/6nxqjm/maximum_of_100...
I posted some more information about it a while ago here: https://www.reddit.com/r/help/comments/7en0uu/my_saved_posts...
Imagine I have a list of max length 5, when I initially fill it up it looks like [5, 4, 3, 2, 1]. If I save one more thing, it adds 6 at the front, then truncates the list and removes the 1 from the end, so now you have [6, 5, 4, 3, 2]. At that point, if I unsave #3, it just removes it from the list so you'd have [6, 5, 4, 2]. #1 is still saved, but nothing happens to pull it back into the list.
1000 probably seemed like a lot at the time or had reasonable performance, and it's just never been changed.
I pulled the dumps from Bigquery and am going to load them into PG at some point here, so I can do arbitrary queries without hitting BigQuery every time. Haven't looked at how to do that in realtime though, retrospective queries are mostly what I'm interested in.
If you don't want to do that, there's also https://redditsearch.io/ - if you want to search back farther, be sure to set the toggle to "all" instead of "day"!
Thankfully, services like pushshift[1] exist, which has a sane API and the option to use plain elasticsearch.
"see all of my own comments": http://api.pushshift.io/reddit/comment/search?author=saurik&... (use &before=[epoch] for pagination)
"see all of the posts made to the subreddit I moderate": https://api.pushshift.io/reddit/search/submission/?subreddit...
I have a subreddit (/r/pushshift) that gives examples on how to use the API. I'm always happy to answer questions about the API and take suggestions to improve it.
I'm very very close to launching the new API which will have a lot more features and a better design than the current one.
Edit: I just realized my username on this site got truncated to uhhh ... wow.
It's not really related to search, most of the cause is a pretty bizarre optimization method that reddit decided to implement fairly early on. The database structure is unusual and not very conducive to indexing (it's similar to an EAV model), so at some point they decided to basically write their own "secondary indexing"-like system that stores the "listing indexes" in memcached/Cassandra (with the data itself in PostgreSQL).
Whenever something happens that affects any listings (new post created, voting, etc), the site figures out all the listings it needs to update, and where in each listing the affected post now belongs and updates them all. So for example, if you make a new submission to /r/pics, it will go through and add the post's ID in the right spot to the "new posts in /r/pics" listing, the "hot posts in /r/pics" listing, the "new posts by yourusername" listing, and so on. As it's going through and updating all these lists, it also trims each one down to 1000 items.
It's conceptually pretty similar to a normal database indexing system, but basically maintains all the indexes "manually" and restricts them all to the top 1000 items.
If you're curious enough to dig around in the code, this is probably the main relevant file: https://github.com/reddit-archive/reddit/blob/master/r2/r2/l...
Why would we think php is slow? PHP is blazing fast certain applications (looking at you sugarcrm) make this into a mockery by rewriting queries and loading unnecessary data into each page request.
Nice to see a php related show and tell.
It's still quite slow compared to C/C++/Rust/Go, more than 10x slower:
https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
That same benchmark shows JS code that’s 6x faster than PHP.
PHP was popular when it had MySQL bindings that were always up to date with MySQL changes and apache PHP setup was simplistic. Thus the LAMP stack, right? It has always been rather fast, in the larger ecosystem of web scripting languages (especially if you avoided particular BIFs) but it got a LOT faster in 7 as the slowest bits were optimized alongside everything else.
Example query which searches for 'f5bot' in the past day and correctly finds the corresponding posts on Reddit:
#standardSQL
SELECT title, subreddit, permalink
FROM `pushshift.rt_reddit.submissions`
WHERE created_utc > TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 1 DAY)
AND REGEXP_CONTAINS(LOWER(title), r'f5bot')> Basically I use the selftext, subreddit, permalink, url and title. The other 95% of it is just wasted bandwidth.
It’d probably be better for Reddit if they allowed for specifying the fields we care about rather than just returning the whole thing…
I mean, I enjoy the idea of human readability as much as Jon Postel, but at certain scales you have to wonder about the hidden cost of petabytes of human-readable data flying over the wire, never to be seen by anything but computers.
(Client must specify compression support)
If I'm right, that would probably speed things up significantly.
Sigh, just use Perl. Writing code with the general regex engine took me only one minute of effort, but it runs already nearly 500× faster than codeplea's optimised special purpose code.
Why 100000 loops and not 10 like in the original code? Otherwise Benchmark.pm will show "(warning: too few iterations for a reliable count)".
----
benchmark.php 100000 loops:
Loaded 3000 keywords to search on a text of 19377 characters.
Searching with aho corasick...
time: 329.3522541523
----benchmark.pl 100000 loops:
Benchmark: timing 100000 iterations of regex...
regex: 0.691561 wallclock secs ( 0.69 usr + 0.00 sys = 0.69 CPU) @ 144927.54/s (n=100000)
----benchmark.pl (fill in the abbreviated ... parts from benchmark_setup.php):
#!/usr/bin/env perl
use Benchmark qw(timethese :hireswallclock);
require Time::HiRes;
my @needles = qw(
abandonment abashed abashments abduction ...
);
my $haystack = 'unscathed grampus ...
heroically';
my $n = join '|', @needles;
timethese 100000, {
regex => sub {
my @found;
while ($haystack =~ /($n)/cg) {
push @found, [$1, pos $haystack];
}
return @found;
},
index => sub {
my @found;
for (@needles) {
my $pos = index $haystack, $_;
push @found, [$_, $pos] if -1 < $pos;
}
return @found;
}
};I get what you're saying, but it's not quite as easy as you imply.
Pulling in an entire programming language is a much bigger dependency and maintenance cost than spending a couple hours writing an algorithm. It would make more sense to just use a C extension.
I did try PHP's regex. It was much, much slower.
I learnt something valuable, thank you for that.
Seems a bit over the top imho. Maybe a better approach is to ask for a 1,000 and look for any missing — which you can grab individually.
I’d be a little annoyed at people not using batch mode and making so many request but that’s just me.
Their default listing mode works very poorly. It would certainly be more requests to use a hybrid system like you're talking about.
https://bigquery.cloud.google.com/table/fh-bigquery:reddit_c...
It's pretty easy to get a firehose.
100 sounds like a typical "max-requests" pipelining limit.
He does not mention CURLMOPT_PIPELINING.
Does this mean he makes 100 TCP connections in order to make 100 HTTP requests?
With the "&limit" parameter he can change how many items he receives per HTTP request. This has nothing to do with a limit on how many HTTP requests he can make per TCP connection (pipelining). Maybe that is the "100" he is complaining about, i.e., 100 items per HTTP request.
However you failed to answer my question: Is he making 100 TCP connections to make 100 HTTP requests?
Does the Reddit server set a limit on how many HTTP requests he can make per connection? (100 is a common limit for web servers)
Sometimes the server admins may set a limit of 1 HTTP request per TCP connection. This prevents users from pipelining outside the browser, e.g., with libcurl or some other method.
She/He's not. One request can supply 100 IDs.
> So here are 2,000 posts, spread out over 20 batches of 100 that we download simultaneously. It assumes you’ve already got the last post ID loaded into $id base-36.
> Here’s how we do it. We find the starting post ID, and then we get posts individually from https://api.reddit.com/api/info.json?id=. We’ll need to add a big list of post IDs to the end of that API URL.
The votes aren't real and they don't matter. This is true for FB likes or whatever else you can imagine. Reddit goes to a LOT of trouble to counter bad actors, but if you're in some subreddit that you think is full of abuse, then it's like any other part of the internet. Go somewhere else. This has little to do with Reddit as a whole, so the proclamation seems unnecessarily volatile.