> Basically I use the selftext, subreddit, permalink, url and title. The other 95% of it is just wasted bandwidth.
It’d probably be better for Reddit if they allowed for specifying the fields we care about rather than just returning the whole thing…
> Basically I use the selftext, subreddit, permalink, url and title. The other 95% of it is just wasted bandwidth.
It’d probably be better for Reddit if they allowed for specifying the fields we care about rather than just returning the whole thing…
I mean, I enjoy the idea of human readability as much as Jon Postel, but at certain scales you have to wonder about the hidden cost of petabytes of human-readable data flying over the wire, never to be seen by anything but computers.
(Client must specify compression support)
Sigh, just use Perl. Writing code with the general regex engine took me only one minute of effort, but it runs already nearly 500× faster than codeplea's optimised special purpose code.
Why 100000 loops and not 10 like in the original code? Otherwise Benchmark.pm will show "(warning: too few iterations for a reliable count)".
----
benchmark.php 100000 loops:
Loaded 3000 keywords to search on a text of 19377 characters.
Searching with aho corasick...
time: 329.3522541523
----benchmark.pl 100000 loops:
Benchmark: timing 100000 iterations of regex...
regex: 0.691561 wallclock secs ( 0.69 usr + 0.00 sys = 0.69 CPU) @ 144927.54/s (n=100000)
----benchmark.pl (fill in the abbreviated ... parts from benchmark_setup.php):
#!/usr/bin/env perl
use Benchmark qw(timethese :hireswallclock);
require Time::HiRes;
my @needles = qw(
abandonment abashed abashments abduction ...
);
my $haystack = 'unscathed grampus ...
heroically';
my $n = join '|', @needles;
timethese 100000, {
regex => sub {
my @found;
while ($haystack =~ /($n)/cg) {
push @found, [$1, pos $haystack];
}
return @found;
},
index => sub {
my @found;
for (@needles) {
my $pos = index $haystack, $_;
push @found, [$_, $pos] if -1 < $pos;
}
return @found;
}
};I get what you're saying, but it's not quite as easy as you imply.
Pulling in an entire programming language is a much bigger dependency and maintenance cost than spending a couple hours writing an algorithm. It would make more sense to just use a C extension.
I did try PHP's regex. It was much, much slower.
I learnt something valuable, thank you for that.
If I'm right, that would probably speed things up significantly.