High-performance .NET by example: Filtering bot traffic
alexandrnikitin.github.io
alexandrnikitin.github.io
Thanks for sharing!
Meaning, yes, the OP got impressive performance improvements but the code is also completely unreadable and utilises unsafe code sections which could expose you to security problems/memory leaks/memory corruption. Not to mention they've recreated and will need to maintain an in-house version of the Dictionary class.
Their first optimisation (from Enumerator to List and Any() to Count()) are something every codebase could use. Most of their other optimisations make the code a maintenance minefield.
Plus programmers are expensive. Hardware is cheap. Why spent time on harder code to write that's also harder to maintain in the medium to long term when instead you could just throw money at hardware and call it a day? Just food for thought, not really a criticism in and of itself.
PS - Please don't take this post too seriously. I am not really being critical, just playing devil's advocate. I actually enjoyed the linked article a lot.
It seems like most git guis are pretty commit-focused. I'm not sure of any way to do it in the command line (though it must be possible) but it would be nice to highlight a section of code in your IDE and have git give you a history of just those lines (or that function) as far back as you want to go.
I cannot believe that you are honestly saying a 2x increase to throughput in production is something you "shouldn't take seriously" because the code isn't as readable as it was before.
Programmers are expensive. Hardware is cheap. That doesn't justify completely throwing out the window any performance increasing changes just because a fresh college grad won't be able to understand what's going on within 10 minutes.
I've worked on large and old codebases for years. I've seen plenty of examples of where a "clever" programmer has optimised the heck out of a section of code, made it completely unmaintainable, and as a result forced a re-write (the resulting code, which was slower, was also easier to maintain and reason about).
"Production" is a meaningless rallying cry. Everything is production sooner or later. Not everything needs to be fast, although there are critical areas of a typical project that do. It is really a question of code quality relative to performance, unless performance itself is actually causing you problems. Typically these overeager performance fixes are done to code preemptively.
This is all fine for your pet and toy projects, go work on something enterprise grade. Big code means clean code. I'd take ten lines of clean easy to consume code over one "clever" line of genius.
The usual caveat of making code correct, readable and then fast in this order of course also applies :)
What it latency for a single request can't be improved with more cores?
What if your product is used by consumers who may have old hardware? or phones, or watches, or laptops and they want their battery to last?
What if you consistently practiced at making high performance code, maybe then it wouldn't seem "unreadable" to you any more?
What if Linq-like higher order functions weren't slow? https://github.com/jackmott/LinqFaster
What if slow software was common today because of modern attitudes, and I wasn't seeing any increase in stability or features to show for it?
Then re-design for map-reduce, and scale horizontally
> What it latency for a single request can't be improved with more cores?
Then look into pre-computing and caching
> What if your product is used by consumers who may have old hardware? or phones, or watches, or laptops and they want their battery to last?
I thought we were talking about server requests? If we are, then offload this work to the server
> What if you consistently practiced at making high performance code, maybe then it wouldn't seem "unreadable" to you any more?
But it's not all about you. Unless you're working on a pet project, or you have the credibility and reputation to be the final call on a significant open source project, you might get hit by a bus tomorrow. Or, if you do a really good job, your company will need you to be a force multiplier to teach a dozen others to try to imitate you. Even if you're a "10x" programmer.
And by the way, when you optimize code THIS much, any refactoring or tweaks to new features cause your optimizations to get tossed out, and you have to start over from scratch.
> What if Linq-like higher order functions weren't slow? https://github.com/jackmott/LinqFaster
https://github.com/jackmott/LinqFaster#limitations
> What if slow software was common today because of modern attitudes, and I wasn't seeing any increase in stability or features to show for it?
Except you are, and you don't even realize it. Optimizations like this blog post matter a LOT on client software. Be it apps, or websites - anything run on the client will need this kind of attention sometimes.
But this guy is writing server software. Micro-optimizing on the server side the way he is doing is silly.
Khm... What if you have millions of requests per second with tight latency requirements measured in milliseconds; and a bunch of business logic to fit into that. Such optimizations aren't so silly. There are different scenarios on both client and server sides.
This is entirely a matter of scale. 1 dev can make faster software that runs on thousands of machines with far more ROI than the hours they spent. This is common in several industries.
The proper way to do this is to block by IP, based on behavior. Block IPs slowing down the site or throw up a captcha like cloudflare does.
Blocking bots sounds great but it just brings Google one step closer to a monopoly. Even good bots just pretend to be people nowadays because lots of people are implementing naive site protection strategies.
Edited: to be less mean
I agree with your sentiment, but you should try to be a little more constructive.
If I was writing a bot, I would set user agent to some well known and very popular value, i.e. newest Chrome on Windows, or something like that.
Disclaimer: Getting rid of those idiots^wmisguides poor souls is part of my job description.
I don't care about this one person who knows to change the user agent sent, those are the ones you can usually talk to and they'll happily throttle their crawlers.
The other 99 are the problem - and they are a problem that can be solved pretty well by simple string matching.
>We won’t cover black bots because it is a huge topic with sophisticated analysis and Machine learning algorithms. We will focus on the white and grey bots that identify themselves as such.
This is about not wasting time & effort showing advertising banners to good bots.
Use the robots.txt to ban the pulling of specific pages. Bots 99% of the time ignore robots, so if they pull it: block
Check how quickly pages are pulled. If passes a threshold: block
I've made a gist[0]; feel free to get in touch via GH if you'd like to discuss it further.
[0] - https://gist.github.com/marklr/ae0c2f1eb61855d13cde6cef6bf63...
Also you may find this talk useful [3] (Particularly slide 11).
Great write up by the way. Really thorough on the benchmarking!
[1] https://lucene.apache.org/core/4_1_0/core/org/apache/lucene/...
[2] http://www.openfst.org/twiki/bin/view/FST/WebHome
[3] https://www.slideshare.net/lucenerevolution/text-tagging-wit...
A lot of manual work with various perf tools.
What's a bit missing is some production performance monitoring (APM) that gives you such data, with no manual interaction.
Technical : if the UserAgent claim to be a regular browser (let say Chrome 43) we will check on network level if the client implement http protocol like Chrome 43 usually do and on the JS side if the Javascript render is correct for Chrome. In case it's a real Chrome, we will check if the Browser is controlled by automation Tool.
Behavior : we will check if the path of requests is regular according to the website usage.
Disclaimer: I'm working at https://datadome.co, a bot protection tool.