Defending a website with Zip bombs
blog.haschek.at
blog.haschek.at
Also worth trying is an XML bomb [1], though that's higher up the stack.
Of course you can combine all three in one payload (since it's more likely that lower levels of the stack implement streaming processing): gzip an XML bomb followed by a gigabyte of space characters, then gzip that followed by a gigabyte of NULs, then serve it up as application/xml with both Content-Encoding and Transfer-Encoding: gzip.
(Actually now that I think of it, even though a terabyte of NULs compresses to 1 GiB [2], I bet that file is itself highly compressible, or could be made to be if it's handcrafted. You could probably serve that up easily with a few MiB file using the above technique.)
EDIT: In fact a 100 GiB version of such a payload compresses down do ~160 KiB on the wire. (No, I won't be sharing it as I'm pretty sure that such reverse-hacking is legally not much different than serving up malware, especially since black-hat crawlers are more likely than not running on compromised devices.)
[1] https://en.wikipedia.org/wiki/Billion_laughs
[2] https://superuser.com/questions/139253/what-is-the-maximum-c...
http://www.aerasec.de/security/advisories/decompression-bomb... has a triple-gziped 100GB file down to 6k, the double-gzipped version is 230k.
I'm trying on 1TB, but it turns out to take some time.
So you don't need to compress the entire terabyte.
The compression finally finished after 3h (on an old MBP), "dd if=/dev/zero bs=1m count=1m | gzip | gzip | gzip" yields a bit under 10k (10082 bytes), and adding a 4th gzip yields a bit under 4k (4004 bytes). The 5th gzip starts increasing the size of the archive.
It still fit on a floppy disk. :)
Sure. But gzip bombs don't do substantial (if any) permanent damage. At most, they'd crash the system. And indeed, that might attract the attention of owners, who might then discover that their devices had been compromised.
Unless you're serving up XML which is itself zipped, such as an Open Office document. But most clients won't be looking for that.
And btw -- when you end up accidentally crashing/DoS:ing your corporate WAF or ISPs DPI, who are they going to call?
I'm certain there's something (unintentionally) perverse in the ranking system, but as it's secret it's impossible to say for sure.
Secret rules that manipulate our behaviour are damaging.
After submitting something to HN I like to watch the HTTP logs, I get a lot of visitors from bots, but it's actually only ca 10-20 real people that actually read your blog. I don't know eneough of statisitcs to explain it well, but as 20 people is so small amount of the total HN readers, it's basically luck. And the representation of those who reads the "new" section might be a bit skewed from those who only reads the front page. If you want to help HN get better with more interesting content, you can help by actually visiting the "new" section.
The Economist once changed their default comments view from "most recommended" to "newest". Suddenly the advantage of being the first to post a moderately good comment vanished. Design matters.
1 http://represent.berkeley.edu/umati/
Edit: tl;dr: Future bot: "I have a submission. It has a topic, a shape, and other metrics. It's from a submitter, with a history. Perhaps it has comments, also with metrics, from people also with histories. I have people available, currently reading HN, who all have histories. That's a lot of data - I can do statistics. Who might best reduce my optimization function uncertainty? I choose consults, draw the submission to their attention, and ask them questions. I iterate and converge." Versus drooling bot: "Uhh, a down vote click. Might be an expert, might be eternal September... duh, don't know, don't care. points--. Duh, done."
Context: The parent observed that with a small number of people up/down voting, the result was noisy. I observed the numbers were sufficient, if the system used more of the information available to it. And that the failure to do so is a long-standing problem.
Details, in reverse order: "Civilization": Does anyone not think the state of decision and discussion support tech is a critical bottleneck in engineering, business, or governance? "AR": A principle difficulty is always integrating support tech with existing process. AR opens this up greatly. At a minimum, think slack bots for in-person conversations. "crowdsourcing": or human computation, or social computing, is creating hybrid human-computer processes, where the computer system better understands the domain, the humans involved, and better utilizes the humans, than does a traditional systems. "ML": a common way to better understand a domain. As ML, human computation, textual analysis, etc, all mature, the cost and barrier to utilizing them in creating better discussion systems declines. "Usenet": Usenet discussion support tooling plateaued years before it declined. Years of having a problem, and it not being addressed. Of "did person X ever finishing their research/implementation of tool Y". "decades": that was mid-1990's, two decades ago. "little changes": Discussion support systems remains an "active" area of research - for a very low-activity and low-quality value of "active". I'm unclear on what else could be controversial here.
For anyone who hasn't read the paper, it's fun - it was a best paper at SIGCHI that year, and the web page has a video. A key idea is that redundant use of a small number of less-skilled humans (undergraduates grading exam questions), can if intelligently combined, give performance comparable to an expert human (graduate student grader). Similar results have since been shown in other domains, such as combining "is this relevant research?" judgments from cheaper less-specialized doctors with more-expensive specialized ones. On HN, it's not possible to have a fulltime staff of highly skilled editors. But it is technically plausible to use the exiting human participants to achieve a similar effect. That we, as a field, are not trying very hard, reflects on incentives.
1. Stop posting to HN and let others do it for you
2. Spend more effort in writing blog and less on HN karma
3. ... (monetize)
4. Profit!!! (+ more blog posts about how you monetized and instantly gain more HN karma than posting will ever do)
</saracasm>=) upvoted for original content.
I presume the ones that gave out sooner were manually stopped by whoever maintains them or they hit some sort of memory limit. Good times.
Keep them on line by being a very dumb customer until they start cursing and hang up on me. : - )
EDIT: Nope, apparently the person who runs it takes it down at night for some reason. Maybe to minimize people using it as a prank call?
When working from home this is one of the few joys/social interactions if the day :D
pretty novel to see it used the other way around though!
And the follow up: http://www.hackerfactor.com/blog/index.php?/archives/763-The...
[0]:https://en.wikipedia.org/wiki/Intrusion_Countermeasures_Elec...
Because it seems in the realm of possibility that if a large botnet hits you and your responses crash a bunch of computers you could do serious time for trying it. I'm hoping there's precedent against this...
Connecting to a server...( A lot)
Putting random strings into forms...( A lot)
Moving your money between banks... (In different countries)
Buying stocks... (With insider knowledge)
A simple act doesn't spell the whole story, and fraud, computer crime, etc laws are written vaguely enough for a country to prosecute someone " sending large files."
Microsoft doesn't take the fall for malware, even if its a fault in SMB or the like.
The intent is damage.
He bases this attack on IP addresses. IPv4 addresses are regularly shared between consumers. He's tossing a knife into a crowd because he thought he saw someone.
> you could make a pretty good fleeing felon argument.
In a nation that allows you to attack, not just restrain, a fleeing felon.
But his attack may hit a nation that doesn't allow that.
You ask for something a vulnerability scanner would ask for? You get a gzip bomb.
[0] https://www.reddit.com/r/PHP/comments/6lfl6p/i_have_created_...
Let's be frank.
He's serving up malware to potential users who hit too many 404s.
> Awesome! My production implementation of the bomb also looks at 404's and 403's per IP and if there are too many of those it will send the bomb. [0]
This could be exploited by a third party, which makes him complicit.
He targets IP addresses, and as the IPv4 world often shares those, he can attack innocent bystanders who happen to be in the same allocation as a miscreant.
Finally, self-defence is established as denial or dropped connections. As he's intentionally avoided established practice, and developed an attack instead, it becomes undue harm.
Let alone if he attacks someone in a nation that has an extradition treaty, but no concept of this sort of "fighting back".
[0] https://www.reddit.com/r/PHP/comments/6lfl6p/i_have_created_...
A farmer here in UK stirred up a whole load of shit when he shot two burglars [1] trying to escape from his property.
Some places in the USA have "stand your ground" laws. These say you aren't required to retreat, that you can "stand your ground", that you can use (legally) leathal force without requiring that your back is against the wall.
As for people running away, the only way I see self defence working is when they still pose an 'imminent threat to life' which seems rather hard to argue.
I've read but couldn't find again the story of someone shooting a tief to get back his VHS player and walk free.
I'm not arguing for actually using the law to shoot people: I don't ever want to be in that situation myself, but I'm saying depending on the situation you do in fact have the law on your side.
That isn't normal, though. It's likely that you were already feuding, and so the law will look askance at you for not bringing authorities into it much earlier.
I think all of those cases are covered by any imminent threat clause, and thus do not need special exemptions. Just like there isn't an exemption that you are not allowed to shoot a retreating person. It simply follows because (with exceptions) retreating people aren't imminent threats.
Florida [1], for example, says:
> ... A person who uses or threatens to use deadly force in accordance with this subsection does not have a duty to retreat and has the right to stand his or her ground if the person using or threatening to use the deadly force is not engaged in a criminal activity and is in a place where he or she has a right to be.
In section 0776.013, the castle doctrine is also noted, but is more expansive, and includes the use of deadly force even if there is no threat of imminent harm.
[1] http://www.leg.state.fl.us/statutes/index.cfm?App_mode=Displ...
The US tends to be a little more prescriptive, leaving a situation where different jurisdictions have more specific requirements for defining what constitutes self-defense.
Juries in the UK tend to have significantly more responsibility for making judgments like these, leading to a system where evolving views of what is right and wrong can result in standards naturally evolving over time, rather than being fixed by what people thought was okay thirty years ago.
Nobody malicious brings down crawlers. It's just unexpected things you find out on the internet.
You're wrong about that. I've more than once brought down crawlers on purpose, especially the ones that didn't respect robots.txt.
I don't see a way to comment on the article itself, but hopefully the author reads this.
'starts_with' is descriptive and language agnostic where '=== 0' is neither.
Might also be used for some kind of reflection attack. Want to kill some service that let's users provide a url (for an avatar image or something) - point it to your zip bomber.
Actually, I don't see how to defend this. Is there any way to ask a gzip file which size it will be once unzipped, without needing to decompress it?
The closest is uncompressing it and counting and immediately discarding the bytes in the output stream.
But of course the proper defense is to give up if you exceed a predefined memory or time budget.
https://www.blockedservers.com/
It's a lot more effective to kill the connection rather than to start sending data if you're faced with a large number of attempts.
I guess if you see them try to use the USB killer, you'd be obligated to report it. Otherwise I don't think its an issue.
If there’s any question about culpability, all they have to do is ask you, “Is there anything in your baggage you think we should know about?” and if you don’t disclose it then, you’re screwed.
What? You say you were never asked such a question? You say you even tried to warn them? Well I have sworn testimony from a TSA officer that says they ALWAYS ask that question, and you’re the guy who was caught carrying a piece of equipment designed for trickery and vandalism. Case dismissed.
The payload could be encrypted text of two chat bots talking jibberish.
Edit: Or even more useful, bbcp. Which is the best file transfer app that I've ever used.
There was a HN story[2] on Chaffinch[2], which is where I came across teh idea.
[1] https://en.wikipedia.org/wiki/Chaffing_and_winnowing [2] https://news.ycombinator.com/item?id=14408757 [3] https://www.cl.cam.ac.uk/~rnc1/Chaffinch.html#Chaffing
Tarpit Action for Fail2ban with rate limit
Permission denied. Please try again.You could, though, write a pam module to trickle data out very slowly. Maybe pam_python would be easier to experiment with.
I use pam_shield to just null route ssh connections with X failed login attempts. There's no retaliation in that approach, but it does stop the brute forcing.
If a large number of hosts treats some behaviour as deserving a slow-service attack, then clients exhibiting that behaviour are faced with a large set of slow-serving servers.
Any given server can monitor how many slow-service attacks it is currently providing. Given that a criterion for an SSA is having already determined that the connection is not a friendly one, then monitoring useful vs. useless (e.g., SSA) connections, and being prepared to terminate (or better: simply abandon) the SSA connections as normal traffic ramps up, is a net benefit.
Meantime, the hostile clients are faced by a pervasive wall of mud, slowing their access.
This is why I want to point out that simply serving content with delay and increasing number of active connections create additional attack vector that more dangerous than script kiddies scanner was in first place.
What if the compressed file is plausibly valid content? How could intent be malicious if a request is served with actual content?
It probably shouldn't be, but law is funny that way.
Intentionally sending a zip bomb could potentially get you in trouble as well. Especially if you're just one private person or a small company without a legal division to brush it off.
There isn't a real black/white interpretation though, at least not outside the US (where there may be history to influence ruling on the subject), and obviously most victims wouldn't report you, but more often than not you wouldn't want to test interpretation of IT related law.
I wonder if there's room to do this with other protocols? Ultimately we want to crash whatever tool the scriptkiddy uses.
I replaced it with a GZIP bomb. It was very satisfying to watch the requests start slowing down, and eventually stop.
That also crossed with another thought about pre-compressing (real!) content so that Apache can serve it gzipped entirely statically with sendfile() rather than using mod_deflate on the fly, so unless I've misunderstood I think that bot defences can be served entirely statically to minimise CPU demand. I don't mind a non-checked-in gzip -v9 file of a few MB sitting there waiting...
I tried visiting the payload site with Tails OS (a Linux distro for privacy minded) and the whole OS is frozen.
With some horrible WebGL code I've crashed the macOS compositor before.
> As 42.zip shows us it can compress a 4.5 peta byte (4.500.000 giga bytes) file down to 42 kilo bytes.
It wouldn't save a scanner from crashing to use a time-out or max read bytes. The defense can send the 100kb zipped data in a matter of seconds. The client then decompresses the zipped data which expands to gigabytes, causing crashes by out-of-memory.
Well actually from memory the author of the blog was doubtful if this exploit actually crashed Eddie or not, but it did crash the other bots (Eddie V1 did go offline, possibly as a crash), so it would appear you are correct. Only truely naive bots might well be affected by this.
I'm actually surprised that many other scanners failed to do this.