HNHacker News
TopNewBestAskShowJobs

gighi

288 karma · joined January 17, 2014

Gianluca Borello

Interested in distributed systems, system programming, performance and scalability.

https://github.com/gianlucaborello

https://www.linkedin.com/in/gborello/

submissionscomments
gighi··on Container Isolation Gone Wrong
> 1. If one of the two containers caused the issue, then the why you needed both of the containers to produce the issue? Why running just the offending one was not enough?

Running just the offending one would have been clearly enough, since its effects would have caused the same increased latency for every other process in the system (including itself). However, using a second container to observe the performance degradation proves the point that one container is able to affect another one, which is sort of the gist of the article, since too many people think containers provide much more isolation than what in reality happens.

> My guess is that "worker" container requested those non-existent files from a volume mounted by the other container, is it right?

No, the containers didn't share any volume, the dentry cache is effectively a singleton within the kernel, so even if the set of volumes is not overlapping, all processes in the system will see a performance degradation, regardless of where the files being accessed reside.

> 2. Kernel hash table implementation. The whole point of hash table is that it's size is O(N), where N is the number of elements it holds.

Your speculation is correct, however, there are sound reasons for doing such a thing in the kernel (and not allowing the main array of the hash table dynamically expand/shrink), so I wouldn't consider it a bug per se. I'll refer you to this excellent comment: https://news.ycombinator.com/item?id=14660954

gighi··on Container Isolation Gone Wrong
OP here

That's correct, I should have included a chart explicitly measuring the I/O activity done by the two containers, but I can assure you there was literally no I/O activity, a dozen open files per second is a very negligible throughput. The bottleneck was solely in the cache.

gighi··on Container Isolation Gone Wrong
Updated, thanks! I am working with ElasticSearch (ES) more than EL these days and my muscle memory tricked me ;)
gighi··on Container Isolation Gone Wrong
OP here

The solution is very simple: as mentioned in the article, just use a newer kernel and always set memory limits for containers, the blog post is based on an older kernel (2.6.32) that quite a few people irresponsibly still use in containerized environments, mostly because EL6 is so popular among enterprises.

In newer kernels, allocations from object pools are now tied to the limits of the memory cgroups that requested them in userspace, if any, so you wouldn't incur in this specific issue and you would just effectively have a container not being able to use more than X MB of dcache entries (although there are probably other minor ones, for example related to sharing global kernel mutexes and such).

gighi··on How we found a bug in Amazon ELB
Definitely a feasible approach. Let's just say that the reality has a bit more color and we have some other practical advantages in controlling the exact moment when we disconnect a particular client :)
gighi··on How we found a bug in Amazon ELB
Yes, every commit gets pulled by jenkins which builds the whole thing, runs unit tests and then starts the deployment once the tests pass.
gighi··on How we found a bug in Amazon ELB
Both approaches are reasonable (and there's also a third one, ship your application in containers and replace containers instead of instances).

We update existing instances because in our test environment we deploy at every single new commit (we absolutely love that), and we have hundreds (or more) a day. At that pace, replacing instances would be more time consuming (again, for our specific use case) and less cost efficient.

Plus, updating existing instances is handled automatically by AWS Code Deploy, which provides a very good deploying pipeline that you can control using the aws cli tool.

There are other minor advantages but those are the two main ones.

gighi··on How we found a bug in Amazon ELB
We definitely observed such drops that we attributed to presumably internal ELB scaling activity, but they happen so occasionally that for the moment they haven't been a real issue, as opposed to this one described in the article which happened consistently at every deployment in our test environment.
gighi··on How we found a bug in Amazon ELB
To be fair, the scenario is less common due to the fact that it happens just when the drained connections are terminated in a certain pattern (as shown in the charts). Still definitely common enough that can be easily replicated and cause real troubles :)
gighi··on Easy, realtime, system-wide Shellshock monitoring
The LD_PRELOAD workaround is meant to "fix" the shellshock bug, sysdig doesn't do that, it just passively monitors every injection attempt up to the point where the injected function is actually read by a newly spawn bash.

With the LD_PRELOAD trick you can definitely secure your system, but you won't be able to see if there's a new service that is being used as an attack vector (for your own curiosity). With sysdig, you can, and if you capture a trace file you can also follow the exact process chain that caused the propagation of the environment variable.

gighi··on Easy, realtime, system-wide Shellshock monitoring
Yes, we most definitely can publish a private brew tap. I'm no expert as I mainly use Linux, but my understanding is that I'd need to create a brand new draios/homebrew-sysdig repo. I'll try to find the time to look into this and update the documentation. If, instead, just a PR to our main repo would suffice, feel free to send it over and we'll merge it in no time :)
gighi··on Easy, realtime, system-wide Shellshock monitoring
We just released 0.1.89 (special release to include the shellshock chisel) a few hours ago, so distribution maintainers aren't that fast: https://github.com/draios/sysdig/releases

Debian is currently at 0.1.88: https://packages.debian.org/sid/sysdig

And Ubuntu periodically merges all the unstable packages from Debian, so that's why they're lagging one version behind at this moment.

gighi··on Easy, realtime, system-wide Shellshock monitoring
Yeah, bummer, I can submit a PR to Homebrew but it would take a few hours/days, and we don't ship OSX binaries from Draios, so why don't you go with:

https://github.com/draios/sysdig/wiki/How-to-Install-Sysdig-...

Assuming you have a C/C++ compiler installed (comes via XCode) it really takes like 2 minutes.

Lazy alternative, in a couple days maximum Homebrew should be updated, unfortunately it doesn't depend on us.

Also, notice that sysdig for OSX doesn't (yet) have live capture, so you'll just be able to run the chisel on a trace file that you previously created on a Linux host.

gighi··on Easy, realtime, system-wide Shellshock monitoring
Fair point, even though:

- At this point sysdig is estimated to have tens of thousands of users, and we haven't gotten a kernel bug in a while, with people (us included) regularly using it a lot in production. Of course, I see the irony of mentioning this in a "shellshock" thread

- the dkms packaging should completely hide all the complexities required in maintaining a kernel module

- Part of the kernel code, if you look at the contributors, has been written/reviewed by gregkh, so we like to think the quality is "high enough"

- There might be plans at some point to try and propose a merge of the code to mainline

gighi··on Easy, realtime, system-wide Shellshock monitoring
What do you get if you run "sysdig --version"?

If you used the official Ubuntu packages, those are a few versions behind upstream (currently at 0.1.87 while we are at 0.1.89): http://packages.ubuntu.com/trusty-backports/sysdig.

What we recommend is uninstalling those ones (sysdig and sysdig-dkms) and just use the binaries that we, Draios, provide, following this: https://github.com/draios/sysdig/wiki/How-to-Install-Sysdig-...

Should be very easy, and sysdig --version should show 0.1.89

gighi··on Ask HN: Do you still use an RSS reader?
I use feedly and this self-hosted application to read all sections of Hacker News: http://gianlucaborello.github.io/rssify/
gighi··on Hiding Linux Processes for Fun and Profit
That's absolutely brilliant!
gighi··on What If You Only Invested at Market Peaks?
Yeah, the condition you mention is the key: what are the odds that the US economy (or the overall world average, if you're diversified with a global index) will continue to grow as gloriously as it has done in the past 50 years? (when the guy started investing)
gighi··on What If You Only Invested at Market Peaks?
What would make this more interesting for me would be a pretty accurate answer to the question: "what are the chances that this could happen again if the guy started in 2014"? :)
gighi··on What If You Only Invested at Market Peaks?
Except that in this case, the guy bought the whole market by indexing, so he wasn't very wary and picky, although the index itself took care of getting rid of bad performers (because they went out of business).
gighi··on Fishing for Hackers: Analysis of a Linux Server Attack
Yes, I mounted the bucket using https://github.com/s3fs-fuse/s3fs-fuse
gighi··on Fishing for Hackers: Analysis of a Linux Server Attack
To answer your questions:

1) Yes, it would have been better but I honestly think this attack was completely botnet-driven and the attacker didn't really mean to cover his footprints too much: in the timespan of 10 minutes, he sent over 800 MB of UDP traffic. That would have been caught even by the most oblivious sysadmin pretty quickly, so these guys are just playing a number game, trying to break in as many hosts as they can knowing that the lifespan of the hacked hosts will be very short, maximizing the short-term profit then.

2) The attacker directly ran these commands on the login shell (no script was copied over scp or something else), so there was no script executed on the host itself, but the whole thing lasted roughly 2 minutes and a lot of commands were "typed", so I am almost sure this was just an automated script ran from another probably compromised host.

3) I didn't check if the build left logs, but by showing every executed process with "evt.type=execve" (which goes deeper than the spy_users chisel) you can see all the processes executed by the build: 99% are just uninteresting sed/gcc/autoconf.

gighi··on Fishing for Hackers: Analysis of a Linux Server Attack
I anonymized the IP addresses in a consistent way before publishing the blog post, so in the very worst case the DoS attack will go towards a completely new host :)
gighi··on Fishing for Hackers: Analysis of a Linux Server Attack
On some providers yes, in fact I explicitly enabled root SSH login for those.

Other providers (such as Digital Ocean) use the root account by default even for Ubuntu, although the password is set to a really secure and random one.

gighi··on Fishing for Hackers: Analysis of a Linux Server Attack
That's an interesting project.

Would it have recorded also statistics like the connection activity?

Seeing all the UDP traffic, and being able to trace its origin to the "@udp1 39.115.244.150 800 300" command, received not via shell but via a TCP connection, was pretty cool.

gighi··on Fishing for Hackers: Analysis of a Linux Server Attack
Yes, I didn't put it in the article because it was getting too long otherwise, but the attacker immediately tried brute-forcing the root account, and after a handful of common passwords ("qwerty", "qwerty123", "pizza" among those) he found "password".

I was able to find all the attempts by looking at the I/O activity of the sshd process, and also the syslog activity recorded every attempt.

gighi··on Fishing for Hackers: Analysis of a Linux Server Attack
Absolutely not, as I said in the article, I went out of the "wise" way and manually enabled:

1) Password authentication 2) Root authentication 3) Changed root password to "password"

All the providers offer fairly safe defaults, either using very random passwords or just enabling SSH keys.

gighi··on Show HN: Sysdig, a tool for Linux system exploration
We use CPack for the moment, so you can just run "make package" inside the CMake build directory and it will generate RPM/DEB.
gighi··on Show HN: Sysdig, a tool for Linux system exploration
That's the wrong way to use it.

"sysdig -w" switch will generate a binary dump (in a pcap format) containing the "raw events" coming from the kernel (plus a snapshot of information gathered from /proc), so it's not supposed to be human-readable, you have to use "sysdig -r" on the dump file to get the output.

If you're used to tcpdump, it's the same thing.