Why strace doesn't work in Docker
jvns.ca
jvns.ca
I had a discussion once with admins at my company that wanted to convert servers to have AD accounts in a naive way. The discussion could be shortened to something like this:
Me: please, don't convert servers to AD because my password will be available to everybody as everybody has root on test environment.
Admins: no, this is perfectly safe.
Me: Guys, once you have root on machine you can do everything including accessing passwords as they are being typed.
Admins: No, not possible. Perfectly safe.
Me: Here, I ran the tool using ptrace from root account, here is a list of all accounts and passwords of all users using test env. This is actually read() call at some point, you just need to be able to filter it out from strace output.
Admins: Wtf! You are going to get fired!
Your company either runs in a model where it trusts the employees or it doesn't. If it doesn't trust the employees they shouldn't be giving ANY of them root on a shared system.
On production servers on the other hand, the amount of people with root is probably much smaller, and there might be bigger wins by being able to centrally enforce password policies or whatever the sysadmins want to do.
90% of the time someone insists AD is too insecure for them, it's because they're unwilling to maintain hygiene (logging into an untrusted machines with overscoped credentials 8 times a day), or unwilling to use the provided security features (what do you mean use kerberized services? ldap binds work just find and give me an anti-ad security straw man!)
With poorly implemented AD (think telnet/ssh login app asking for password), on the other hand, you get password transmitted to the other box. At some point some application performs read() to actually receive it.
If you're running X you don't even need root :(
I do this about as frequently as I need to install something on remote server that has graphical installer. Not sure why people think thought this is a good idea but a huge amount of commercial software in 2000s had that kind of installation process.
In the 90's, firewalls were uncommon, so literally anyone on the Internet could sniff your keyboard or pop up windows on your display. Or, better yet, xmelt it.
This resulted in LOTS of mayhem.
I've gotten used to just going to read the source to see what an option does.
Wait what? I didn't know that. This sounds terrible. Is the root in the container the same user who runs the container, or the same root who is root on the host machine?
Root inside a non-userns container is the same as root on the Linux host, but it is constrained by security policies like seccomp and apparmor.
Yeah much easier to run everything as root. Also, we don't need /etc/password and stuff for UID/GID resolution and credentials in the container anyway; we just use ad-hoc auth and crypto from a random third-party lib (that isn't vetted, is never updated, runs in the same address space as your app, and hasn't access to meaningful entropy since running in a container) and supply root credentials on the docker command line or the Dockerfile checked in to github, or both.
That's what we've been saying for years: Docker doesn't solve anything, it merely hides problems from you (and helps your cloud provider's bottom line). Good luck, and yes, PHBs should be worried for civil/criminal gross negligence if the shit hits the fan. Your cloud provider is happy to take your checks, but will shrug-away and point out they're just providing the infrastructure; it's up to you to competently configure your ever-changing 12-factor k8s.
getrandom(2)/getentropy(3) should be using kernel randomness generation, which isn't affected by containers I thought.
This is one of the many reasons why giving non-root users access to the Docker daemon (so they can start containers) is dangerous: if, as a non-root user, I can start a container that’s running as UID 0, there’s a lot of possibility for misuse.
User Namespaces enable Docker to use separate UID pools for the containers, which enables a container to run as “UID 0 / root” but have the host actually map that as some arbitrary other UID, so the host can treat it differently than actual UID 0.
https://docs.docker.com/engine/security/userns-remap/
Also, Kubernetes doesn't support it.
Insane defaults, but easy to get started without knowledge
Edit: Note that userids are the same, but usernames are not managed by the kernel (handled by /etc/passwd or LDAP or something), so userid 1001 inside the container is userid 1001 outside, but they might have different names if you "ls -l" from different places.
1. Or, if they didn't require trusted ports, any account at all using https://github.com/NetDirect/nfsshell
A container is a process with some kernel isolation mechanisms setup around it. Unless one of those mechanisms is a user namespace with uids/gids mapped to an unused set of users, you get the same user.
Well it is. Most production web sites today are exploitable because of this.
It's even worse than it sounds.
How so? To exploit this, you need to already have RCE on a container. But generally you get that RCE by exploiting the site (the application code) in the first place.
In which scenario does an attacker have code execution privileges in a container, but needs this root privilege to exploit the site?
Here's what the linked commit to docker[0] says:
> 4.8+ kernels have fixed the ptrace security issues so we can allow ptrace(2) on the default seccomp profile if we do the kernel version check.
This commit itself links a commit in the linux kernel[1], which says:
> x86/ptrace: run seccomp after ptrace
> This moves seccomp after ptrace on x86 to that seccomp can catch changes made by ptrace. Emulation should skip the rest of processing too.
This doesn't give us much more information, and I'm not familiar enough with this code to understand the changeset. Thankfully, by looking up the changeset name, we can easily find the email that proposed the changeset[2], which explains the issue more in-depth:
> There has been a long-standing (and documented) issue with seccomp where ptrace can be used to change a syscall out from under seccomp. This is a problem for containers and other wider seccomp filtered environments where ptrace needs to remain available, as it allows for an escape of the seccomp filter.
So the basic idea is to use ptrace to swap an allowed syscall with a forbidden syscall. Seccomp used to run before ptrace, so it would see the allowed syscall and allow the code to continue, but then ptrace would swap the syscall to something forbidden! The details are a bit thin, however, they don't explain how we can do that. After a bit more digging, I found this pdf[3] which contains a simple POC showing the issue, and explaining it.
The TLDR is to use ptrace with `PTRACE_SYSCALLS` to execute until a syscall is hit. When a syscall is hit, seccomp would first check the syscall is allowed, and then pass execution to the tracing system of the kernel, that would stop the process for ptrace to inspect. From ptrace you can then modify the registers with `PTRACE_SET_REGS` to change the syscall being called (the syscall number is in register RAX on amd64) and resume execution. The kernel will then happily execute the modified, unchecked syscall!
So what changed is that seccomp will be run after ptrace now, so there isn't a way to modify the registers before the syscall is run anymore.
[0]: https://github.com/moby/moby/commit/1124543ca8071074a537a15d...
[1]: https://github.com/torvalds/linux/commit/93e35efb8de45393cf6...
[2]: https://www.mail-archive.com/linuxppc-dev@lists.ozlabs.org/m...
[3]: http://asm.rajiska.fr/misc/container_security.pdf page 9 for the explanation, page 16 for the POC code
I don't think the above would actually be all that useful, but still it's kind of cool that it's possible to do at all.
For this to work you'd also need a way to share various namespaces (like file descriptors) which is doable via the clone2 syscall.
There's a ton of cool tricks you can do with linux, it's a really powerful kernel.
ptrace: Operation not permitted. You can't do that without a process to debug. The program is not being run. gcore: failed to create core.36
I found this SO answer and it worked out for me.
https://stackoverflow.com/questions/42029834/gdb-in-docker-c...
(Change my mind?)
Do you have a technical argument to put forward?
For the difference of Docker and reinventing OSes, it depends on what you establish as reinventing OSes. Docker provides some form of separation/containerization, a function that is also implemented by some OS. But it's discussable of having the functionality in user or kernelspace or having a wrapper application with a consistent api across multiple operating systems isn't worth it.
Then why? I mean, I put it in parentheses for a reason man.
I mean, you seem to be saying that I can't tell the difference between automobiles and, what? horses?
Surely if the difference is that great it should be easy to come up with lots of solid points to "change my mind"? Since you're taking the trouble to comment anyway?
In your sib comment https://news.ycombinator.com/item?id=23071602 you seem to be making some solid points, but from my POV it still sounds like "learn the thing" where "the thing" does what I can already do without it, or it enables me to do stuff that I'm not actually interested in doing.
Bottom line: if you want to argue please do it with facts and not dismissive condescension. (I've got plenty of that already without your help, thanks.)
In fact I recently completed a migration away from Docker/Kubernetes to an AWS stack that doesn't use anything higher level than autoscaling groups and it has been a positive experience. The same application is now faster (less layers necessary), easier to deploy (we own all the moving pieces), easier to debug in production (I can use strace and tcpdump without trying to install them in an already running Docker container), etc. All of the "complexity" that Docker and Kubernetes were handling for us has been replaced with a ~ 500 line Python script that's mostly comments (you can see an early prototype of this script in the repository below, named deploy_build.py, which omits some error handling but weighs in at ~ 150 LoC). Furthermore, now that we have to think about how to package and deploy our application to production, the incentives are there to follow proper Python best practices around packaging (i.e. use packages, don't make assumptions about paths, etc) which has been a nice bonus.
Containers (and Docker, Kubernetes, etc) are all tools that have their use cases and corresponding tradeoffs. Anyone who can't explain when or why you wouldn't want to use a tool is not prepared to explain when or why you should use it.
Obviously, Microsoft also saw your same vision of containers being so unnecessary that they also decided to build the concept into their OS as well.
This concept can't be done on bare metal. Additionally, because it is built into the kernel and they share the same kernel, they generally startup at over 10x performance of traditional VMs. There are solutions now of course to boot directly into the kernel from a hypervisor, but these are generally paired with container solutions as well due to the orchestration and ecosystem that exists to make container orchestration easy.
Most container solutions use a unified image solution that allows you to rapidly reiterate and test your changes in multiple environments. Doing this on bare-metal or VMs takes considerably more time and money for fixed infrastructure costs.
With container solutions such as Kind and Helm, you can rapidly deploy a local cluster to pre-test cloud rollouts on a single machine. The orchestration can also autoscale clusters horizontally and vertically, and is robust enough to directly compete with other bulkier VM orchestration solutions such as OpenStack.
With automatic certificate rotation, automated service discovery, etc there is no need to try to create tasks for every little thing you'd need to do with Ansible or another IaC component. It is all baked into the Kubernetes ecosystem.
Abstraction away from cloud-specific solutions can only be seen as a good thing. People that are invested in AWS generally can't be trusted to create holistic solutions.
In particular, it has been useful to move our engineers to developing using it. We'll probably experiment with using it in production for a few things at some point, just for consistency, but it buys us zero value there otherwise. (We will never run "in the cloud" because it would be cost-prohibitive, at least with current cloud revenue models.)
There will be another new shiny in a few years, at which point the True Believers will move on and tell you how you'll lose your job if you don't get on board that train, too.
Good to know, this should open up some troubleshooting
by the same author. It looks "popular" but it's written much more thoughtfully and competently than most of the content found on the web. Like this article, it's the result of real personal research, it is never just rewording of some manual or already existing material. In short, really worth reading, and also worth using inside of the companies.