Podman: A Daemonless Container Engine
podman.io
podman.io
- podman layer caching during builds do not work. When we switched back to docker-ce, builds went from 45 minutes to 3 minutes with no changes to our Dockerfile
- fuse-overlayfs sits at 100% CPU all day, every day on our production servers
- podman loses track of containers. Sometimes they are running but don't show in podman ps
- sometimes podman just gets stuck and can't prune or stop a container
(edit: formatting)
Now I understand that btrfs is not really a requirement.
But I've been thinking about using podman to replace docker-in-docker on our fleet of gitlab runners where we build our images, and running btrfs is a deal breaker. We really don't want to add complexity to the mix. Fuse-overlay burning through the whole CPU is not really good looking either. Mh.
I still have to properly dig deep into this, but I'm keeping my hopes high.
Why? Storage is literally the backbone of technology. From databases to webservers there are unique and real requirements. There's a reason why the enterprise storage market is still worth several billion dollars, and it's not because storage is easy.
In 2021, with billions of dollars of R&D at their disposal - neither Amazon, Google, or Microsoft have even a mediocre NAS stack in comparison to the likes of Dell/EMC (Isilon) or NetApp (ONTAP). Heck they can't even compete with startups like Qumulo or Vast.
They have "competition" in the market, it's just horrible. Because storage is hard.
>Storage may be a multi-billion dollar market - but it's also razor thin margin compared to where those brands focus their R&D spend.
I'm going to let you go back and review the earnings reports from the aforementioned companies. Razor thin margins? You wouldn't survive in the market with razor thin margins.
> I'd argue those you've named don't have interest selling storage to the enterprise in the form of legacy boxes.
I'd argue Azure Stack, AWS Outpost, and Google Anthos readily prove you wrong.
> Cloud storage is also recurring revenue where enterprise storage is considered perpetual.
Just... no. Enterprise storage has a shelf life of 3-5 years before maintenance or technology make it obsolete.
>The latter is generally frowned upon from the Street's perspective these days (good, bad or otherwise).
https://www.google.com/finance/quote/NTAP:NASDAQ
5 years ago they were at $20/share, today they're at $69/share. Reality doesn't appear to match your claim as to the street's opinion of enterprise storage.
Yes, razor thin margins - comparatively. The cost of buying physical storage is commoditized. I've worked in technology from the pre-sale engineering side for a number of years. Margins on hardware are fractional compared to software and subscription offerings.
Stack, Outpost and Anthos are not revenue drivers today. They're vendor lock in tools.
3-5 years is a horrible argument in the position you're attempting to make considering it's an eternity in technology. Consider that you sell the storage once in that 3-5 years. Your cloud provider bills you monthly and you don't own it. It's recurring revenue vs a singular sales event.
NTAP over 5 years is a decent return. 215% over that timeframe. Amazon returns 543% in the same timeframe. And let's use a hot space that's 100% subscription based, like Crowdstrike: 265% in less than a year. Software and subscription margins crush the legacy perpetual hardware model.
Storage is an old, stable, and commoditized market. There's no huge growth (any storage earnings report show that very clearly). And while the demand for storage continues to increase, the NetApps of the world aren't capturing where that growth is.
I'm sorry but you're just flat out wrong and didn't bother to look up margins like I suggested you do.
NTAP 2020 margins: 66% https://www.macrotrends.net/stocks/charts/NTAP/netapp/gross-...
AMZ 2020 margins: ~41% https://www.macrotrends.net/stocks/charts/AMZN/amazon/profit...
An all time high... and not even close.
>Storage is an old, stable, and commoditized market.
Which is why... per my original post... the cloud providers struggle to provide basic features in 2021 that have been available to enterprise storage customers for 20+ years.
Second thing is gross margin is not the same thing as net margin. Net margin is a far better indicator in the case of pure profitability / revenue. Just because a company shares $1.42B in revenue doesn't mean they have $1.42B in profits. If you look at NetApp net profit margin it's off 45% year over year. Latest net profit margin is 9.68%. So on that 1.42B we have $137M in profit. Not too hot.
The reality is hardware sales have thin margin (comparatively) and NetApp's software model is much weaker than comparative storage offerings in the cloud. It's very expensive to design, build, ship and house inventory of hardware for purchase. There's no way around this, which chews through that profit margin.
Yes, I'm aware Amazon continually buries their AWS profitability. It's 57% of the overall margin, it's not significantly better than NetApp.
>Second thing is gross margin is not the same thing as net margin. Net margin is a far better indicator in the case of pure profitability / revenue.
Net margin is a dial they can and do turn by shifting money in and out of R&D and sales staff. Furthermore you can't simultaneously say that for Amazon we only take into account AWS and not the rest of the business... which is a part of the engine that drives the AWS business. That's no different than saying we shouldn't account for hardware in NetApp's number because they also sell software.
>It's very expensive to design, build, ship and house inventory of hardware for purchase. There's no way around this, which chews through that profit margin.
What exactly do you think Amazon's datacenters are full of? Hardware they design, build, and ship inventory around the country for. They are working with the exact same ODM's that NetApp or any other storage vendor works with to design their custom hardware variations.
All balance sheets can and are manipulated to differing extent. The reality though is net margin takes better into account than what you were proposing, which is a cherry picked angle to make a point. Growth companies move profits to R&D. NetApp isn't highlighting R&D on earnings calls. Amazon highlights R&D very publicly multiple times per year. Low R&D costs and investment are indicators of a stagnant market.
Amazon's hardware model is not the same as NetApp's. You realize Amazon consumes almost 100% of their own hardware offerings for AWS, right? They have a much more palatable JIT model. They can build as they need. NetApp has to house dead weight inventory for potential sales so they can recognize revenue, especially for end of quarter / end of year sales pushes. Amazon doesn't have this problem.
Also NetApp has to package their product for customer distribution. You may ignorantly shrug this off but it's an additional cost that adds up quickly. NetApp also has to maintain documentation and support staff for customer facing hardware related issues. Amazon has a much lighter requirement because they have specialists within their DCs.
It's not even remotely comparable. Amazon isn't working with "the exact same ODMs". Many ODMs NetApp is forced to work with Amazon doesn't need. Take a look at what Amazon is doing internally and it's clear their stack continues to evolve more towards in house designs and build. NetApp is far more reliant on external support than Amazon is given their positions in the market and have far greater control over the hardware stack from top to bottom. Maybe peruse the career openings at AWS vs NetApp for some insights.
Netapp is love, Netapp is life.
I have fond memories of cloning a volume, from a snapshot, in seconds. On EBS it takes minutes.
Layer caching seems to work for me. Note that rootless stores images (and everything else), in different place from rootfull. It may be that you're caching the wrong directory.
I want to look at the problems. But I don't want to do a guessing game with the 159 open issues, which one may relate to one of these points or not.
I find it very strange that your reaction is, "the software is infallible, it is _YOU_ who have failed the software by not logging bugs!" It is perfectly reasonable to have a conversation about the quality or issues with a piece of software SEPARATE TO the logging or finding of bugs in said software.
I can sympathize with feeling overwhelmed when searching for similar Issues in a bug database, as I was in a similar situation in the past.
What happend was this: I went to the central starting point of the repo's database (e.g., https://github.com/containers/podman/issues), removed the filters like `is:open` from the search bar, and entered one or two keywords that for me sounded similar to the problem I faced.
As I didn't find any Issues that seemed similar, I just created a new one. Of course, I was worried about creating a possible duplicate, but I figured that it would still be a valuable contribution by adding more keywords (with the maintainers linking my duplicate to the ticket where the problem gets resolved) for others to search for. And last but not least, I thought this would also add additional debugging information for the developers, as well as an indication of the community-wide impact of the problem for the release management team.
In the end, I was quite happy I went this way, because I didn't just make a small impact on a project that would be useful for me, but it also enabled me to connect to the amazing people that build the tools I use on a daily basis.
I hope that you can also find a way that gives you some happiness in dealing with your situation, yobert.
I've taken a look at podman from time to time over the years but it seems like it's just never formalized, never been polished and almost always has been sub-par in execution. On this list the builds and container control are things that I've run across. I guess - what's the point? The rootless argument leaned on so heavily is pretty much gone, the quality of Podman hasn't (seemingly) improved and now IBM owns Red Hat (subjective, but a viable concern/consideration given what's recently happened with CentOS).
You're more than safe leveraging Docker and buildkit (when and where needed). Quite honestly, given the relatively poor execution of Red Hat with these tools over the years, I don't see the point. I'm sure there are some niche use cases for Podman/Buildah, but overall it just seems to come up as an agenda more than an exponentially better product at this point. Red Hat could have made things better, instead they just created a distraction and worked against the broader effort in the container ecosystem.
[1]: Apparently initial support planned in version 3.0?
Sadly it doesn't feel as polished as docker-compose, I can only assume it's due to the lack of API. Specifically, podman-compose just translates all the docker-compose.yml file directives into podman cli commands, which doesn't seem to be handled gracefully.
This is a misrepresentation. The situation was that Docker didn't take patches, as some were very specific changes for systemd and lack of unionfs, etc but over time it applied to most patches from RH associated people.
Edit: You also may have different perspective given you work for Red Hat / IBM.
It's just minor difference to me except it doesn't work well.
Not sure what the laziness was about from RedHat on podman development.
You have root in the container without having root in the host system. That takes care of a lot of issues as well.
What benefits would toolbox add ?
I use bindfs to mount the volume. I have a $HOME/Dev folders with WPProjectA, WPProjectB folders. Each has a volume subfolder mounted like that (the script has more variables but that's the gist of it):
/usr/bin/bindfs \
--force-user=johnchristopher \
--force-group=johnchristopher \
--create-for-user=www-data \
--create-for-group=www-data \
/var/lib/docker/volumes/WPProjectA-web/_data \
$HOME/Dev/WPProjectA/volume
This setup allows using VSCode+xdebug and editing the code in the mounted volume while running the container and remote debugging.Toolbox emphasises "keeparound" containers since it's intended to be the primary command-line environment for image-based systems like Silverblue or CoreOS. Such systems try to keep a small, atomically updated rootfs and push users to install everything in containers.
On other occasions, they just hijack the name. Like they did with dstat.
Making that assumption was reasonable - moving forwards without at least making an attempt to contact the maintainer "just in case" was not. It would have been courteous to at least try.
But it's not exactly the deliberate hostile takeover that you make it out to be.
https://bugzilla.redhat.com/show_bug.cgi?id=1614277#c9
>> To my knowledge, there is no need (legally) to obtain consent to use the name 'dstat' for a replacement command providing the same functionality. It might be a nice thing to do from a community perspective, however - if there was someone to discuss with upstream.
>> However, dstat is dead upstream. There have been no updates for years, no responses at all to any bug reports in the months I've been following the github repo now, and certainly no attempt to begin undertaking a python3 port.
>> Since there is nobody maintaining the original dstat code anymore, it seemed a futile exercise to me so I've not attempted to contact the original author. And as pcp-dstat is now well advanced beyond the original dstat - implementing features listed in dstat's roadmap for many years, and with multiple active contributors - I think moving on with the backward-compatible name symlink is the best we can do.
> Runtime: We use the OCI runtime tools to generate OCI runtime configurations that can be used with any OCI-compliant runtime, like crun and runc.
Red Hat literally took a memory safe Go program and rewrote it in C for performance, in 2020.
It’s been years since I’ve been able to get excited about rewriting systems from a memory safe language into a memory unsafe language for “performance reasons”. As an industry, we have too much evidence that even the best humans make mistakes. Quality C codebases still consistently run into CVEs that a memory safe language would have prevented.
Rust exists, if performance is so critical here and if Go were somehow at fault... but it sounds like most of the performance difference is due to a different architecture, so rewriting runc in Go probably could have had the same outcome.
I wish I could be excited about crun... I’m sure the author has put a lot of effort into it. Instead, I’m more excited about things like Firecracker or gVisor that enable people to further isolate containers, even though such a thing naturally comes with some performance impact.
gVisor is also very exciting due to how it handles memory allocation and scheduling. Syscalls and I/O are more expensive, though.
I agree completely!
I suspect it has been a bottleneck in some of the systems redhat has worked with, or they would have no benefit in writing it. As the authors of openshift I’m sure they’ve seen some pretty crazy use cases. Just because it isn’t your use case doesn’t make it invalid.
Also, small c projects can be written properly with discipline. This is a very talented team on a very small and well scoped project. It can be written properly.
The C++ approach is to offer a feature-rich language so the programmer doesn't have to reinvent common abstractions. The C approach is to offer a minimal and stable language and let the programmer take it from there. It's not obvious a priori which approach should result in fewer memory safety issues. If I had to guess my money would be on C++ being the better choice, as it has smart pointers.
Please find me a well-known C code base that doesn't suffer from memory safety CVEs. You're the one making the claim that they exist... I can't prove that such a thing doesn't exist, but I can point to how curl[0][1], sqlite[2][3], the linux kernel[4][5], and any other popular, respected C code base that I can think of suffers from numerous memory safety issues that memory safe languages are built to prevent.
Writing secure software is hard enough without picking a memory unsafe language as the foundation of that software.
"Just find/be a better programmer" isn't the solution. We've tried that for decades with little success. The continued prevalence of memory safety vulnerability in these C code bases shows that existing static analysis tools are insufficient for handling the many ways that things can go wrong in C.
I think you mentioned railcar elsewhere in this thread, and I would find it easier to be interested in a runc alternative that is written in a memory safe language... which railcar is.
[0]: https://curl.se/docs/CVE-2019-5482.html
[1]: https://curl.se/docs/CVE-2019-3823.html
(among numerous others)
[2]: https://www.darkreading.com/attacks-breaches/researchers-sho...
[3]: http://cve.mitre.org/cgi-bin/cvekey.cgi?keyword=SQLite
(the second link has quite a few memory safety issues that are readily apparent, including several from SQLite that affected Chrome)
Yeah and C++ would be a way better language to write critical system daemon in 2020 than C. Both safer and more productive while keeping the exact same portability and performance as C when necessary.
Most safety issues of C (buffer overflows, use-after-free, stack smash) are not a problem anymore in modern C++.
Yes, writing new userland software in C in 2020(1) is non-sense.
- Use at least C++ if you are conservative.
- Use Rust if you aim for absolute safety.
- Use Go if you accept the performance hit and do not need libraries.
There is zero excuses to C in userland in 2020.
The only excuse is some Red Hat folks seems to practice C++ hating as some kind of religion.
That exactly what give us "beautiful" monstrosities like systemd or pulseaudio with their associated shopping list of CVEs [^1]
Even GCC maintainers switched to C++, by the sake of god, do the same.
----
[^1]: https://www.cvedetails.com/product/38088/Freedesktop-Systemd...
I wouldn't go quite that far. I'm generally in the C++ camp (rather than C, that is) but there are significant advantages to a compact and stable language, and real disadvantages to a sprawling disaster like C++ that grows still more monstrously complex every few years.
This topic turned up a year ago: https://news.ycombinator.com/item?id=21946060
Your point stands though: well written modern C++ should be much less prone to memory-safety issues than well written modern C.
You can ensure memory-safety in C by enforcing strict rules since the first commit.
Talent does not preclude you from making mistakes
Most projects start out small and well scoped. That doesn’t mean it’s going to last.
See also: podman; cri-o; buildah.
“Grudge?” (proceeds with details of grudge).
I don't think it's a stretch to look at the events as described and see them as Red Hat doing what they need to work around someone else's grudge. It could be any number of things, but I don't think the information presented so far lends itself towards Red Hat having a grudge.
If I carpool with someone and they like to go around and loudly proclaim "I refuse to listen to anything kbenson has to say", maybe I find someone else to carpool with or drive myself? That only seems prudent. That doesn't mean I have a grudge against that person, it just means I don't want to trust them with getting me to and from work anymore.
The drama you’re talking about is from that period - 2015-16. But “crun” was launched in 2020, a full four years later. Frazelle left Docker years ago, as well as most of the people involved in those feuds. Runc, containerd and docker are chugging along drama-free, and are maintained by multi-vendor teams. There is no interest from anyone outside of Red Hat in forking or re-writing those tools. It’s just a tremendous waste of energy and in the end will be problematic for Red Hat, because they will be building their platform on less used and therefore less reliable code.
The world’s container infrastructure runs on containerd, runc and docker. Not crio, crun and podman. Someone at Red Hat clearly is having trouble accepting that reality, in my opinion because they’re having trouble moving on from old grudges. I really wish they did, because all of that wasted engineering effort could be deployed to solve more pressing problems.
> The world’s container infrastructure runs on containerd, runc and docker
I don't consider alternate implementations of widely used software to be wasted evergy almost ever. Was Clang a waste of energy because everyone was using GCC? There are many reasons why multiple implementations may be useful, competition and being able to cater to slightly different common use cases obvious ones.
I'm not sure why any end user would wish for there not to be multiple open source projects looking to satisfy a similar technical need.
> in the end will be problematic for Red Hat
That may be, but I don't really care about whether it's good for Red Hat. I care that it increased user choice and maybe at some point features and capabilities offered to users (whether through podman or pressure on docker).
Thanks for the evidence though, it's always best to have, even if it can be a pain to drum up when it might not really be called into question anyways. :)
For reference, this was the shirt that Docker Inc was giving away at Redhat summit years ago. I was there. I saw the shirt, and thought it was in very poor form. https://twitter.com/SEJeff/status/1125871424126767104
As Rambo said, "But Colonel, they drew first blood."
Containerd was donated (also by Docker) to CNCF, a different organization, and I believe a few years later.
You are correct that Docker did those things because of pressure to be more interoperable.
Reading your linked tweet (“don’t mess with an engineering behemoth”) you seem to agree that Red Hat is rewriting away Docker code because of a past grudge?
OCI would have made docker inc entirely irrelevant if they didn’t join it. They donated runc (I was wrong originally, dyslexia sucks sometimes) as the reference implementation so docker continued to stay at the forefront.
They donated a very minimal implementation of containerd to the CNCF when kubernetes wanted to do things with the container runtime interface (CRI) to make it more pluggable. As a more purpose built CRI, crio is better suited for k8s.
Docker did not create OCI any more than Redhat did. Redhat and a collection of other orgs did. Docker and Redhat were besties after Alexander Larson added device mapper support to docker. This is what allowed docker to run on any Linux distro that didn’t include the out of tree aufs. Look it up. Redhat then made a lot of money doing container multihost orchestration with openshift v3, which deprecated their own tool, geard. When Docker inc accepted a ton of VC money, they realized they actually had to monetize things. This is the part where the relationship started to seriously sour. The company that helped make them so popular was now suddenly a huge competitor. This is what lead to the problem.
It wasn’t a grudge from Redhat I mentioned in the linked tweet. It was docker inc thinking docker was anything other than commodity. It just is some very nice UX and cli tooling with a container registry around Linux namespace and Linux control groups. They just put it all together very nicely for developers to ship code faster. Podman is also a commodity. At its core, it just is some json parsing and pretty wrapping around Linux namespace and cgroups. Redhat knows this, and they don’t hide it. They monetize openshift and quay, not podman.
Sorry if this is a bit scrambled. Long posts on mobile are difficult.
Neither Red Hat nor CoreOS were involved in any of those steps, nor were they even aware of them until the last minute. When they were invited it was on a “take it or leave it” basis. They took it, then tried to pre-empt the Docker/LF announcement by a few hours to make it look like they were launching it. Those were tense days and there was very little trust between those companies.
Source: I was involved in the process of creating OCI.
I actually have a slightly different explanation of what set it off. I don’t buy the “docker is suddenly under pressure to make money” explanation. Docker was a VC funded startup since 2010. They already had a business when they pivoted to Docker. As far as I know they kept the same investors and board. So it seems unlikely that they suddenly remembered that they needed to make money. More likely they planned to make Docker as ubiquitous a platform as they could, at which point there are many known avenues to monetize. Also, the CEO they brought on board, Ben Golub, had just spent a year at Red Hat after selling his previous startup to them. So he was certainly familiar with Red Hat’s business and strategy, and probably on friendly terms with their leadership (this is speculation on my part).
And in the early days Red Hat and Docker were in fact partners. Docker engineers even implemented features specifically to please Red Hat; for example devicemapper support and storage drivers, which were developed by Docker for Red Hat.
The problem is that the partnership was built on a misunderstanding. Docker was hoping for a distribution deal between the two businesses, with revenue share etc. and intended to keep control of their open-source project. Whereas Red Hat was hoping to make Docker their “new Linux” which involved no partnership with Docker and instead taking or at least sharing control of the project as they did with Linux. Each side was slow to realize the true intentions of the other, whether by naivete or deception I don’t know. This still surprises me because both side’s intentions were rational and entirely predictable. I think a lot of wishful thinking, and possibly some individual incompetence in leadership were involved. For example Red Hat has never actually partnered with a smaller startup and shared revenue with them in the way Docker was hoping, so I’m not sure what made them so confident it would happen. In any case, when the fog lifted, bitterness and conflict quickly followed.
In the end Red Hat shifted gears to Kubernetes as their “new Linux” which was a much better match for them. Then proceeded to rewrite history to minimize Docker’s role; which is unfortunate but understandable, I guess. I know a lot of Docker employees are personally hurt by it, to this day. It doesn’t feel good to see your work swept under the rug for reasons of competition and ego.
https://github.com/moby/moby/pulls?page=2&q=is%3Apr+is%3Aclo...
In fact, Solomon's own words from https://github.com/moby/moby/pull/2609 "Integrate lvm/devicemapper implementation by @alexlarsson into a driver", so unless you're more knowledgeable about the history than Solomon, I'm going to have to discount your take on this entire thing.
Regardless, I think things are at a much better place now. It is unfortunate that there were tensions and more so that Solomon had such a tough time through it all. He's a super nice guy. But this is tech, and there will always be lots of personality conflicts. It is one of the precious few constants.
You can see all this from the early history of the devmapper directory: https://github.com/alexlarsson/docker/commits/a14496ce891f1f...
That seems like a serious allegation. Can you back that up?
I think OP wanted to say that Podman hates Docker what is not I feel when I'm interacting with the community there. People who use Podman do it because of it's additional features that Docker does not have, like starting an Container from a rootfs or mounting the currect directory in a container using "." as path. Also building containers using a shellscript is supported. No need to learn the Dockerfile syntax. There are lot of small things that make Podman better integrated into Linux systems: running the container via systemd is a lot more intuitive.
[0]: https://www.weave.works/blog/linux-namespaces-and-go-don-t-m... [1]: https://github.com/drahnr/railcar
Portability seems like an unlikely reason here, but I would pick memory safety over portability every single day of the week. Rust is far from the only language that offers memory safety, but it is certainly one of the most compelling ones for projects that would otherwise be written in C.
"runc" is battle tested, written almost entirely in Go, which offers memory safety, and the performance delta under discussion is only about 2x for a part of the system that seemingly only has a measurable performance impact if you're starting containers as quickly as possible.
This just isn't a very compelling use case for C.
If I worked for RedHat I would still seek to write as portable as possible just on general principle or for the sake of my own future self.
But I agree it's merely a possible reason and may not be a major one.
This was addressed in 2017/2018 [0], it's no longer a poor choice.
[0]: https://github.com/golang/go/commit/2595fe7fb6f272f9204ca3ef...
[0]: https://github.com/opencontainers/runc/tree/master/libcontai...
Yes it’s easier to shoot yourself in the foot, but with practice you can learn to miss every time.
djb wrote remotely exploitable C for qmail, applications written by the OpenBSD team have had exploitable memory issues, Microsoft still doesn't get it right, neither do Linux kernel devs.
I'm very glad languages like Go and Rust exist, but saying you shouldn't write C because you might create memory leaks is kind of like saying you shouldn't write multi-threaded code because you might create race conditions. Yeah, it adds complexity to your code, but it's sometimes worth the overhead it saves. Whether that trade-off is worth it is always up for debate.
I work on avionics and we use C/C++ now and then. We have a ton of rules regarding memory management (pretty much everything stays on the stack) and I can't recall anything I've ever been involved with suffering from a memory leak.
Granted, I learned to write C 25+ years ago, and have worked as a C programmer for 15 of them, writing mostly embedded software for mobile phones (pre smartphone), airport sorters, but have also written financial software (mostly for parsing NASDAQ feeds), but the point is that most of the software I’ve written has had close to a decade of “runtime”, and while I started out making the same mistakes as everybody else, you learn to put ranges on your heap pointers like strncpy instead of just blindly doing strcpy.
Checking the size of your input vs your allocated memory takes care of a lot of it. As for memory leaks, it’s not exactly hard to free a pointer for every time you malloc one.
People are terrified of C, and yes, Go, Rust, Java, C# makes it harder to make _those_ mistakes, but that doesn’t mean it’s impossible to write good C code.
And it’s not like projects written in Go or Rust are error free. They just struggle with different errors.
As for good C projects, check stuff like postfix, nginx, and yes Linux or FreeBSD.
Writing in C doesn’t create complexity but lots of traps to fall in. Writing in C doesn’t save overhead over Rust. Your rust code can be very low level and safe and C isn’t that close to hardware as it used to be anyway.
https://github.com/opencontainers/runc/blob/4d4d19ce528ac40c...
https://github.com/opencontainers/runc/blob/4d4d19ce528ac40c...
This seems like a manageable amount to very carefully maintain and not at all like writing something entirely in C.
A huge part of a container runtime is pretty much just issuing Linux syscalls - and often the new ones too, that don’t have any glibc bindings yet and need to be called by number. Plus, very careful control of threading is needed, and it’s not just a case of calling runtime.LockOSThread() in go - some of these namespace related syscalls can only be called if there is a single thread in the process.
It’s _possible_ to do this in go, of course; runc exists after all. But it’s certainly very convenient to be able to flip back and forth between kernel source and the crun code base, to use exactly the actual structs that are defined in the kernel headers, and to generally be using the lingua franca of interfacing with the Linux kernel. It lowers cognitive overhead (eg what is the runtime going to do here?) and lets you focus on making exactly the calls to the kernel that you want.
I would very much like to use Podman as a finally proper container launcher in production (non-FAANG scale - at which you maybe start to need k8s), but having an unnecessary daemon moving part in thousands lines of C makes me frown so far.
We make heavy use of Podman in our infrastructure and it's mostly a pleasure. My current pet peeves are that:
1) Ansible's podman_container module is not as polished as docker_container. I regularly run into idempotency issues with it (so lots of needlessly restarted containers).
2) Gitlab's Docker executor doesn't support Podman and all our CI agents run on CentOS 8. I ended up writing a custom executor for it and it's working quite well though (we're probably not going back to the container executor even if it supported Podman, since the custom executor offers so much more flexibility).
3) GPU support is easier/more documented on Docker. For this reason, the GPU servers we have are all Ubuntu 20.04 + Docker since it's the more beaten path.
4) Podman-compose just needs more work. Luckily for us, it seems that Podman 3.x will support docker-compose natively [1].
As mentioned, our CI environment is very dependent on Podman. The first step of every Gitlab pipelines is to build the container image in which the rest of the jobs will run. I find that it's simpler to have a shell executor in a unprivileged, restricted environment (i.e. can only run `podman build`) than setting up dind just for building images. All jobs that follow are ran in rootless containers, for that nice added layer of security.
Wishing all the best to the Podman, Buildah and Skopeo teams.
I see in the docs [1]:
"Podman is a tool for running Linux containers. You can do this from a MacOS desktop as long as you have access to a linux box either running inside of a VM on the host, or available via the network. You need to install the remote client and then setup ssh connection information."
And yes, I know that Docker (and any other Linux container-based tool) also runs a VM—but it sets it up almost 100% transparently for me. I install docker, I run `docker run` and I have a container running.
For other use cases I imagine it’s just fine.
Why does that seem 'kind of crazy' to you? A container is really just a namespaced Linux process, so … you either need a running Linux kernel or a good-enough emulation thereof.
What seems kind of crazy to me is that so many folks who are deploying on Linux develop on macOS. That's not a dig against macOS, but one against the inevitable pains which arise when developing in one environment and deploying in another. Even though I much prefer a Linux environment, it would seem similarly crazy to develop on Linux and deploy on Windows or macOS. Heck, I think it is kind of crazy to develop on Ubuntu and deploy on Debian, and they are generally very very close!
That's exactly not true, considering you said Linux.
The utility of Linux containers is that they share the same OS instance, but have an embellished notion of Unix process group where, within a container i.e. embellished process group, they see their own filesystem and numbering for OS resources, AS IF they were on individual VMs, but they're not.
I wish they had tried to collaborate with Docker and contribute upstream instead of this project.
On Windows, the WSL2 feature gives you that easy wrapper in a different way. It's sets up and manages the underlying VM although you have to choose your distro etc. Once this is running after a simple setup you are just using Linux from that point on. It's less specialized than how Docker on macOS seems to work.
If someone knows of something that follows the WSL2 approach without VirtualBox or VMware Fusion I'd be all ears. That would be more versatile than how Docker seems to work right now. Docker's business interests probably aligned well with getting a workflow running, so unless someone is motivated to do similar for Podman, you are going to be out of luck. At least recognizing this deficiency would be a start though.
Comments like this help reinforce the stereotype of lazy macos developers ;)
Containers use Linux namespaces to decouple from the main OS. MacOS doesn't support those, so no matter what you do, you can't run them directly on MacOS. That's why you need a VM, and you need the WSL Windows subsystem for Linux on Windows.
BSD has jails, but they don't have the same functionality. In particular, I sorely miss the network namespace that Linux has on MacOS.
There are ways around it: raise the ulimit for your user and run new enough podman to raise limit for fuse-overlayfs, use the 'vfs' driver (it has other perf issues). I heard (but haven't tested yet) that the 'btrfs' driver avoids all these problems and works from userspace. Obviously requires an FS formatted as btrfs...
There are also compatibility issues with Docker. E.g. one container was running sshfs inside Docker just fine, but fails with a permission error on /dev/fuse with podman.
Also, docker runs as root, so it won't have permissions problems. You can change the permissions of /dev/fuse if you want to allow podman containers to access it or update the group of the user launching podman.
The lack of 'cannot connect to dockerd' mysteries makes for a much-improved developer experience if you ask me.
All I had to do was 's/docker/podman/g' and remove the chown hack and it works fine: https://github.com/sevagh/pq/commit/6acf6d05a094ac2959567a9a...
It understands Dockerfiles and can pull images from Dockerhub.
Also, Docker does support live restore if you want to keep containers running over daemon restarts https://docs.docker.com/config/containers/live-restore/
Similarly, the docker security model is that there isn't a security model. If you can talk to the docker socket you have what ever privileges the daemon is running as.
Second point, yep if you run Docker as root and someone can access the socket file they get root.
If that's a concern, you can run Docker rootless.
And as we're talking file permissions on a local host to allow that access, the same applies to podman containers does it not? If there are permission issues allowing another user to use access the container filesystems, you have the same problem.
As to swarm, I was comparing Docker rootless to podman, which is more a developer use case than prod. container clusters.
[0] I suppose you could have a user-level daemon that runs for each user that needs to run containers, but that's even more overhead.
There's some tradeoff I guess though, between rootful setup and per user, as images duplication per user could add up.
I used Arch Linux and followed their docs[1], mainly because I wanted it to also run without root. Big mistake. Attempting to run a simple postgres database w/ and w/o root permissions for Podman resulted in my system getting into an inconsistent state. I was unable to relaunch the container due to the state of my system.
I mucked around with runc/crun and I ended up literally nuking both the root and non-root directories that stored them on my computer. I reset my computer back to Podman for root only, hoping that I had just chosen a niche path. It still would leave my system broken, requiring drastic measures to recover. No thanks.
After much debugging, finally switching back to Docker, I realized my mistake: I had forgotten a required environment variable. Silly mistake.
Docker surfaced the problem immediately. Podman did not.
Docker recovered gracefully from the error. Podman left my system in an inconsistent state, unable to relaunch the container and unable to remove it
Again, I'm not sure how much of my experience was just trying to force Podman into a square hole? I'm sure there are other people that make it work just fine for their use cases
Edit: I should note that I used it as a docker-compose replacement, which is probably another off-the-beaten path usage that made this more dramatic than it should have been
That makes it much more convenient for build environments, without having to hard-code user IDs both in the container and on the host.
So all you need to do is create a very simple base container layer that just installs sssd-client, and wires up /etc/nsswitch.conf to use it (your package manager will almost surely do this automatically). Then just bind mount the sssd socket into your container and boom, all your host users are in the container.
If you already log in with sssd you're done. But if you only use local users then you'll need to configure the proxy provider so that sssd reads from your passwd on your host. In this case the host system doesn't actually have to use sssd for anything.
(Also I'm not sure how that's better (and not just different), except maybe it allows more than one host user in the container, but I haven't had a use case for that).
Totally eliminates dependency hell, e.g. ROS heavy workflows where it wants to control every part of your environment.
* The newer version of Docker+BulidKit supports bind-mounts/caching/secrets/memory filesystem etc. we used it to significantly speed up builds. You could probably find a way in buildah to achieve the same things, but it's not standard.
* Parallel builds - Docker+BulidKit builds Dockerfile in parallel, in podman things run serially. The combination of caching and parallel builds and Docker being faster even for single-threaded builds, Docker builds ended up an order of magnitude faster than the Podman ones.
* Buggy caching - There were a lot of caching bugs (randomly rebuilding when nothing changed, and reusing cached layers when files have changed). These issues are supposedly fixed, but I've lost trust in the Dockerfile builds.
* Various bugs when used as a drop-in replacement for Docker.
* Recurring issues on Ubuntu. It seemed all the developers were on Fedora/RHEL, and there were recurring issues with the Ubuntu builds. Things might be better now.
* non-root containers require editing /etc/subuid, which you can't do unless you have root access.
More information about the new BuildKit features in Docker:
We also migrated back to Docker from Podman because of Buildkit and the general bugginess of Podman.
If your environment would benefit from smaller, isolated containerization tools then my two cents would be that it's worth keeping an eye on these as they mature, and perhaps perform early evaluation (bearing in mind that you may not be able to migrate completely, yet).
The good:
- Separation of concerns; individual binaries and tools that do not all have to be deployed in production
- A straightforward migration path to build containers from existing Dockerfiles (via "buildah bud")
- Progressing-and-planned support for a range of important container technology, including rootless and multi-architecture builds
I feel like the podman experience is binary, either it works perfectly or it sucks balls. My suggestion is to give it a try, if it fails, then maybe file a bug report, fail fast and fall back on docker.
I was researching this few hours ago and according to https://github.com/containers/podman/issues/6114#issuecommen... it just works when you add another network.
Docker registry having no IPv6 is another fun story tho.
IPv6 support is always like this: half-baked at best, someone got it to work somehow, developers declare it done since it worked for someone. Then crickets...
IPv6 support isn't done until you can replace IPv4 with it, with no changes to your config except addresses. Even docker isn't there yet. And podman's is still in a larval state.
Through it has similar drawbacks as using docker-compose with docker.
As far as I know there is currently no rootless way on Linux to setup custom internal networks like docker compose does (hence why podman-compose doesn't support such things).
But this can't be done rootless.
So with rootless podman all container map to the same ip address but different ports.
This is for some use cases (e.g. spinning up a DB for integration testing) not a problem at all. For others it is.
More over you can run multiple groups of docker containers in separate networks, you can't do so with rootless podman.
Through you can manage networks with rootfull podman (which still has no deamon and as such works better with e.g. capabilities and the audit sub system then docker does).
Through to get the full docker-compose experience you need to run it as a deamon (through systemd) in which case you can use docker-compose with podman but it has most of the problems docker has.
https://www.redhat.com/sysadmin/podman-docker-compose
That says, you literally can use the same docker-compose app shipped from docker and it works using podman.
I am not impressed with podman. It's buggy and slow. Documentation is very basic and the underlaying mechanics are not covered where needed. Want to reconfigure the command parameters your start your container with to add a exported variable? Gotta remake the entire setup because there's no edit command.
- image unpacking issues (invalid tar header)
- linking containers isn't supported
[EDIT] - I wrote up the post real quick[2]
[0]: https://vadosware.io/post/rootless-containers-in-2020-on-arc...
[1]: https://vadosware.io/post/bits-and-bobs-with-podman/
[2]: https://vadosware.io/post/back-to-docker-after-issues-with-p...
* Manual config tweaking may be required.
You still can run podman as root without damon.
Or you can run a podman deamon (through systemd) which is even compatible with docker-compose but has most of the drawbacks of running docker.
# /etc/cni/net.d/testnet.conflist
{
"cniVersion": "0.4.0",
"name": "testnet",
"plugins": [
{
"type": "bridge",
"bridge": "br0", # main host interface is part of this bridge
"ipam": {
"type": "host-local",
"subnet": "10.0.0.0/16",
"gateway": "10.0.0.1",
"routes": [{ "dst": "0.0.0.0/0"}]
}
}
]
}
You can then start a container and operate on its network namespace for added flexibility: podman run -it --net testnet --ip 10.0.0.2 ...
ns=$(basename $(podman inspect $id | jq -r '.[0] .NetworkSettings .SandboxKey'))
ip netns exec $ns ip route add ...
[1]: https://github.com/containernetworking/cniAFAIU, Podman v3 has a docker-compose compatible socket and there's a daemon; so "Daemonless Container Engine" is no longer accurate.
"Using Podman and Docker Compose" https://podman.io/blogs/2021/01/11/podman-compose.html
What does being "daemonless" actually buy me in practice?
rhel ~forces you to do podman in rhel 8+. They make it hard enough to do a docker install that corp policies will reject docker (non-standard centos 7 repos etc), while allowing ibm's thing. Ex: Locks podman over docker for a lot of US gov where rhel 8 is preferred and with only standard repos, and anything else drowns in special approvals.
subtle and orthogonal to podman's technical innovations.. but as a lot of infra oss funding ultimately comes from what enterprise policies allow, this a significant anti-competitive adoption driver for podman by rhel/ibm.
We are fine because we use Ubuntu for the generally better GPU support, but I don't have control over the reality of others. I repeatedly get on calls with teams where this comes up, and I feel sorry for the individuals who are stuck paying the time etc. cost for this. Docker for RHEL 8 is literally installing Docker for Centos 7, so this isn't a technology thing, but political: A big corporate abusing their trusted neutral oss infra position to push non-neutral decisions in order to prop up another part of its business. This is all bizarre to operators who barely know Docker to beginwith, and care more about say keeping the energy grid running. ("The only authorized thing will likely not work b/c incompatibilities, and I have to spend 2w getting the old working thing approved?")
And again, none of this is technical. Docker is great but should keep improving, so competitive pressure and alternatives are clearly good... but anti-competitive stuff isn't.
But this is the case for all software on RHEL. It's all older and more stable. Try installing any programming language, database, etc... it's always a version or two older in the official RHEL repositories. RedHat written stuff is always going to be more "current" because they control the repositories, but this isn't any different from how any other package is treated. RHEL has always been slow to update. (Which is why Docker created their own repositories in the first place).
It makes sense though... because RHEL has to check all packages that are part of their repository. Specifically because things in their repository are safe and stable. There is only so much time to go around testing updates to packages.
I'd totally expect a trade-off for old & stable over new and shiny for RHEL: They're getting paid for that kind of neutral & reliable infra behavior. That'd mean RHEL is packaging docker / docker-compose, and maybe podman as an optional alternative. Instead IBM/RHEL is burning IT dept $ and stability by not allowing docker and only sanctioning + actively promoting IT teams use their alternative in unreliable ways.
Might not be relevant if it’s your own machine though but for $BIG_CORP desktop that might be useful.
What's the justification for this... while at the same time allowing the use of containers?
I don't see how that contradicts the use of containers? When we only used docker for the CI we didn't have access to them, because escalating to root from within docker is still pretty easy. But with podman they're running with our privileges (and due to gid/uid mapping, I can run our installer "as root" in the container and see if it sets up everthing correctly - while in reality the process runs under my uid).
Disclaimer: This is in a professional enviroment. At home I despise containers and prefer to have all my service installed and easily updated by the system's package manager.
The big corp I work for has no issue with any of this so I genuinely don’t understand the motivation for this restriction.
Your machine shouldn’t have anything important on it and if it does I’m not sure that root/non-root will protect you from that.
I would guess developers must be using a lot of "workarounds" without management's knowledge.
However, it maybe should not be a requirement. AFAIK nix allows you to install stuff as a normal user (and so does dnf on Fedora, but it depends on the package).
Only thing that annoyed me was when I did some work on an internal web tool (php+js) and had no php-fpm & httpd. But by now I've got a container for that.
I would guess, though, that having to change the prefix and default locations in all tarballs you download must be quite a pain. You can't install DEBs or any other tools either?
We use some external stuff like sqlite, qt,... but that's all versioned in our gits as well (-> easily reproducible builds). Since we sell a commercial product we can't just add random deps anyway. Plus, I think there is very little code that would benefit from being replaced by an external libs.
Relevant tools are on a global package list. We can get stuff added there on a short notice, and with little questions asked. Eg when I needed some lib for a FPGA dev kit that was a matter of minutes.
Uaaah, Firefox on Android is messing up the input again (can only append, not edit). I guess I should just hit that "reply" button then).
Why not just give everyone an isolated space on a server with stuff like systemd-nspawn and let people do whatever they want including using docker inside and if they've screwed up, that's 1 lesson for them and redeploy his environment from a daily backup of the entire space or from a base image quickly.
And please note that getting a CI like env is usually not necessary. I only need that exact env to prebuild things that should run in the CI, e.g. target-specific GCCs. And we can't upgrade the CI because some customers pay good money to have our product run on ancient Linux versions (safety critical industries, once something is certified, you use it a looooong time).
Manjaro is great. Best Linux experience I've had so far. It's good that people recognize it.
I recognize it's a minor pet peeve and would not try to convince anyone though. It's just that I found Manjaro's release policy more to my liking (and unwillingness to fix a broken installation which granted, was easy every time yet took time nonetheless).
Right about then systemd came along (as in, my distro at the time adopted it) and essentially required a service to monitor (some process, tcp port, or file to monitor), but my efficient little script had no service by design, because none was needed and I hate making and running anything that isn't needed to get a given job done. (Plus about that time my company went whole-hog into vmware and a managed service provider and so my infant home-grown lxc containers on colo hardware system never grew past that initial tiny poc)
I laugh so many years later reading "Hey! Check it out! You don't need no steeenkeeeng daemon to run some containers!"
Buildah (`podman buildx`, `buildah bud --arch arm64`) just gained multiarch build support; so also building arm64 containers from the same Dockerfile is easy now. https://github.com/containers/buildah/issues/1590
IDK what BuildKit features should be added to Buildah, too?
Cache mounts are what we rely on for incremental rebuilds.
I used --squash, cleaned up the various apt-gets but boy it made me realise that for every abstraction we get leakage - sometimes serious leakage
Podman doesn’t have that. It spawns containers without the help of a controlling daemon, and can spawn containers both as root and rootless.
Rootless is of course a fairly big deal considering if you run docker containers as root, and runc has a vulnerability, you could potentially escape the container and become root, where a rootless installation would just let you escape to whatever user is running the container.
It's doing a better job for most of the use cases I had recently.
You can run it rootless which makes it a far better use cases for a lot of dev use cases (where you don't need to emulate a lot of network aspects or heavy filesystem usage).
Also you can run it as root without deamon which is much better for a bunch of use cases and eliminates any problems with fuse-overlayfs. This is the best way to e.g. containerize systemd services and similar (and security wise still better as e.g. it still works well with the audit subsystem).
If worse comes to worse you can run it as a deamon in which case it's compatible with docker compose, and has more or less the same problems.
In my experience besides some integration testing use-cases is either already a good docker replacement without running it as a deamon, or the use case shouldn't be handled by docker either way (sure there are always some exceptions).
Lastly I had far less problems with podman and firewalls then with docker but I haven't looked into why that's the case. (But it seems to be related to me using nftables instead of iptables).