Manager: "Fine, then we will ship your machine"
And thus docker was born.
Manager: "Fine, then we will ship your machine"
And thus docker was born.
Dev would get the thing working on their machine configured for a customer. We'd take their machine and put it in the server room, and use it as the server for that customer. Dev would get a new machine.
Yes, I know it's stupid. But if it's stupid and it works, it isn't stupid.
DLL Hell was real. Spending days trying to get the exact combination of runtimes and DLL's that made the thing spring into life wasn't fun, especially with a customer waiting and management breathing down our necks. This became the easiest option. We started speccing dev machines with half an eye on "this might end up in the server room".
nm -D is not hard. Debugging missing symbols isn't even that hard.
Biggest damn challenge I see is we've completely dropped the ball at educating people about how linkers, loaders, and dynamic library symbols work.
...and so people understand something here... I've transplanted/frankensteined userspaces that were completely hosed/disjoint into working before. In fact, every now and again I do it just to remain in practice. It's actually gotten to the point I don't even worry about dependency hell anymore. I just find the right version of the library/source and expand my archives just in case I need it down the road.
I had one of these years ago where QA had an issue that I couldn't reproduce.
I walked over to his desk, watched the repro and realized that he was someone who clicked to open a dropdown and then clicked again to select, while I would hold the mouse button down and then let up to select.
I didn't know the latter was possible, but I also don't know which of these I do.
I'm guessing it may be one of these things where it only starts to make sense after something has matured enough to warrant being replicated en-masse in a data-center environment.
Then again, I tend to live in a world where I'm testing the ever-loving crap out of everything; and all that instrumentation has to go somewhere!
Simple answer: with configuration. In Docker Compose or Kubernetes. Less often in Mesos. Maybe I want to run it on a fleet of VMs, maybe on bare metal.
> and you're doing what you'd be doing on VM anyway
Right. But with containers I can have different apps using different dependency versions. Some things use nginx, some use some other web server. Some things run with node 12, some with node 16, some things use MySQL, some other things need Postgres. This is so easy with containers.
And like, that's just minimum: I generally care deeply about the filesystem that PostgreSQL is running on and I will want to ensure the transaction log is on a different disk than the data files... and I now have to configure all of this stuff multiple times.
At some point I am going to have to edit the configuration file for PostgreSQL... is it inside of the container? Did I have to manually map it from the outside of the container?
The way you access PostgreSQL locally--maybe if you find yourself ready to add a copy of pgbouncer, but also just to run the admin tools--is via unix domain sockets. Are those correctly mapped outside of the container, and where did I put them?
I honestly don't get it for something like PostgreSQL. I even use containers, but I can only see downsides for this particular use case. You know how easy it is to run PostgreSQL in some reasonable default confirmation? It is effectively 0 commands as you just install it and your distribution generally already has it running.
And how is that different from running directly on the OS?
this makes sense when you're trying to be deployable universally, but it increases the amount of onboarding that someone needs to receive before they're proficient with the container system; onboarding they may have not actually needed to get the software working and well understood, simply 'docker overhead'.
from personal experience : i'm a long time old linux person, the insistence on going 'all in' on Docker (or whatever) just to run a python script that has two or three common shared dependencies gains me nothing but the hassle of now having to maintain and understand a container system.
if you're shipping truly fragile software that is dependent on version 1.29382828 rather than version 1.29382827 then I understand the benefits gained, but just to containerize something very simple in order to follow industry trends is obnoxious, increasingly common, and seemingly has soured a lot of people on a good idea.
p.s. : I can also understand the idea of containerizing very simple things as parts of a larger mechanism; I just don't get it with the promise that it'll reduce end-user complexity, it isn't that simple.
1. I can start it and stop it automatically with Compose. That's a big increase in ergonomics.
2. I don't have to write my own scripts to set up and tear down the database in the dev environment. The Postgres image + Compose does all that for me.
3. Contributors who don't know much about Unix or database admin stuff (data analysts learning Python & data scientists contributing to the code) don't have to install or mess with anything. It works on everyone's machine the same way.
Volume and port mapping are basically trivial concepts anyway. There's been zero downside for me in using it, even though I personally have the skills and knowledge to not "need" it. Why would I go without it? It saves me time and effort that could be significantly better used elsewhere.
However, we live in the world where the choice we have for new hires is: a) teach them all of those OS fundamentals, b) give them Docker.
This doesn’t mean we shouldn’t strive to teach said new developers all those underlying concepts, but when we talk about training juniors and have them contribute relatively quickly, it’s a much smaller surface area to bite through.
That's what I said.
Furthermore, let's be pragmatic. The operations team needs to know, yes. We want new hires to contribute and feel productive. They'll naturally learn while working on software. A junior person can contribute almost immediately with a limited surface knowledge.
Less friction: no need to understand systemd/upstart/rc what have you, /etc, /opt, /usr, mount, umount, differences between various distros, build-essentials, Development Tools, apt, dpkg, rpm, dnf, ssh keys, ...
The right time will come but give them an easy way in. Containers provide exactly that.
All they need to know: it's somewhat isolated so under normal circumstances whatever you do in the container doesn't affect the host, how to expose ports, basics of getting your dependencies in, volumes, basic understanding of container networks - for things to talk to each other they need to be in the same network. Enough to start.
- Where is the data stored? Your compose file will tell you.
- Is the configuration file in that container or mapped? File tells you. (If you didn't map it, it's container-local)
- How do you get to it with admin tools? If you mapped those sockets outside, file tells you.
Solution: config file for the config files.
The amount of work you are complaining about is, objectively, trivial. It's no harder than learning how to deploy on another distro or operating system, except in this case, your new knowledge is OS-agnostic.
As a user, this is why I love Docker. The configuration is explicit and contained, and it's well documented which directories and ports are in play.
I don't need to remember to tweak that one config file in /etc/ which I can never remember where is. Either it's a flag or it's a file in a directory I map explicitly. And where does _this_ program store its data? Don't need to remember, data dir is mapped explicitly.
That said I haven't tried to use PostgreSQL myself directly, just tools that uses it like Gitea.
As someone else said, docker isn't about you, it's about everyone else. The extra complexity up front is so worth it for the rest of the team.
You bind a mount to where the data is onto the external system. That way only the important data is exposed. It's very clean, although it requires you understand docker a bit to know to do this. But for things like postgres and similar sprawling software that basically assume they're a cornerstone of your entire application and spread out as though this was the case, it's actually a very neat way of using them a bit without having them take over your machine.
This is something that can be useful for a developer too. Like my search engine software assumes it owns the hardware it runs on. It assumes you've set up specific hard drives for the index and the crawl data. But sometimes you just wanna fiddle with it in development, and then it can live in a pretend world inside docker where it owns the entire "machine", and in reality it's actually just bound to some dir ~/junk/search-data.
Although I guess an important point is that using docker requires you to understand the system better than not using docker, since you both need to understand the software you run inside the container, as well as docker. It's sometimes used as though you don't need to understand what the container does. This is a footgun that would make C++ blush.
Postgres is a poor example; it's far far easier to do `apt-get install postgresql` than to run it from a container.
The latter needs a container set up with correct ports, plus it needs to be pointed at the correct data directory before it is as usable as the former.
It was refreshing being able to say: this works perfectly from a base Ubuntu 20.04 container, launching this and that commands to install and run the software. Anything that diverges from there, you're on uncharted territory and/or is a problem in your system, not a defect in the software.
Instead of keeping some old machine with old dependencies because project is on life support and doing update to latest version of language/tools is unviable, you can just keep a docker container for it while machine it is running on is up to date.
It doesn't solve the problem of application, but it is no longer ops problem that app is old.
You can also do that partially, like keep old PHP version running with just fcgi socket exposd in a container, while rest of the app lives on "normal" VM (or other container, if that's what you want).
Granted, a lot of THAT benefits from the fact that the container actually has to be a documented flavor of good solid *NIX, unlike the "Windows Image X" which is usually just a big "whatever" of old iron, management fad spoor, and seven different kinds of antivirus software. Hell, last time I did this, IT couldn't even find me any kind of description of what the "official" windows image actually was. So no fault of Windows there, it's just that, as the big bus everyone rides in, it gets all the goop from everyone wrenching on it.
The last time I counted up all the different persistent, security-related agents running on my corporate Windows box, it was actually 14.
My theory is eventually things will come full circle and there will entire app ecosystems running in a unikernel executable which is running on physical hardware and the hardware itself is segmented. I can imagine companies like netflix and cloudflare having racks of servers with 1TB ram with no disk, booting a unikernel from the network for maximum throughput and latency and then everyone in tech follows along.
What people want is the ability to do all these things with a single command that they put inside a bash script.
Another solution would find a senior developer who can manage his dependencies properly, but I can't afford him. However, I can pay a junior ML developer wage and just throw what he develops in a docker image instead of wasting time and money on doing things the hard way for a prototype.
As engineers, we're supposed to be rational people, but unfortunately, we tend to forget that time and money are often happen to be the most important constraints we have to work around.
For toy/evaluation use, it's hard to beat "tweak compose file, docker-compose up". You now have a golden source of truth for how the entire application and all of its moving parts were set up.
Going back to deploying applications on bare metal feels positively medieval by comparison. From the admin side, configuration drift is effectively not a thing anymore. That's huge.
Provided all that, I don’t get why “devops” is even a high-paying job at most places except few really web scale ones. They are basically former sysadmins who reject anything except plugging colored squares into square sockets.
I can run that container in CI, in tests (with stuff like testcontainers), on my laptop, in production, on Kubernetes, in Mesos, Lambda, whatever.
Instrumentation usually ends up in Prometheus and Jaeger.
One company came close, they had someone dedicated to maintaining the dev tooling including the containers. It still didn't "just work" but they handled the troubleshooting, and the fix made it into the repo for the container so it wasn't just an unwritten adhoc fix I needed to remember. Close enough.
I totally get the concept, I wish it worked so seamlessly that the promise was realized. But since I usually work with tooling I am familiar with, there is no point for me personally. I can stand up a local env faster than I can troubleshoot a broken container. Since I am not interested in the infrastructure and just want to get to work, that is what I do.
[0] We don’t have concise terms for the types of distributed, complex systems which have become standard.
This is the most concise and exact definition I've heard lately.
My only problem with the consequences of this approach is that the amount of overhead for performing the same operations is astonishing.
Messages pass on a network instead of stsying in RAM, CPU context switches every other ms to handle IO for what could have been a simple function call, etc.
It's amazing when it's needed, but seeing this approach becoming the standard for a big part of the industry almost turns it into an environmental problem. How much energy is wasted juggling bits around?
However, I was specifically answering to the different case highlighted by the parent comment: how a potentially cohesive application is cut along some of its internal APIs, and some functionality is allocated to different processes living in different containers.
It is an extreme point in the continuum "single thread" -> "multi thread" -> "multiple processes" -> "fully distributed". In that continuum scalability increases, while efficiency progressively decreases.
Cornering oneself to a specific point in the design space is problematic, and for some cases has direct implications on how many resources are wasted.
A very didactic experience is, for example, running a simple local application under a microarchitecture profiler, such as Intel Vtune. It is not uncommon seeing that even straightforward C/C++ programs use a core resources less than 10%.
What I am reflecting about is that the choice of fragmenting that program among tens of systems (maybe in a scripting langiage) should be conscious, and done after encountering performance or scalability bottlenecks.
How much of the resulting total system workload would be useful work?
The quantity of potentially wasted resources is astonishing if you think about it.
That's just one example, though. It also lets you standardize the dev environment (work directly out of the build container directly, or use your host OS and run unit tests in a local build), and it allows for easy, standardized testing within the same environment.
How so? A container is the same architecture as the host, how does that help with cross compiling?
The power of docker compose for the development process is unparalleled. Being able to declare my entire local environment in a single file in a consistent way, create and destroy it with a simple command, and share it with my co-workers eliminates so many issues.
Then being able to package my application in way that is repeatable and predictable is great. Then that same artifact is run across multiple environments spanning data centers around the world and on other developers’ machines via their compose files if needed.
We test the crap out of all of this as well. Docker helps there to generate ephemeral environments to run the bulk of the integration tests against every time we push to git.
Its just so damn powerful compared to what we had before…
There's a lot of code I've written that's not deployable any more because the versions of the dependencies it was using don't exist on my OS any more, and are too old to install.
Also, from my misadventures with Ubuntu VMs I always ran into the problem of them breaking over time.
Exact reproducibility is nice for two scenarios: 1) academic research, and 2) very large-scale applications and deployments. For regular people writing boring small web apps, choosing a stable base image and pinning dependencies is good enough.
Consider also that your preferred programming language will also very likely not provide particularly reproducible package builds.
This just means you don't.
Try using `sudo apt install git=1:2.39.2-1ubuntu1`
That pins it to a particular version so that it should be reproducible.
I've never looked seriously into it, but my feeling is that distros will delete old versions as newer ones are uploaded: When I run "apt-cache policy git" in my Ubuntu, I only see a couple versions available to install, often other packages show only a single one (so, the latest).
However, much in the same way that if you actually take your build system seriously you'll store your application dependencies in a local proxy, you can run a mirror or proxy to hold these historical packages too.
Take a look at something like apt cacher, however it is a proxy cache so you can reproduce builds using the exact same package versions but if upstream delete old packages, and you want to roll back to one you haven't previously downloaded, then you are out of luck.
I was cobbling together build scripts for a mail system for Raspi/Armbian the last couple weeks. Very similar packaging stuff, but the number of little subtle differences in install/postinstall/prerm/postrm scripts took a generous level of spackling over to get just right.
Hell, once I get everything nailed down, I'm writing test frameworks for my bloody build scripts if you can believe it.
So...uh... Last week?
The freedom given in ability to limit your thinking space by reproducibility should not be under stated.
For almost all non-Google use cases, repeatable is generally good enough.
But then, we have none of those problems at work and people still believe it's a panacea. I don't know what people believe they gain by it.
In general, without some kind of qualification, deploying an image is way worse than a flat binary, or a tar from some interpreted language.
Yes you do need to configure some environment variables like in almost any other form of deployment.
Messing around with tar files or flat binaries doesn't work for me. I tried that. I build a python app into an executable with Nuitka but it generates a folder anyway. Ok throw it in a tar file. Nope, doesn't run because the glibc version is too new on the developer machine. Nice try. I have to build it on the target machine. Amazing.
I got stuck creating VMs for testers for a distributed workflow (several services, several tools), and keeping those VMs up to date and working was a PITA, but far less painful than dealing with them filing bug reports based on thinking they were doing X when they were doing Y.
I ended up creating a workflow that felt like layers, so when Docker came along I didn't need a salespitch. I could just do what I'd done with a script instead of a runbook. Where do I sign up?
Containers are packaging for executables.
MacOS makes static linking quite difficult.
And Windows is infamous for "DLL Hell."
Even windows reinvented the exe with msi's, and python has dists and wheels, and *nix/bsd has jails and vms have existed forever, and java has the jvm with jars and previously, applets.
Of course, none of those are actually as user friendly nor cross platform compatible as docker. Heck even docker isn't perfect and will leak abstractions once you start hitting syscalls (network, filesystem, etc), or if you're like 99.9...% of developers who don't host their own artifactory and cryptographically tag every step in their dockerfile
Show me an .exe that can self deploy an elk stack across windows, Darwin, and Linux with minimal system environment issues
Honestly, getting some desktop software in the form of one of those is wonderfully refreshing. Just take an AppImage file or whatever, make it executable and run it.
Sometimes you don't care about the benefits of shared libraries and want something to just run regardless of what's going on with your distro, for which these technologies are great. Ideally, in addition to the proper package manager approach for those that value different things at different times.
But for shipping things like WebApps, OCI containers and Docker/Podman is amazing.
There is a long, bloody history of packaging executables with only their direct dependencies.
Docker doesn't fix that. In fact, sometimes it makes it even harder for my testers to keep track of what is happening where. A shockingly high number aren't even able to visualize the logical boundaries between physically cohabiting, but logically partitioned systems.
It's too many damned layers.
And developers do awful things with docker containers to the point I end up having to rip apart multiple projects Dockerfiles to rebuild those, but with different values to work around the fact that oh .. look at that, this one spins up a postgres database that wants to bind to ports we're already using. Looks like I get to go dig into the base image to change that port and rebuild...
Again, not saying it doesn't work. As someone who put in the work to become a generalist though, I put in an asinine amount of time working around other peoples things that quote "make things easier".
Maybe for the person writing it. Certainly not for someone learning from it or having to analyze it.
The whole app can be started with a single command and it works on most Linux distributions. I can't imagine wasting the time of users or newbie developers demanding that they install all these things separately and with no easy way to clean it up if they want to undo everything.
The OS, the development environment and the application (both code and live objects) where one and the same thing. To ship an "app" you would export the image and the user would load it into their Smalltalk VM.
Turns out they did ship a 1:1 image of his machine.
After he spent 3 months building a web app, I asked him how he wanted to deploy it.
Perfectly straight face he said we would take his developer machine to a local data center and plug it in. We could then buy him a new developer machine. It went downhill from there.
I ended up writing the application from scratch and deploying it that same evening.
Owner hired a lot Of strange people.