Why I joined Heptio
dave.cheney.net
dave.cheney.net
"I remember, thinking back to when I started to use Puppet, and imagining about what it would have been like to have those tools in previous jobs, where automation involved SVN repositories full of perl scripts, and crontab entries lovingly copy pasta’d between machines"
Is this really the problem people have been trying to solve? Is it that people were just using hacked together perl scripts and copying files around, then they moved to Puppet to automate this same process. And now they're moving is to containers?
Because it always seemed like building your own deb or rpm was a reasonably neat, and easily deployable solution to these kinds of problems. Whenever I've had to do it, that's what I've done. I guess there are some limitations to this. But are there other compelling reasons to use containers in general?
With Docker I can take a base userland, install pip, use pip to add libraries, then add my own files, and once I've done that and created an "image" it's just a tarfile: it will unpack everywhere and run.
I can do this without any expertise in pip, rpm, yum, npm, etc., etc. Just follow a couple of recipes once, and if they work then my containers always work. That's a key attraction.
Hate to be pedantic: it will run on an OS that shares the same kernel or newer (assuming backwards compatibility).
This type of "run anywhere" statement caused a lot of confusion for me re: Docker when I first started looking into it.
[edit] backstopped by Docker requiring at least 3.10
It's just not the type of "run anywhere" I was expecting based on the marketing/hype. And the more people kept saying "package once, run anywhere" the more I began to wonder how Docker was able to accomplish that without being a VM. Short answer: they don't.
I'm not saying Docker isn't valuable. It is. I just wish it was documented more readily, without having to sift through jargon, that Docker containers must run on a compatible kernel.
Essentially, I found it confusing how Docker was sold as a lightweight VM.
The difference is the language is more friendly to entry level usage.
You still had to learn something. And they made 1 file-ish which traditionally most package managements weren't.
I have this idea that deb is more file-oriented. Never looked into it properly.
You will get out a control.tar.gz, data.tar.gz, and debian-binary file which is a text document with a version for something in it. (Not the package version. Maybe DPKG api version.) Now, just look at how simple this structure is and marvel at how quickly you can demystify debs!
If you wanted to run "pip install" you could do it in the postinst, a part of control.tar.gz, but this is not the traditional way in a deb. No reason you couldn't do it. Traditionally you should list those pip dependencies as dependencies in control and let debuild or your preferred deb building tool pack it up this way for you, then they can be frozen and will be reproducible through only the package mirrors. (The dependencies are listed as dependencies, and they are also packaged into debs.)
I am looking at my heroku_3.99.4_all.deb that I happened to find first, and data.tar.gz contains a single empty directory usr/local/heroku/ for the place that postinst will dump things into (postinst does a wget to some path at https://cli-assets.heroku.com, and moves a few things around after that.)
IMHO this is a terrible deb, because if heroku goes away it will completely bomb out and fail to install.
(Then again if heroku does go away, what do you need the heroku client for exactly?)
Your suggestion of doing pip install has the same problem, but what's worse? I am a rubyist and I would never question whether you should use bundler to manage your dependencies. Absolutely you should, the package versions in stable distros are atrocious and usually horrifically outdated, and you almost certainly don't want to use anything but a stable distro in production.
So, to recap, yes you can but it is not usually done this way.
Recalling that I am speaking in a Dockerfile context, no it doesn't. The 'pip install' runs once, on my machine when I create the Docker image.
Thank you for the detailed explanation. You have clarified that "just build a deb" is an entirely different exercise.
That's all! :) Thanks for writing back.
(In the Heroku deb example, you can't even guarantee that you are going to get the version back from the server that the deb file says it contains. That's probably by design, because Heroku does not intend to break backwards compatibility with new releases of the client, and wants to force upgrades so that they know that you have the latest version of the client at all times.)
If Heptio can really deliver "at every price point", then this would be amazing. I just want to be able to deploy up a side project for free or very cheap, and seamlessly scale all the way up to thousands of dollars per month.
Launching a side project on Heroku and AWS seems to cost a minimum of $69.00 per month. I typically need at least a hobby web and worker dyno ($7 each), Redis Micro for $5, and a Standard 0 Postgres database for $50. (Backups and continuous protection are important.)
Actually, I just looked at AWS, and it's about the same price to get started. You could probably start with RDS instance on db.t2.small, for $26. And run your server and workers on a t2.medium server for $34.
Maybe the only way to start from $0 is to use AWS Lambda and DynamoDB. But it would be really nice if Kubernetes was as easy to use as Heroku, and you could also start from $0 with a production-ready service.
I disagree with your assessment of "minimum". the $7/mo is more than enough for starting out; enough to validate your idea and serve some customers. If you need more services, you can get a digital ocean box for $10, with less "seamless" scaling. or, move to heroku once you have validation.
Rather than work around the limitations of Debian-like package management, I find it easier, and more worth while to just start using more advanced package managers like Guix or Nix.
"modulo" in this sense does make sense, because it's abstracting that part out and mapping it to a equivalence relation.
Also, there are (and will continute to be) plenty of companies/projects that can't use public clouds for a variety of security, compliance, performance and business inertia concerns ('concerns', not necessarily 'reasons').
Edit: typo
See, can spin it either way :-)