Rails on Docker
fly.io
fly.io
It isn't explicitly explained, but the reason why it must be in this command and not separated out is because each command in a dockerfile creates a new "layer". Removing the files in another command will work, but it does nothing to decrease the overall image size: as far as whatever filesystem driver you're using is concerned, deleting files from earlier layers is just masking them, whereas deleting them before creating the layer prevents them from ever actually being stored.
FROM the_source_image as builder
RUN build.sh
FROM the_source_image
COPY --from=builder /app/artifacts /app/
CMD ....
i'm not sure if you can really call it the new best practice though, its been the default for ... a very long time at this point.For example, say I have my own Ubuntu image that is based on one of the official ones, but adds a bit of common configuration or tools and so on, on which I then build my own Java image using the package manager (not unlike what Bitnami do with their minideb, on which they then base their PostgreSQL and most other container images).
So I might have something like the following in the Ubuntu image Dockerfile:
RUN apt-get update && apt-get install -y \
curl wget \
net-tools inetutils-ping dnsutils \
supervisor \
&& apt-get clean && rm -rf /var/lib/apt/lists /var/cache/apt/*
But then, if I want to install additional software, I need to fetch the package list anew downstream: FROM my-own-repo/ubuntu
RUN apt-get update && apt-get install -y \
openjdk-17-jdk-headless \
&& apt-get clean && rm -rf /var/lib/apt/lists /var/cache/apt/*
As opposed to being able to just leave the cache files in the previous layers/images, then remove them in a later layer and just do something like: docker build -t my_optimized_java_image -f java.Dockerfile --purge-deleted-files .
or maybe
docker build -t my_regular_java_image -f java.Dockerfile .
purge-deleted-files -t my_regular_java_image -o my_optimized_java_image
Which would then work backwards from the last layer and create copies of all of the layers where files have been removed/masked (in the later layers) to use instead of the originals. Thus if I'd have 10 different images that need to use apt to install stuff while building them, I could leave the cache in my own Ubuntu image and then just remove it for whatever I want to consider the "final" images that I'll ship, which would then alter the contents of the included layers to purge deleted files.There's little reason why these optimized layers couldn't be shared across all 10 of those "final" images either: "Hey, there's these optimized Ubuntu image layers without the package caches, so we'll use it for our .NET, Java, Node and other images" as opposed to --squash which would put everything in a single large layer, thus removing the benefits from the shared layers of the base Ubuntu image and so on.
Who knows, maybe someone will write a tool like that some day.
https://docs.docker.com/develop/develop-images/dockerfile_be...
> Official Debian and Ubuntu images automatically run apt-get clean, so explicit invocation is not required.
RUN apt-get update -qq && \
apt-get install -y build-essential libvips && \
apt-get clean && \
rm -rf /var/lib/apt/lists/* /usr/share/doc /usr/share/man
I'd do: RUN apt-get update -qq && \
apt-get install -y build-essential libvips && \
rm -rf /var/lib/apt/lists/*rm -rf isn't configuration management, it's system entropy increasing leaving users scrambling to reinstall the world.
If people don't want docs, then distro vendors should package them separately and make them recommended packages.
"I recommend avoiding Ubuntu because it violates the principle of dev - prod parity."
What exactly is the problem? Would you please like to provide sources to this issue / some explanation / any helpful content? THANKS!
It's good to have a slim script but also to have an unslim when you need manpages or locales.
It is really weird looking at Dockerfiles for the first time seeing all of the `&% \` bash commands chained together.
FROM foo AS builder
.. build steps
FROM foo
COPY --from=builder generated-file target
(I hope I got that right; on a phone and been a while since I did this from scratch, but you get the overall point)Say `apt install build-essential libvips` from the OP, it's not obvious to me what files libvps is adding. I suppose there's probably an incantation for that? What about something that installs a binary? Seems like a pain to chase down everything that's arbitrarily touched by an apt install, am I missing some tooling that would tame that pain?
It is of course possible to do for a single project with a bit of effort: build each stage with a remote OCI cache source, push the cash there after. But... that sucks.
What you want is the `max` cache type in buildkit[1]. Except... not much supports that yet. The native S3 cache would also be good once it stabalizes.
It worked wonders for us, on a cache hit the build time is reduced from 10 to 1,5 minutes.
https://www.docker.com/blog/introduction-to-heredocs-in-dock...
This is great and really makes things way simpler. Thanks!
Many other dockerfile build tools still don't support it, e.g. buildah (see https://github.com/containers/buildah/issues/3474)
Useful now if you have control over the environment your images are being built in, but I'm excited to the future where it's commonplace!
You can put stuff in a script file and just run that script too.
Tracking for this: https://github.com/docker/buildx/issues/1104
this has been a serious pain in my side for a while, both for my own debugging and for telling people I try to help "you're gonna have to start over and do X or this will take hours longer".
If you want to reduce layer/image size, then I think [1] "multi-stage builds" is a good option
You can use HEREDOCS to combo together commands that make sense in a layer, ensure your layers are ordered such that the more-frequently changing ones are further on in your Dockerfile when possible (this will also speed up your builds, ensuring as many caches as possible are more likely to be valid), and use mutli-stage builds on top of that to really pare it down to the bare necessities.
Is it though? From the post:
RUN <<EOF
apt-get update
apt-get upgrade -y
apt-get install -y ...
EOF
It may be due to my ninja level abilities to dodge learning more advanced shell mastery for decades, but to me it looks haphazard and error prone. Are the line breaks semantic, or is it all a multiline string? Is EOF a special end-of-file token, or a variable, if so what’s it’s type? Where is it documented? Is the first EOF sent to stdin, if so why is that needed? What is the second EOF doing? I can usually pick up a new imperative language quickly, but I still feel like an idiot when looking at shell.
<<XYZ
...
XYZ
syntax for multi-line strings is worth learning since it is used in shell, ruby, php, and others. See https://en.m.wikipedia.org/wiki/Here_document . You get to pick the "EOF" delimiter.> > Nice syntax
> Is it though?
Before the heredoc syntax was added, the usual approach was to use a backslash at the end of each line, creating a line continuation. This has several issues: The backslash swallows the newline, so one must also insert a semicolon* to mark the end of each command. Forgetting the semicolon leads to weird errors. Also, while Docker supports line continuations interspersed with comments, sh doesn't, so if such a command contains comments it can't be copied into sh.
The new heredoc syntax doesn't have any of these issues. I think it is infinitely better :)
(There is also JSON-style syntax, but it requires all backslashes to be doubled, and is less popular.)
*In practice "&&" is normally used rather than ";" in order to stop the build if any command fails (otherwise sh only propagates the exit status of the last command). This actually leads to a small footgun with the heredoc syntax: it allows the programmer to use just a newline, which is equivalent to a semicolon and means the exit status will be ignored for all but the last command. The programmer must remember to insert "&&" after each command, or use `set -e` at the start of the RUN command, or use `SHELL ["/bin/sh", "-e", "-c"]` at the top of the Dockerfile. But this footgun is due to sh's error handling quirks, not the heredoc syntax itself.
> Are the line breaks semantic, or is it all a multiline string?
The line breaks are preserved ("what you see is what you get").
> Is EOF a special end-of-file token
You can choose which token to use (EOF is a common convention, but any token can be used). The text right after the "<<" indicates which token you've chosen, and the heredoc is terminated by the first line that contains just that token.
This allows you to easily create a heredoc containing other heredocs. Can you think of any other quoting syntax that allows that? (Lisp's quote form comes to mind.)
> Where is it documented?
The introduction blog post has already been linked. The reference documentation (https://docs.docker.com/engine/reference/builder/, https://github.com/moby/buildkit/blob/master/frontend/docker...) explains the syntax using examples. It doesn't have a formal specification; unfortunately this is a wider problem with the Dockerfile syntax (see https://supercontainers.github.io/containers-wg/ideas/docker...). Instead, the reference links to the sh syntax specification (https://pubs.opengroup.org/onlinepubs/9699919799/utilities/V...), on which the Dockerfile heredoc syntax is based.
> This actually leads to a small footgun with the heredoc syntax: it allows the programmer to use just a newline, which is equivalent to a semicolon and means the exit status will be ignored for all but the last command.
This sounds like a medium-large caliber footgun to me, and while I don’t expect Docker to fix sh, it could perhaps either set sane defaults or decouple commands from creating layers? Or why not simply support decent lists of commands if this is such a common use case?
> This allows you to easily create a heredoc containing other heredocs.
Hmm, what’s the use-case for that? The only effect for the programmer would be to change the escape sequence, no?
Ha ha, I guess footgun sizes are all relative. The quirky error handling of sh is "well-known" (usually one of the first pieces of advice given to improve safety is to insert `set -e` at the top of every shell script, which mostly fixes this issue). So I don't think of Dockerfile heredocs themselves as a large footgun, but rather as a small footgun that arises out of the small interaction between heredocs and the large-but-well-known error handling footgun.
I don't know why Docker doesn't use `set -e` by default. I suppose one reason is for consistency -- if you have shell commands spread across both a Dockerfile and standalone scripts, it could be very confusing if they behaved differently because the Dockerfile uses different defaults.
I also don't know why the commands are coupled to the layers. Maybe because in the simple cases, that is the best mapping; and in the very complex cases, the commands would be moved to a standalone script; so there are fewer cases where a complex command needs to be inlined into the Dockerfile in a way that produces a single layer.
It would be really nice if the Dockerfile gave more control over layers. For example, currently if you use `COPY` to import files into the image and then you use `RUN to you modify them (e.g. to change the ownership / permissions / timestamps), it would needlessly increase the image size; the only way to avoid this is to perform those changes during the COPY, for example using `COPY --chown`; but COPY has very limited options (namely: chown, and also chmod but that is relatively recent).
Regarding native support for lists of commands, I don't really see much value since sh already supports lists (you "just" need to correctly choose between "&&" and ";"/newline).
> > This allows you to easily create a heredoc containing other heredocs.
> Hmm, what’s the use-case for that? The only effect for the programmer would be to change the escape sequence, no?
It can be useful to embed entire files within a script (e.g. when writing a script that pre-populates a directory with some small files). With most quoting schemes, you'd have to escape special characters that appear in those files. But with heredocs, you just have to pick a unique token and then you can include the files verbatim.
(Picking a token that doesn't appear as a line within the files can be a little tricky, but in many cases it's not a problem; for example if the files to be included are trustworthy, it should be enough to use a token that includes the script's name. On the other hand if the data is untrusted, you'd have to generate an unguessable nonce using a CSPRNG. But at that point it's easier to base64-encode the data first, in which case the token can be any string which never appears in the output of base64, for example ".".)
1) a layer is essentially a docker image in itself and
2) a layer is as static as the image itself
3) Docker images ship with all layers
Thanks jchw!
Example: in the last few years docker compose files have gone from version 2 to version 3 (missing tons of great v2 features in the name of simplification) to the newest, unnamed unnumbered version which is a merging of versions 2 and 3.
Externally managed caches don't have a lifecycle controlled or invalidated by changes in Dockerfiles, or by docker cache-purging commands like "system prune".
That means you have to keep track of those caches yourself, which can be a pain in complex, multi-contributor environments with many layers and many builds using the same caches (intentionally or by mistake).
Caveat: it doesn't work on Fly.io. They seem to be having some issue with OCI manifests: https://github.com/containers/skopeo/issues/1881 . They're also having issues with new docker versions pushing from CI: https://community.fly.io/t/deploying-to-fly-via-github-actio... ... the timing of this post seems weird.
FWIW the article says
> create a Docker image, also known as an OCI image
I don't think this is quite right. From my investigation, Docker and OCI images are basically content addressed trees, starting with a root manifest that points to other files and their hashes (root -> images -> layers -> layer configs + files). The OCI manifests and configs are separate to Docker manifests and configs and basically Docker will support both side by side.
But I've been really enjoying it as a way of just telling a PaaS "hey here's the compiler/runtime my code needs to run", and then mostly not having to worry about it from there. It means these services don't have to have a single list of blessed languages, while pretty much keeping the same PaaS user experience, which is great
Have you ever worked somewhere where in order to run the code locally on your machine, it's a blend of "install RabbitMQ on your machine, connect to MS-SQL in dev, use a local Redis, connect to Cassandra in dev, run 1 service that this service routes through/calls to locally, but then that service will call to 2 services in dev", etc?
And 50% of the time it actually works. The other 50% of the time, it doesn't, and I spend a week talking to other engineers and ops trying to figure out why.
Not a jab at Docker, btw.
This way each developer can/should tear down and rebuild their entire development environment on a semi-frequent basis. That way you know it works way more than 50% of the time. (Stretch goal: put this in CI.... :)
The other side benefit: if you reset/restore your development database frequently, you are incentived/forced to add any necessary dev/test data into the "rake generate_development_data", which benefits all team members. (I have thought about a "generate_development_data.local.rb" approach where each dev could extend the generation if they shouldn't commit those upstream, but haven't done that by anymeans....)
IMO Docker for local dev is most beneficial for python where local installs are so all over the place.
If anyone is curious about the details, I simply reused the existing `bin/dev` script set up by Rails by adding this to `Procfile.dev`:
docker: docker compose -f docker-compose.dev.yml up
The only issue is that foreman (the gem used by `bin/dev` to start multiple processes) doesn't have a way to mark one process as depending on another, so this relies on Docker starting the Postgres and Redis containers fast enough so that they're up and running for Rails and Sidekiq. In practice it means that we need to run the `docker compose` manually the first time (and I suppose every time we'll update to new versions) so that Docker downloads the images and caches them locally.For your issue could you handle bringing up your docker-compose 'manually' in bin/dev? Maybe conditionally by checking if the image exists locally with `docker images`. Then tear it down and run foreman after it completes?
If someone updates nodejs, just pull in the latest dockerfile. Someone updates to a new postgres version? Same thing. It's so much better than managing native dependencies.
I don't totally understand the maintenance story. On heroku with buildpacks, I don't need to worry about OS-level security patches, patches to anything that was included in the base stack provided by the PaaS, they are responsible for. Which I consider part of the value proposition.
When I switch to using "Docker" containers to specify my environment to the PaaS... if there is a security patch to the OS that I loaded in one of my layers... it's up to me to notice and update my container? Or what? How is this actually handled in practice?
(And I agree with someone else in this thread that this fly.io series of essays on docker is _great_, it's helping me understand docker better whether or not I use fly.io!)
But it sounds like for patch/update management purposes, now I need to add something like that in? Another platform/host to maintain/manage, at possibly additional price, and then we add dealing with the specific mechanisms for scanning/updating too...
Bah. It remains mystifying to me that the current PaaS docker-based "best practices" involve _quite a bit_ more management than heroku. I pay for a PaaS hoping to not do this management! It seems odd to me that the market does not any longer seem to be about providing this service.
It is odd, but the general direction of the market (at least at the top) is becoming less opinionated, which means you need to bring your own. Not sure I like that myself.
Re: security updates. It's not handled. There are companies that will scan your infrastructure to figure out what's in your containers and find out of date OS base images.
One thing I'm currently experimenting with is a product for people who would like an experience slightly closer to traditional Linux. The gist is, e.g.
include "#!gradle -q printConveyorConfig" // Import Java server
app {
deploy.to = "vm-frontend-{lhr,nyc}{1-10}.somecloud.com"
linux {
services.server {
include "/stdlib/linux/service.conf"
}
debian.control.Depends = "postgres (>= 14)"
}
}
Then you run "conveyor push" from any OS and it downloads a Linux JVM, minimizes it for your app, bundles it with your JARs, produces a DEB from that, sftps it to the server, installs it using apt, that package integrates the server with systemd for startup/shutdown and dynamic users, it healthchecks the server to ensure it starts up and it can do rolling upgrades/downgrades. And of course the same for Go or NodeJS or whatever other server frameworks you like. So the idea is that if you use a portable runtime you can develop locally on macOS or Windows and not have to deal with Linux VMs, but deployment is transparent.SystemD can also run containers and "portable services", as well as using cgroups for isolation and sandboxing, so there's no need to use debs specifically. It just means that you can depend on stuff that will get upgraded as part of whole OS system upgrades.
We use this to maintain our own servers and it's quite nice. One of the things that always bugged me about Linux is that deploying servers to it always looks like either "sftp a tarball and then wire things up yourself using vim", or Docker which involves registries and daemons and isolated OS images, but there's nothing in between.
Not sure whether it's worth releasing though. Docker seems so dominant.
So... what do actual real people do in practice? I am very confused what people are actually doing here.
LOTS of people seem to have moved to this kind of docker-based deploy for PaaS. They can't all just be... ignoring security patches?
I am very confused that nobody else seems to think the "handle patches story" is a big barrier to moving from heroku-style to docker-style... that everyone else just moves to docker-style? What are they actually doing to deal with patches?
I admit I don't really understand your solution -- I am not a sysadmin! This is why we deploy to heroku and pay them to take care of it! It is confusing to me that none of the other PaaS competitors to heroku -- including fly.io -- seem to think this is something their customers might want...
The Dockerfile has in it:
> ARG RUBY_VERSION=3.2.0
> FROM ruby:$RUBY_VERSION
Which it says gets us "gets us a Linux distribution running Ruby 3.2." (Ubuntu?). The version of that "linux distribution" isn't mentioned in the Dockerfile. How do I "bump" it if there is a security patch effecting it?
Then it has:
> RUN apt-get update -qq && apt-get install -y build-essential libvips && \ ...
Then I just run `fly launch` to launch that Dockerfile on fly.io. How do I "bump" the versions of those apt-get dependencies should they need them?
If it's not, you just need to trigger a build every so often. Maybe this could be a feature PaaS offers in the future.
There are a few problems with this, as noted elsewhere in the thread:
1. You have to do an uncached rebuild and repush all of your images, using some ad-hoc company specific process. There's nothing that can do this for you at the end points or service levels, because Docker images are meant to be immutable after build and don't come with the scripts or inputs used to build them.
2. The default is to use caching, so devs may not notice that they didn't refresh their base OS for a while.
3. You don't get notified when updates are available or applied. There's nothing like the unattended-upgrades package that comes with Debian normally, which will apply upgrades and then tell you what happened.
4. Because of (3) the latency is very high. There is no story (other than third party scanners) for getting notified about urgent upgrades. If there's another zero day in OpenSSL then with a standard Linux install you'll get patched as soon as a new package is released and your machines update, so pretty quick (a day or so). With Docker images, it'll get patched on an app-by-app basis if and when people get around to doing an uncached rebuild and repush of the image.
5. Kernel upgrades are a whole can of eels in container-world. People like to think of the container as being a self-contained OS but it isn't. There is a largely unstated and untested assumption that any Linux distro user space can run on any kernel version or configuration, regardless of whether the OS originally shipped in that configuration, and everything will just automatically do something sensible. Mostly this assumption is OK because servers are very simple, but it's not actually guaranteed by anything. A lot of people misunderstand the "stable Linux syscall interface" guarantees and what that means.
It's for reasons like this that I prefer the slightly older way of running real binaries that are exposed to the OS and which use OS specific packages. I configured unattended-upgrades and use LTS versions of the OS, so that security patches just stream in without me doing anything. There are a few downsides to this too:
1. You have to either restart your servers from time to time to force security patches to actually get reloaded into memory, or use the needrestart Debian package - however that only works if you're using package metadata properly.
2. You do need to understand at least a bit of Linux sysadmin. Enough to know how to ssh in as root, use apt-get and so on.
3. The tooling story is poor, hence my musings above about demand for something better. Without Docker, today you're going to be manually copying files to the server, having to learn systemd and how to start/stop/enable services, how to restart them on upgrades etc. That's why I've written something that does it all for you.
They can and they are, IME. I've seen images that literally run some version of Alpine Linux from 3 years ago.
Right. I didn't explain it very well then.
Basically it's a heroku-ish solution but as a tool rather than a service, and where you build locally instead of pushing source code to some remote cloud. You say "here's my build system, go push to these plain Linux VMs". Now, someone needs at least a bit of sysadmin knowledge - you need to know how to obtain Linux VMs, set up access to them, and make sure they're (self-)applying security updates. From time to time you'll need to roll to new OS releases and do restarts for kernel fixes. But that stuff isn't all that hard to learn.
Still - I'd be curious to know where your threshold is for touching Linux. With Heroku you never did, right? Dynos could have run Windows for all you knew? If you had a tool that e.g. you gave your cloud credentials to, and it then spun up N Linux VMs, logged in, configured automatic updates and some basic monitoring then let you push servers straight from your git repo, would you be interested in that? How much did you rely on Heroku support going in and helping you fix app-specific problems live in production?
I wouldn't be surprised if I'm wrong and newer tools bridge the gap for a lower price and/or time investment, but I also wouldn't be surprised if I'm right and many places using Docker could save time/money offloading the maintenance to something more like a managed PaaS.
There's also netlify, vercel and similar sites but I think they're mainly geared toward all-javascript apps.
If you're really concerned, just have a CI job that rebuilds and tests with newer base image versions.
It's worth noting that the Docker experience is very different across platforms. If you just run Docker on Linux, it's basically no different than just running any other binary on the machine. On macOS and Windows, you have the overhead of a VM and its RAM to contend with at minimum, but in many cases you also have to deal with sending files over the wire or worse, mounting filesystems across the two OSes, dealing with all of the incongruities of their filesystem and VFS layers and the limitations of taking syscalls and making them go over serialized I/O.
Honestly, Docker, Inc. has put entirely too much work into making it decent. It's probably about as good as it can be without improvements in the operating systems that it runs on.
I think this is unfortunate because a lot of the downsides of "Docker" locally are actually just the downsides of running a VM. (BTW, in case it's not apparent, this is the same with WSL2: WSL2 is a pretty good implementation of the Linux-in-a-VM thing, but it's still just that. Managing memory usage is, in particular, a sore spot for WSL2.)
(Obviously, it's not exactly like running binaries directly, due to the many different namespacing and security APIs docker uses to isolate the container from the host system, but it's not meaningfully different. You can also turn these things off at will, too.)
Though also, the added complexity to my workflow is an order of magnitude more important to me than the memory usage/background overhead (especially now that we've got the M1 macs, which don't automatically spin up their fans to handle that background VM)
Therefore, if you use Docker 'idiomatically', configuring with environment variables, communicating over the network, and possibly using volumes for the filesystem, it doesn't make your actual backend code any less portable. If you want to doubly ensure this, don't actually tether your build/CI directly to Docker: You can always use a standard-ish shell script or another build system for the actual build process instead.
To that end- I prefer to just stick with modern languages whose first-party tooling makes them not really have to care what OS they're building/running on. That way you can work with the code and run it directly pretty much anywhere, and then if you reserve Dockerfiles for deployment (like in the OP), it'll always end up on a Linux box anyway so I wouldn't worry too much about it being Linux-specific
Agree that it's a strong argument for using newer languages with good tooling and dependency management.
You can just bind the dependencies that are running in Docker to local ports and run your app locally against them without having to docker build the app you're actually working on.
The closest I've come in years to having to really wrestle with it was the URL hacking needed to find the latest version willing to run on my 10-year-old MBP that I expect to run badly on anything newer than 10.13 - the installer was there on the website, just not linked I guess because they don't want the support requests. Once I actually found and installed it, it's been fine, except that it still prompts me to update and (like every Docker Desktop install) bugs me excessively for NPS scores.
It made sense at one point to use Macs but these days pretty much everything is electron or web based or has a Linux native binary. IMHO backend developers should use x64 linux. That's what your code is running on and using something different locally is just inviting problems.
I've used Ubuntu, WSL2 and currently a M1 mac and if I need to be mobile AT ALL with the machine I chose a Mac any day. For a desktop computer Ubuntu works great though
That’s quite the assumption. Graviton is very popular. I haven’t touched x64 stuff in a very long time. Perhaps such generalization is a bad idea.
It may be less than ideally efficient in processor time to have everything I work on that uses Postgres talk to its own Postgres instance running in its own container, but it'd be a lot more inefficient in my time to install and administer a pet Postgres instance on each of my development machines - especially since whatever I'm building will ultimately run in Docker or k8s anyway, so it's not as if handcrafting all my devenvs 2003-style is going to save me any effort in the end, anyway.
I'll close by saying here what I always say in these kinds of discussions: I've known lots of devs, myself included, who have felt and expressed some trepidation over learning how to work comfortably with containers. The next I meet who expresses regret over having done so will be the first.
Maybe you've dealt with Python deployments and been bitten by edge cases where either PyPI packages or the interpreter itself just didn't quite match the dev environment, or even other parts of the production environment. But still, it "mostly" works.
Maybe you've dealt with provisioning servers using something like Ansible or SaltStack, so that your setup is reproducible, and run into issues where you need to delete and recreate servers, or your configuration stops working correctly even though you didn't change anything.
The thing that all of those cases have in common is that the Docker ecosystem offers pretty comprehensive solutions for each of them. Like, for running containers, you have PaaS offerings like Cloud Run and Fly.io, you have managed services like GKE, EKS, and so forth, you have platforms like Kubernetes, or you can use Podman and make a Systemd unit to run your container on a stock-ish Linux distro anywhere you want.
Packaging your app is basically like writing a CI script that builds and installs your app. So you can basically take whatever it is you do to do that and plop it in a Dockerfile. Doesn't matter if it's Perl or Ruby or Python or Go or C++ or Erlang, it's all basically the same.
Once you have an OCI image of your app, you can run it like any other application in an OCI image. One line of code. You can deploy it to any of the above PaaS platforms, or your own Kubernetes cluster, or any Linux system with Podman and Systemd. Images themselves are immutable, containers are isolated from eachother, and resources (like exposed ports, CPU or RAM, filesystem mounts, etc.) are granted explicitly.
Because the part that matters for you is in the OCI image, the world around it can be standard-issue. I can run containers within Synology DSM for example, to use Jellyfin on my NAS for movies and TV shows, or I can run PostgreSQL on my Raspberry Pi, or a Ghost blog on a Digital Ocean VPS, in much the same motion. All of those things are one command each.
If all you needed was the static binary and some init service to keep it running on a single machine, then yeah. Docker is unnecessary effort. But in most cases, the problem is that your applications aren't simple, and your environments aren't homogenous. OCI images are extremely powerful for this case. This is exactly why people want to use it for development as well: sure, the experience IS variable across operating systems, but what doesn't change is that you can count on an OCI image running the same basically anywhere you run it. And when you want to run the same program across 100 machines, or potentially more, the last thing you want to deal with is unknown unknowns.
If you freelance and work on different project, sure rvm is a great thing, but docker will contain it even better and you won't litter your work machine with stuff like mine is after a few years.
That said, I recently discovered that VS Code has a feature called "Dev Containers"[0] that ostensibly makes it easy to develop inside the same Docker container you'll be deploying. I haven't had a chance to check it out, but it seems very cool.
[0] https://code.visualstudio.com/docs/devcontainers/containers
web:
image: rubylang/ruby:3.0.1-focal
volumes:
- .:/myappCurrently, I'm working in 3 projects at once + my own. Even though it's all PHP, the projects are PHP 8.1 / Symfony 6.2, PHP 8.2 / Symfony 6.2, PHP 8.1 / Laravel 8 and the newest project I joined is PHP 7.4 / Symfony 4. I have to adhere to different standards and different setups; I don't want to switch the language and the yarn or npm versions or the Postgres or MySQL versions and remember each one. Some might use RabbitMQ, another Elasticsearch.
Docker helps me tremendously here.
Additionally, I add a Makefile for each project as well and include at least "make up" (to start the whole app), "make down", "make enter" (to enter the main container I'll be working with), "make test" (usually phpunit) and perhaps "make prepare" (for phpunit, phpstan, php-cs-fixer, database validate) as a final check before adding a new commit.
Locally, I had additional aliases on my shell. "t" for "make test", "up" for "make up".
So now I just have to go to a project folder, write "up" and am usually good to go!
Modern languages with modern tooling can automatically take care of dependency management, running your code cross-platform, etc [0]. And several of them even have an upgrade model where you should never need an older version of the compiler/runtime; you just keep updating to the latest and everything will work. I find that when circumstances allow me to use one of these languages/toolsets, most of the reasons for using Docker just disappear.
[0] Examples include Rust/cargo, Go, Deno, Node (to an extent)
If the new project without any tests with a PHP 7.2 and MySQL installation and I need to upgrade it to PHP 8.2, I first need to write tests and I can't use PHP 8.2 features until I've upgraded.
Composer is the dependency manager, but I still need PHP to run the app later on. And a PHP 7.2 project might behave differently when running on PHP 7.2 or PHP 8.2. And sometimes PHP is just one part of the equation. There might be an Angular frontend and a nginx / php-fpm backend, maybe Redis, Postgres, logging, etc. They need to be wired together as well. And I'm into backend web development and do a bit of frontend development, but whoa, I get confused with all the node.js, npm, yarn versioning with corepack and nvm and whatnot. Here even a "yarn install" behaves differently depending on which yarn version I have. I'd rather have docker take care of that one for me.
I feel like "docker (compose)" and "make" are widely available and language-agnostic enough and great for my use cases (small to medium sized web apps), especially since I develop on Linux.
Something language-specific like pyenv might work as well, but might be too lightweight for wiring other tools. I used to work a lot with Vagrant, but that seems to more on the "heavy" side.
Edit: I just saw your examples, unfortunately I've only dabbled a bit with Go and haven't worked with Rust yet, so I can't comment on that, but would be interested to know, how they work re: local dev setup.
Yeah- so in Rust, the compiler/tooling never introduces breaking changes, as (I think) a rule. For any collection of Rust projects written at different times, you can always upgrade to the very latest version of the compiler and it will compile all of them.
The way they handle (the very rare) breaking changes to the language itself is really clever: instead of a compiler version, you target a Rust "edition", where a new edition is established every three years. And then any version of the Rust compiler can compile all past Rust editions.
Node.js isn't quite as strict with this, though it very rarely gets breaking changes these days (partly because JavaScript itself virtually never gets breaking changes, because you never want to break the web). Golang similarly has a major goal of not introducing breaking changes (again, with some wiggle-room for extreme scenarios).
> I get confused with all the node.js, npm, yarn versioning with corepack and nvm and whatnot. Here even a "yarn install" behaves differently depending on which yarn version I have.
Hmm. I may be biased, but I feel like the Node ecosystem (including yarn) is pretty good about this stuff. Yarn had some major changes to how it works underneath between its major versions, but that stuff is mostly supposed to be transient/implementation-details. I believe it still keys off of the same package.json/yarn.lock files (which are the only things you check in), and it still exposes an equivalent interface to the code that imports dependencies from it.
nvm isn't ideal, though I find I don't usually have to use it because like I said, Node.js rarely gets breaking changes. Mostly I can just keep my system version up to date and be fine, regardless of project
Configuring the Node ecosystem's build tools gets really hairy, but once they're configured I find them to mostly be plug and play in a new checkout or on a new machine (or a deployment); install the latest Node, npm install, npm run build, done. Deno takes it further and mostly eliminates even those steps (and I really hope Deno overtakes Node for this and other reasons).
> maybe Redis, Postgres, logging, etc. They need to be wired together as well
I think this - grabbing stock pieces off the shelf - is the main place where Docker feels okay to use in a local environment. No building dev images, minimal state/configuration. Just "give me X". I'd still prefer to just run those servers directly if I can (or even better, point the code I'm working on at a live testing/staging environment), but I can see scenarios where that wouldn't be feasible
What if you want to start a new project using the latest postgres version because postgres has a new feature that will be handy, but you already maintain another project that uses a postgres feature or relies on behaviour that was removed/changed in the latest version? You're going to set up a whole new VM on the internet to be a staging environment and instead of setting up a testing and deployment pipeline you're going to just FTP / remote-ssh into it and change live code?
you define an apps entire chain of dependencies including external services in a compose file / set of kube manifests / terraform config for ecs. Then in the container definition itself you lock down things like C library and distro versions: maybe you use specially patched imagemagick on one project or a pdf generator on another, and fontconfig defaults were updated and it changed how aliasing works between distro releases and now your fonts are all fugly in generated exports... stick all those definitions in a Dockerfile and deploy onto any Linux distro / kernel and it'll look identical to it does on local
nevermind this, check out this thread to destroy your illusion that simply having node installed locally will make your next project super future proof https://github.com/webpack/webpack/issues/14532 and note that some of the packages referencing this old issue in open new bug reports are very popular!
if you respond please do not open with "yeah but rust", I can still compile Fortran code too
I have to assume you work by yourself or in an extremely small company to be able to handle project complexity without docker.
It includes running Rails and also Sidekiq, Postgres, Redis, Action Cable and ties in esbuild and Tailwind too. It's all set up to use Hotwire as well. It's managed by Docker Compose. The post also includes a ~1h hour ad-free YouTube video. The example app is open source at https://github.com/nickjj/docker-rails-example and it's optimized for both development and production. No strings attached. The example app has been maintained and deployed a bunch over the years.
It's a super productive framework to develop in, but deploying an actuals Rails apps - after nearly 20 years of existance, still seems way more difficult than it should be.
Maybe it's just me.
Makes deployment super easy.
That is in fact, I think, why fly.io is investing in trying to make it easier to deploy Rails, on their platform.
But also contributing to the effort to make Rails itself come with a Dockerfile solution, which Rails team is accepting into Rails core because, I'd assume/hope, they realize it can be a challenge to deploy Rails, and are hoping that the Dockerfile generation solution will help.
Heroku remains, IMO, the absolute easiest way to deploy Rails, and doesn't really have a lot of competition -- although some are trying. Unfortunate because heroku is also a) pricey and b) it's owners seem to be into letting it kind of slowly disintegrate.
I'm really encouraged that fly.io seems to be investing in trying to match that ease of deployment for Rails specifically, on fly.io.
It's still going to be some what challenging for some folks as this Dockerfile makes its way through the community, but this is a really small step in the right direction for improving the Rails deployment story. There's now at least a "de facto standard" for people to build tooling around.
Fly.io is going to switch over to the official Rails Dockerfile gem for pre-7.1 apps really soon, so deploying a vanilla rails app will be as simple as `fly launch` and `fly deploy`.
Checkout Cloud 66!
Deploying Rails to PAAS solutions is also easy but used to be expensive with Heroku. I very recently started using DigitalOcean Apps for a personal project and it's been very easy, while costs look acceptable.
The issue is how much RoR does and how tightly it is with its build toolchain - gems often require ways of building C code, whatever you use to build assets has its own set of requirements, there is often ffmpeg or imagemagik dependency. In my opinion, a lot of the issues are from Ruby itself being Ruby.
I agree that it's silly for such productive framework to be such PITA to deploy. To be fair, I pick RoR deployment over node.js deployment any day. I still not sure how to package TS projects correctly.
- productions servers should not have C compiler installed (why???) - compiling assets on every single VM that runs service is twice as silly
I'd still take that over something like AWS Beanstalk.
Being able to generate a war or jar as a released binary is something that would be cool to see in the ruby world.
Whenever you've covered your bases, Rails grows in complexity. Probably necessary complexity to keep up with modern world.
Solved asset pipeline pain? Here's we packer. No, let's swap that out for jsbundle.
Finally tamed the timing of releasing that db migration without downtime? Here's sidekiq, requiring you to synchronize restarts of several servers. Oh, wait, we now ship activejob.
Managed to work around threads eaten by websocket connections in unicorn? Nah, puma is default now. Oh, and here's ActionCable you may need to fix all your server magic.
Rails is opinionated. And that is good at times. But it also means a lot of work on annoying plumbing having to be rebuilt for a new, or shifted opinion. Work on plumbing, that is not work on your actual business core.
To be fair, this was already an issue whenever you have more than one instance of anything. Whether it's an extra sidekiq or two web servers or anything else, you have a choice of: stop everything and migrate, or split the migration into prepare, update code, finalise migration.
But async workers have the tendency to be busy on long running processes, whereas a web server typically has connections that last at most seconds. Their different profile makes restarting just a tad harder.
With Ruby and Python, most instructions for getting something going is just a series of "install this or that", "modify this file over there", "you could do this or that", "call this, than that, and then that", etc. These instructions tend to be developer focused. Virtual environments are usually left as an exercise to the reader and failing to use those leaves you with a big mess on your filesystem. Just pretend your production server is a snowflake developer laptop and you'll be fine seems to be the gist of it. Except of course that doesn't quite work like that anymore in many places and you need to take some steps to prevent that.
I spend some time face palming myself through the Apache Air (python) documentation trying to figure out a sane way to get that on a production environment. As it turned out that involved jumping through quite a few hoops. My conclusion was that whoever wrote that, was not used to dealing with production environments.
So, good that they are tackling this in the rails community. Stuff like this should not be an afterthought. With docker, you don't really need any virtual environments anymore. And you can also use them for development. That actually simplifies getting started instructions for both developers and operations people. Just use this container for development and run this command to push your production ready image to your docker registry of choice. No venvs, no gazillions of dependencies to install, etc.
I'm super stoked about the included Docker config and all the blog posts it will shortly inspire. Finding best-practices for Docker based deployments has been anything but fun. I'm still not sure how we'll implement the equivalent of `cap production deploy:rollback` and the like with Docker. Not that we use that basically ever, but knowing it's available is great.
> We auto-magically add, configure, and deploy your services described in the procfile.
…but that’s an outright lie and they caveat that statement by saying they only support single processes which completely defeats the entire purpose of Procfiles.
https://docs.railway.app/deploy/builds
Maybe Fly does this better and I’ll give them a try.
- https://gitlab.com/sdwolfz/docker-projects/-/tree/master/rai...
Haven't spent the time to document it. But the general idea is to have a `make` target that orchestrates everything so `docker-compose` can just spin things up.
I've used this sort of thing for multiple types of projects, not just Rails, it can work with any framework granted you have the right docker images.
For deployment I have something similar, builds upon the same concepts (with ansible instead of make, and focused on multi-server deploys, terraform for setting up the cloud resources), but not open sourced yet.
Maybe I'll get to document it and post my own "Show HN" with this soon.
The downside of capistrano was that you'd be responsible for patching dependencies outside of the Gemfile and there would be occasional inconsistencies between environments.
That said, the cap approach was easy for a static target, but a horizontally scalable (particularly with autoscale) was an utter nightmare. That's where having docker really shines, and is IMHO why cap has largely been forgotten.
It's fine for people who are just starting out or want a repeatable environment.
Also, if you want stability and fewer headaches long-term:
- use a RHEL-derived kernel and customize the userland (container or host) quay.io has a good cent 9 stream. Ubuntu isn't used at significant scale for multiple reasons, and migrating over later is a pain.
- consider podman over docker
- use packaging (nix, habitat, or rpms) rather than make install (and use site-wide sccache)
- container management (k8s or nomad)
- configuration management (chef) because you don't always have the luxury of 12factor ephemeral instances based on dockerfiles and need to make changes immediately without throwing away a database cluster or zookeeper ensemble
- Shard configuration and app changes, with a rollback capability
- Have CI/CD for infrastructure that runs before landing
- Monitoring and alerting
- Don't commit directly to production except for emergencies. Require a code review signoff by another engineer. And be able to back out changes.
- Have good, tested backups that aren't replication
- Don't sweat the small stuff, but get the big stuff right that doesn't compound tech debt
If (one of the) front-facing servers do regular http caching (a good idea anyway to play nice with rails caching[1])", you can probably "serve" the static assets straight from rails and let your proxy serve them from cache (if you don't have/need a full cdn).
[1] https://guides.rubyonrails.org/caching_with_rails.html#condi...
I replaced it with nginx intercepting the image URLs to serve them without even letting puma know and it was instantly zippy. I still use sendfile to support when I'm doing dev and not running nginx in front of it, and I'm not happy with the kind of leaky abstraction there, but damn are the benefits in prod too difficult to ignore.
...but... it's set to True....
The assumption here is that there is a caching proxy in front of the container, so Rails will only serve each assets once, so performance isn't critical.
> RAILS_SERVE_STATIC_FILES="true"
> RAILS_SERVE_STATIC_FILES - This instructs Rails to not serve static files
That is not what it does when you put `="true"`, nope.
In fact both, settings are common. On heroku you typically use `RAILS_SERVE_STATIC_FILES="true"` but put a CDN in front. But other people do false and have eg nginx serving them. If this is what fly.io means you to do... where is the nginx, not mentioned in the tutorial?
Whichever they intend to do, their code does not match their narrative of what it does.
That comment isn't in the Rails dockerfile, it was added by the OP.
https://github.com/rails/rails/blob/4f3af4a67f227ed7998fed57...
Standard heroku Rails deploy instructions are to turn on `RAILS_SERVE_STATIC_FILES` but also to put a CDN in front of it.
My guess is this is what fly.io is meaning to recommend too. But... yeah, they oughta fix this! A flaw in an otherwise very well-written and well-edited article, how'd it slip through copy editing?
It's super cool rails 7.1 is including the Dockerfile by default - not that rails apps need more boilerplate though..
[1]: https://pragprog.com/titles/ridocker/docker-for-rails-develo...
Deploying it was still always a hassle and involved searching for existing Dockerfiles and blog posts to cobble together a working one. At the beginning I always thought I'm doing something wrong as it's supposed to be easy and do everything nicely out of the box.
And dhh apparently agrees (Now at least: https://dhh.dk/posts/30-myth-1-rails-is-hard-to-deploy) as there's now a default Dockerfile and also this project he's working on, this will make things a lot nicer and more polished: https://github.com/rails/mrsk
The idea is I could run a command like `bundle packages --manager=apt` and get a list of all the packages `apt` should install for the gems in my bundle.
Since I know almost nothing about the technicalities and community norms of package management in Linux, macOS, and Windows, I'm hoping to find people who do and care about making the Ruby deployment story even better to give feedback on the proposal.
This is one of those problems that sounds easy but gets really fiddly. I had a quick run at it from a slightly different direction a looooong time ago: binary gems (https://github.com/regularfry/bgems although heaven knows if it even still runs). Precompiled binary gems would dramatically speed up installation at the cost of a) storage; and b) getting it right once. The script I cobbled together gathers the dependencies together into a `.Depends` file which you can just pipe through to the package manager, and could happily use to strap together a package corresponding to the dependency list.
I've never really understood why a standard for precompiled gems never emerged, but it turns out it's drop-dead simple to implement. The script does some linker magic to reverse engineer the dpkg package dependency list from a compiled binary. I was quite pleased with it at the time, and while I don't think it's bullet-proof I do think it's worth having a poke at for ideas. Of course it can only detect binary dependencies, not data dependencies or anything more interesting, so there's still room for improvement.
My primary use case for this is PCI Compliance. While PCI DSS and/or HIPAA do not specifically rule out Docker, the principle of isolation leans heavily twoard the principle that web hosts must be running on a private virtual machine.
This rules out almost all docker based-PaaS (including Fly.io, Render.com, AWS App Runner, and Digital Ocean), as these run your containers on general Docker hosts. In fact, the only PaaS provider that I can find advertising PCI compliance is Heroku, which now charges +$1800/month plus for Heroku Private to achieve it.
I would love to share my configuration with anyone that needs it.
> Fly.io doesn't actually run Docker in production—rather it uses a Dockerfile to create a Docker image, also known as an OCI image, that it runs as a Firecracker VM
By the same token, I think this reinforces the point that Docker itself is not considered PCI Compliant, unless we are simply treating it's config files as a DSL. And in that case, if you want PCI Compliance and go with the Docker DSL, then you are locked into providers that offer this same "transmorgification" from Docker to Firecracker.
Happy to hear that there are still other Capistrano users out there! I will push up my config later today.
"Ruby on Whales: Dockerizing Ruby and Rails development"
https://evilmartians.com/chronicles/ruby-on-whales-docker-fo...
Previously posted to hn - but without any comments.
Still, I'd much rather use Nix/Guix over Dockerfile.
I can chose my hosting provider and switch each other.
May I suggest you take a look at Roda or Hanami?
Use Alpine liberally for local images if you like, but don't use it for production.
We take the exact opposite approach: default to Alpine based images, only use another base OS if Alpine doesn't work for some reason. The majority of our underlying code base isn't C-based, so maybe that's why Alpine has been successful for us, but as always, everyone's situation is different and YMMV.
That said, keep using Alpine! There’s no reason for folks to stop doing what they’re already doing if it’s working for them.
The new dockerfile is meant more for people who are just getting started that aren’t familiar with Docker or Linux.
We support Kubernetes and Nomad as deployment platforms, and folks are welcome to build their own builder plugins if they need to support others.
You should run at least 2 servers for redundancy, regardless of size. You just can't lean on a single server, even if you can squeeze thousands of RPS out of it (big doubt).
It will inevitably fail, and you will have downtime.
It makes your local development environment incredibly resilient.
There are a lot of dockerfile / compose examples on github.
* https://github.com/docker/awesome-compose * https://github.com/jessfraz/dockerfiles
Getting a Docker compose dev env working is another can of worms. Maybe I should write about that next?