CapRover: Build your own PaaS
caprover.com
caprover.com
Basically, you can create a bare git repository on your server (`git init --bare`), and put a `hooks/post-receive` script within it that will clone sources in a temporary directory, build the docker image and rotate containers. That way, you can `git push` to build and deploy, and it's easy to migrate server.
The added bonus is that you now have a central git repos that can act as backup, so you don't need github or gitlab.
The main painpoint, which I find dokku interesting for (and I assume caprover too) is zero-downtime deployment. But well, if this is critical, you probably need something more extensive.
#!/usr/bin/env bash
export APP=appname
export DOCKER_OPTS=""
unset GIT_DIR
rm -rf /home/username/apps/$APP
cd /home/username/apps && \
git clone /home/username/git/$APP && \
cd $APP && \
echo building image && \
docker build -t $APP .
if [[ "$?" != "0" ]]; then
echo "error while building image."
exit 1
fi
echo "Stopping previous container..."
docker stop $APP
echo "Starting new container..."
sleep 1
docker run -d --name $APP --rm $APP rm -rf /home/username/apps/
Guess how I learned this lesson? :) GIT_WORK_TREE=/home/username/apps/$APP git checkout masterNginx is made to load a configuration: you don't have the auto-configuration that comes with service discovery. Service discovery is doable with a standard HTTP API through /var/run/docker.sock or /var/run/podman/podman.sock in more advanced systems.
As such, service discovering HTTP servers are more reliable because it's built with a service isolation from the ground up: if one service has some poor value then it won't work, but it won't block the other services.
Nginx is too far behind now, they might have some service discovery module, but even then the thing that happens when your configuration autogenerates (like with snapshot testing) is that you still have to read the configuration it generates. Traefik offers a great dashboard for this so it's even more pleasant than reading a configuration file that you didn't even write ;)
For sure, I bet that in a patch into something like CapRover (or your own solution), changing nginx to traefik would end up removing quite a lot of code ;)
I'm not really sure what you mean "acheiving ZDD", ZDD is complicated any time there's a data schema migration, not to mention that containers deployment traditionally is "delete a container: KILL a process" and "create another one like cattle". uWSGI for example, could gracefully renew every worker process on SIGHUP, but re-creating the uWGSI process in another container defeats that. Maybe you have some kind of blue green deployment, maybe even canary, in this case I wonder if basing a container platform on configuration files such as those for nginx would really make it to ZDD. Would love to read more about your setup
Ingress management is done with the very useful nginx-proxy[0] service that loads virtual host definitions directly from the docker daemon and sets virtual hosts based on env vars set on the container. Configuration changes are loaded using an nginx reload, so even if there was an error in the configuration (which I personally have never run into though is likely possible), it wouldn't take effect. LE is then handled using the nginx-proxy-letsencrypt-companion[1]. My goal was to abstract away reverse proxy+cert management and I think any solution (traefik, caddy, etc) would work here and I'm more than happy to change it. More or less just went with nginx since it was easy and I didn't have to do any configuration other than adding it to a docker-compose file.
I guess my goal wasn't to handle ZDD for stateful applications. As you mentioned, there's a plethora of issues that arise and make that type of application much more difficult to do ZDD. I tend to write a lot of stateless web apps for simple use cases and like to have an easy way to deploy them. In the primitive sense, creating a new container, waiting for it to be ready, and then swapping the upstream used for the reverse proxy with the new pointer would be ideal but isn't supported directly with docker-compose (as mentioned).
Happy to talk this through also, especially if you'd be interested in contributing!
[0] https://github.com/nginx-proxy/nginx-proxy [1] https://github.com/nginx-proxy/docker-letsencrypt-nginx-prox...
Traefik and Caddy are not comparable in my opinion because Traefik was literally made for self-configuration based on service discovery, see an interresting discussion here: https://www.reddit.com/r/selfhosted/comments/gq90aw/traefik_...
I completely agree with you about ZDD, 99.9% of uptime is plenty enough for 99.9% of the projects, and trashing the container to start a fresh process from a fresh system build does come with other advantages. Sure, any kind of blue/green deployment or canary would be really nice to see, but it wouldn't seem to create a lot of value for 99.9% of the projects, and for the rest well there's k8s that deals with clusto
Currently I'm just using a bunch of ansible roles with an ansible command line wrapper, so I'll do `bigsudo yourlabs.netdata @somehost` and it'll auto-install yourlabs.traefik if not already there, which will auto-install yourlabs.docker if not already there, and basically just leave me with `https://netdata.somehost.fqdn`
Thank you for the invitation to contribute ! As you probably guessed, I'm a bit like you in the sense that I cannot live without making my own system, and I have made different design decisions:
- Python for server side, I find it more fun than JS, nothing we can do about that
- Python for client side, because we maintain our crazy isophormic component library in python
- Not docker, but podman, which can run rootless and daemonless (thought we need the daemon that provides a docker compatible API to have Traefik service discovery)
- Not docker build, but something I'm cooking on my own ("shlax") that I find a lot better for my taste, and that uses buildah which can build rootless
- Not docker-compose, but shlax, which aims to support a broader range of use cases (such as backup/restore)
- The thing I'm building is first a really KISS Sentry alternative, then also a GitLab alternative, and I'm in the process of adding CI into it ... but I stopped doing that until I have finished my little Python lib ("shlax") that replaces docker/compose and ansible to have something to put in the CI test that's not tech that I'm trying to move away from, and so that it can build/test/deploy itself,
So, I suppose our goals and design decisions are a bit too different, but I can assure you that I'm always happy to see CapRover featured in social media, and I'm always happy to discuss rare passions like that, if you're looking for a crazy friend recoding his entire little world to just talk about these kind of things feel free to send me an email or give me a call ;)
For those noting "why don't you just use Linux / k8s / ...", that feels close to the original complaints re: Dropbox on Hacker News[2]. I've run clusters hundreds of nodes in size myself but CapRover gives me the pleasure of not having to sweat the small details. You can get this from other platforms but usually there's a dollar cost tied to each option. When I'm experimenting I don't want to have a dollar cost attached.
Deploys are trivial. The default nginx setup is most of what I'd want to do. LetsEncrypt is a single button click. Monitoring is included by default. If I need to scale up, everything I'm pushing is Docker containers. If I want to experiment, there's great fun in looking at the included "One click apps / databases" and just playing around.
CapRover is just a lovely freeing experience that will do what you need :)
I think CapRover has the best experience out of the Dokku/Flynn/CapRover "group". Not a huge Dokku fan. Would use Flynn over Dokku again, but I'd rather use CapRover over both.
If you're intent on using Dokku, there's a useful web console:
CloudRon [1] is actually really great, but its pricing is prohibitive (and changes frequently) for side-projects, which is when I most want to use something like this rather than just deploying/managing the k8a myself.
[1]: https://cloudron.io/
Does CapRover manage a shared storage or do you still manage it yourself (ie. Ceph, gluster ...) ?
It's not the best for hosting many static pages, as you'll need a HTTP server for each site anyway.
But my main gripe is that there is only single factor authentication and you can't easily secure it more other than using a strong password and a hidden subdomain. (because of webhooks, acme, etc. I guess)
I haven't yet gone through the discussion but would you mind letting me know if you're satisfied with the outcome? You appeared to be fighting for increased privacy, so thank you.
Also you'll see that in my comment, I raised another issue RE: lack of two factor auth. I'm curious, why do you think single factor auth is fine? Simply because brute force for a 30 char password is not practical on todays hardware? Or is there something I'm missing?
- sneak and I have fundamental differences in what we call spyware. The issue that was brought up in that thread is standard analytics events - nothing like stealing passwords or etc.
- Regardless, CapRover uses NetData 1.8 [1] . According to NetData's github page, they added analytics in NetData 1.12 [2] , so even if you're concern with analytics events, this issue won't apply to you anymore.
Regarding two factor auth: CapRover blocks brute-force attacks by limiting number of wrong passwords per minute.
[1] https://github.com/caprover/caprover/blob/48440db14aa115aca1...
RE: 2fa. Brute force protection is a step in the right direction, but passwords can leak in various ways, brute force isn't the only attack vector. I'll comment in the actual two factor auth discussion on the CapRover GitHub issue though.
Brute forcing a 30 char (or even 20 char) password over the network is infeasible. Do the math. Regardless, as the CapRover developer pointed out in a sibling comment, it rate limits attempts, but in the case where you are using a long, random password, it would be fine even if it didn’t.
Brute forcing this, you would have to try every combination. Which means for a four-letter long password: 26x26x26x26 = 456 976 possible passwords.
For a 10 letter long password: 26^10 = 141167095653376 possible passwords.
Clarifying it further is “number of days to brute-force if you can try (eg) 10k requests/sec”.
I kept it to 26 letters to keep the math simpler (or rather - the numbers smaller, for myself, really).
Number of days to brute-force if 10k requests/sec (26 letters still...):
4-length password = 45 seconds
10-length password = 453 years
Please give me a heads up if my math is off.
isn't that what virtualhosts are for?
You could try and mount the static site's container files to the local filesystem and serve from there I suppose, but there's currently no easy way to do so.
This product is undoubtedly the P in PaaS, but there is no service behind it. If your company uses this as an alternative to a real Heroku/AWS/xyz PaaS, you must have engineers at hand for 24/7 ops, scaling servers and fixing bugs. In my opinion, this is quite risky for anything running in production and should not survive a cost-benefit analysis.
I completely disagree, the difference of price between dedicated servers and even EC2 instances is completely amazing.
This is what you get for less than $200/month with a dedicated server:
1× AMD EPYC 7281 CPU - 16C/32T - 2.1 GHz, 2 × 1 To NVMe, 96 Go DDR4 ECC, unmetered 750 Mbps
In one of my companies the AWS bill is just completely insane, we have like half that hardware, with a really small bandwidth, which is metered, for more than $800/month, which is fine while we're on free credits.
I love working for cloud companies, it's a lot of fun, but when it comes to my money then I never go for anything but a dedicated server.
When you got applications that don't require high availability while needing a very low cost per CPU, dedicated servers just make sense. We are running a cluster of a few high-CPU dedicated servers for our data-science team, and it just makes sense: we don't need 99.99%+ availability, and the servers we rent are cheaper than the equivalent AWS storage cost alone ... The op cost of managing these is exactly the same as managing equivalent EC2 instances. We don't need backups either.
On the other side, we got some low-CPU web services that require high availability, redundancy and reliable backups. For these I just use Heroku. It's extremely reliable and easy to operate, while only costing about $100/month (a few hobby dynos + a fully managed PgSQL DB). Sure it's probably 5x more expensive than a dedicated server with 10x the performance, but I don't have to worry about backups, availability and scalability. And these apps just don't need this 10x faster CPUs anyway.
How do you handle Heroku outages then?
I am not affiliated nor haven't tried CapRover (for special reason of: coding my own for my tastytastes), but I would bet any standard system administrator could get a 99.9% uptime after the second month of production without particular effort (unless they don't know underlying technologies ie. "what a container" "what http" "what is namespace" "what iptables" ...)
As an example of the latter bit, if you are running your own hardware and need to add another host and you do not have a spare lying around, then you need to order one. It has to be shipped. Someone has to unpack it. Someone has to make sure that the data centre has sufficient power. Someone has to install it, its power and its network cables. Each of these steps takes time, but also each step is an opportunity for friction.
By contrast, with a service, you would just add a new host. Five minutes later you are up and running. That gives you an operational nimbleness that you wouldn't otherwise have had.
Servers, for the most part, just work. In DC climate-controlled environments, hardware failures is exceedingly rare. Apart from harddrives, most hardware will happily tick along for a decade, if not longer.
Sane production-grade OSes (read: not Ubuntu) will also happily run for literal years with zero human intervention. For obvious reasons, it's a bad idea to not patch your systems, but things will continue to "just work" pretty much forever unless you're running really shitty code.
For renting vs buying servers, there's upsides and downsides. Buying gear is far far cheaper if you plan to be around for more than a year, but renting dedicated servers gives you a lot more flexibility -- to provision a new server, you hit a button in their online panel, wait 15 minutes, then let your deployment strategy take care of the rest.
I find it almost mind-boggling that AWS and friends have convinced people that it's normal to spend ridiculous amounts of money for fairly "meh" service specs in what's essentially VMs.
Not to mention that at that scale you have plenty of redundancy and, if your ops team knows what they're doing, automagic failover / HA. Anything that happens can easily "wait till Monday", no need for 24/7 anything.
And my experience from providing devops services to clients on a contract basis is that the clients who use cloud services tends to need more, not less, devops assistance.
The idea is that you ask OpenStack a VM and it will give it to you, dealing with the lower level details for you.
PaaS means that you ask it to deploy a service and it will deploy it for you, dealing with lower level details for you.
You're probably more likely to see OpenStack called a private cloud or on-prem cloud than "IaaS" these days. And OpenShift is usually called a Container Platform rather than a PaaS.
"Container platform" seems pretty vague to me, PaaS means something I know right away.
I mean, k8s is a container platform too isn't it ? But you'll need to build what we called a PaaS on top of it yourself (or use something like Kelproject, OpenShift ...)
PaaS isn't a verboten term or anything like that. But it turns some people off because it was most associated with services/products/projects that mostly focused on a simplified developer experience at the cost of flexibility.
k8s for me is a framework, OpenShit, Rancher, KelProject would be "distributions" of k8s, just like Linux kernel and distributions including it.
As a person who writes technical requirements and implementation document, it strikes to me when I'm asked to document implementation of a "SaaS" that there will be paid accounts and billing.
Maybe CapRover will provide paid accounts on managed servers in which case they would be creating a SaaS with their PaaS solution.
But again I'm not talking from a "managerial" perspective of the definitions, rather from a technical one. I suppose at this stage CapRover is trying to attract technical users rather than managerial ones (unless they have something to sell for cash but I didn't see it on their site or just missed it)
Many IT depts would do themselves a massive favor to deliver actual services instead of “just infra and some stuff thrown on top” and call it service delivery.
Tools like in this link can help, but a big part is simply about automation and delegation/self provisioning.
Platform as a Service or anything "as a Service" means someone else provides it as a service (ie subscription). The Platform part is all this is offering. So it is not a Platform as a Service.
It is not necessarily hard tied to a business model, but of course I understand that this is the common usage.
It’s really about abstractions and consumability.
This is my interpretation of the NIST meaning of aaS.
My read on whether something is a service or not is, can I make a request of the thing in simple terms, and have the thing carry out all the messy details on my behalf?
Dokku on the other hand has support for buildpack deployment as well as Procfile support for running multiple processes.
I prefer Dokku. The main reason is that I only need a single server for my apps and running Docker Swarm adds complexity.
I wrote about some other differences in my blog: https://www.mskog.com/posts/heroku-vs-self-hosted-paas/
Swarmlet seems to be very young and not production ready. Still, I am exited to see how this project will evolve with time.
When it comes to operations I often feel overwhelmed, even though I've done DevOps and automation work in the past.
Most of the things I'm working on professionally don't need the "scale" part, but the "robustness" and especially "ergonomics" parts. When I look at most infrastructure solutions, then I often get a combination of "this is too complex" and "I don't need this".
So I was drawn to solutions like Heroku at some point, but there you cannot even do the most basic thing: persistently writing to the filesystem. So you are forced to introduce system level complexity and coordination for such a fundamental feature.
Naturally I tend to prefer simple tools that enable things rather than constrain them.
Side note: I think when the "code has to run on some computer" problem is finally solved, then we likely see an explosion in productivity in our industry.
Migrating the docker image building from the dokku server to a CI would be easier to do without this. On top of that, deploying an existing software into your machine would be easier.
[1] http://dokku.viewdocs.io/dokku/deployment/methods/images/#de...
So basically, you could put a Dockerfile file container just FROM and MAINTAINER, referring the image you want to use in the FROM, and dokku will download and execute it on `git push` (provided it can access to the image repository).
I wonder if there's a ticket about this on dokku already
[EDIT] - Couldn't find anything... Some tickets about how the containers are built and changing the base image but not much about.
I wonder if you could jury rig something like kraken[0] and make sure wherever your building images is a peer or something... Of course the simpler solution might be to add a CI step that just pushes the image (via the working `docker save` method) to the deployment machine(s)? Maybe if you have a staging environment, let CI push there, then if that machine is peered (via something like kraken) with production, production will get the image (though it may never run the image).
You will learn k8s and you will get the same thing as they do but with open components, industry standards and a whole industry moving in this direction.
I have already microk8s running at home with argocd. I have never had IaC that quick and that simple setup.
With traefik you can have your domains as well. Then just go to gitlab (or now to github, haven't checked out yet if i wanna migrate back) and register your microk8s cluster as a buildrunner.
Thats it you are set. Quite future proof setup, modern, stable, easy to use.
Deploying a simple app with a database with Dokku is something like: 1. Run command to create a database of your choice(Postgres, MySQL, Redis etc) 2. Run command to create application 3. Run command to link the database to the application 4. Push to the Dokku repo to deploy the application.
Kubernetes is just the future, used by much more people and you have the additional benefit of learning kubernetes which might help you in your job/day to day business etc.
If you are already thinking of operating CapRover/Dokku, i would strongly considering using kubernetes instead.
That being said, as long as it works, it works. And if your app is small enough never to get into the grey waters, all the better.
Are there any tools you recommend looking into, were I to take the next step? I don't plan on depending on CapRover to fill gaps in my knowledge for too long, but for now this product really is a good start for me.
Ironically it's best learned "on the job" (for me at least); just try to deploy your app from scratch. Play around with nginx/apache, letsencrypt, your db stack, packages installation etc. and get a working product.
I'm no expert by far in any of this, but think that knowing "just enough" about these tools really helped along the way. Up to the point where I can now use CapRover like tools with some degree of confidence, closing the full circle ;)
The same thing applies with these turnkey admin panels like cPanel or Plesk and which is why I don't recommend getting anywhere near those.
I was hoping to move over, but I won't just yet. Was hoping for two factor auth support [2]. The dashboard was publicly facing and only guarded by a password.
There was also an issue which concerned me: 'netdata image in use is spyware' [3] - however I have not digested that thread and its related discussions to understand if it's a genuine issue yet.
Finally, I was hoping to understand more about the motivations of the project. Who's funding this project? The OpenCollective [4] page shows an annual budget of $529.55 USD? What are the long term goals? How can they sustain themselves?
I use PM2 for running some of my Node.js apps, but I can see there's also PM2 Plus and PM2 Enterprise [5] which helps me understand how they're able to sustain the free version.
I don't mean to be pessimistic towards a piece of software which actually works extremely well. Just want to better understand its long-term suitability for deploying production-grade applications.
[1] https://news.ycombinator.com/item?id=23278095
[2] https://github.com/caprover/caprover/issues/493
[3] https://github.com/caprover/caprover/issues/553
[0] https://github.com/dokku/dokkuCombined with portainer (which u can install with caprover) I'm improving my docker knowledge. I'd recommend it for someone starting out with containers and "home labs".
As I said in that thread - this looks interesting, but the installation instructions put me off a bit. Open a port on your server, and don't change the default password `captain42` - then run a cli tool from your dev machine.
Maybe it's because there's so much arduous research required to finally figure out what magic commands to run to get something to work. Would having a set of HOWTOs that just explain the steps to set up each component work as well for you as a turn-key solution? (It would be great if we could start a trend of people writing a HOWTO.md after writing their README.md)
Honestly, I do a lot of the setup work for my actual job — including documenting/creating demos and examples for others — when I’m doing my own side stuff, I really just don’t want to bother, especially if it isn’t in production and it’s just on the home lab.
To use a crude analogy, I could build my own robust NAS with hardware components and a BSD or Linux distro optimized for storage and acting as a home server with better performance at a lower price than a Synology system. Or I could continue to use my 8-bay Synology NAS (that I really want to upgrade), because the appliance nature is worth the extra cost and pure performance deficits. There was a time I took great pleasure in maintaining all that stuff myself but honestly, I just want to plug it in and know it’ll work with all my machines without having to think about it.
Very cool!
However, I wish the caprover had built this experience on top of kubernetes (or k3s) instead of Swarm. The future of Swarm is really unknown and the ecosystem is undoubtedly behind k8s.
You can expose some ports on different nodes and point your external LB (for ex. cloudflare)
You can create complex containers that could update with security fixes without restarting. But it is easier to update an image e.g. once per week/day and auto restart the containers.
I wonder about the underlying instance's OS, though... in the past, for home servers, I've set up cron jobs to get OS updates and reboot, but that seems wrong for a web server I'd like to be always up.
Maybe create a new instance, update the OS, install the app, switchover? Is there automation for this kind of thing?
Spin up a cheap instance (DO has a preconfigured image ready to go), git pull and caprover deploy to test. I am pretty sure even the cheapest ones will be able to run that.
E.g. your database can be an app that doesn't have any web frontend.
On the second point, I vary between Caddy and Traefik, depending on the use case.
All 3 of these are open source on github with widely used code bases that are free to view and read as you want.
I'll admit that the open source nature limits the damage a state actor can do, though.
If you are going to go full paranoid, you can't pretend that the media is 100% trueful at every aspect, especially when it comes to internal US affairs.
Also, while every media source is biased, if you review enough angles on a given story, you can arrive at some semblance of a true account.