What One-person SaaS Healthchecks.io uses for hosting, hardware and software
blog.healthchecks.io
blog.healthchecks.io
I do use it for a purpose it probably was not designed for :) In my summerhouse, one of the circuit breakers trips every now and then (1-2 times per year) for no apparent reason, and since the fridge is connected to that breaker, i very much fear arriving at a summer house where the fridge has been unpowered for a couple of weeks.
The solution to that of course, was to setup a systemd timer on a Raspberry Pi (that also runs HomeBridge), which pings Healthchecks.io every 15 minutes. If it misses 2 pings it raises an alert.
So far i've had 2 notifications in the 2+ years it's been running, and both have been internet connection related, so monitoring the situation has apparently changed something :)
I set up a phone with Tasker, if the phone stops charging or fails to check in, it sends an alert.
Luckily, and as you say, monitoring changed the situation, as it never failed since.
Usually there's little any single person can do, but this waste of electricity is very selfish considering climate change.
You have that effect in many areas of life. A quite obvious example is working out. Of course it won't make much of a difference if you skip today's workout or abort your 40 min run today halfway cause you don't feel like it. But it's the mindset/attitude. It makes you a quitter. It enables you to do the same thing again. And again. And in the end, the cumulative damage of the change in mindset is magnitudes worse than the single event.
Ironically that is what kept me from setting up solar power. Being in Scandinavia, solar has a somewhat limited potential given that days are 7 hours long during winter, and December often has less than 20 hours of sunshine in total. My calculations for the "payback" time of the solar panels said that _MAYBE_ they would have saved something before their "expiry date" some 20 years into the future.
Considering the "energy and material waste" involved in creating solar panels, this task can be handled by energy companies much more efficiently than i ever can.
I am still considering a share in a windmill though :)
Everyone should be calculating payback on these investments every couple of years, because the calculus changes depending on incentives/tech/other factors (e.g. EV purchases, or massive cost of living changes)
First Google result: https://www.energy.gov/eere/solar/articles/busted-common-sol...
We have an annual power consumption of 6-8k kWh, with about half during winter (heat pump, yay). More energy is required during night time due to dropping temperatures and it being dark, so realistically, I could probably save about half.
At normal energy prices of €0.3/kWh, that means I would save €1200/year, and that’s without taking into account reduced electricity taxes for electric heating.
Taking those into account, all electricity used beyond 4000 kWh/year would be around €0.12/kWh. Taking that into account, I could save €480/year, still assuming I could save 4000 kWh/year out of 8000 total.
A solar panel installation _without_ batteries is about €13500, and saving €480/year means it would take 28 years to get back the investment.
I’m aware batteries can change that equation somewhat in my favor, but it’s ever harder to calculate anything usable with those as it depends if the battery can get charged during the day. In any case, a solar panel installation with batteries is around €20000.
Since I bought an electric car I've been drooling over the idea of having the car powered by the sun. I go into the office only on occasion so I don't need to charge fast either. But you need an available roof and an okay topography.
But in Scandinavia, I bet most energy goes to heating so improving insulation and adding ground thermal will probably be more efficient. The idea of a sun powered car is just so enticing, however. It's kinda sci-fi.
This is Denmark, where a large part of our power grid is based on renewables, and backup power is based on natural gas. There is no coal involved. When we import power it's from Norway/Sweden or Germany, so either renewables or nuclear. We're not at Norwegian CO2 levels (yet), but it's getting better :)
You can check for yourself here : https://app.electricitymap.org/map
> Since I bought an electric car I've been drooling over the idea of having the car powered by the sun.
Denmark has outlawed sales of new ICE cars from 2030, so at that point solar will make MUCH more sense, as i can charge my car for "free", and the time to ROI will be much shorter.
> But in Scandinavia, I bet most energy goes to heating so improving insulation and adding ground thermal will probably be more efficient.
I thought so to, but currently, in no small part due to the Ukraine situation, lots of people are replacing oil and natural gas heaters with heat pumps or central heating, and the central heating providers are installing large heat pumps as well.
As part as getting a heat pump, my house (built 1970's and renovated 2010's) has a total heat demand of roughly 13000 kWh. Assuming a heat pump with a SCOP of 3.5, that means i will need about 3700 kWh of electricity to heat it, which means that our total electricity consumption ends up around 7500 kWh, so about half the energy (in my case) being used for heating.
In my experience, anyway.
We tested it, and with our usual usage pattern of using the summerhouse every 2 weeks, the brand new fridge uses less power while powered for the 12 days between departure and arrival, than it does cooling down when powered on at arrival.
Despite being brand new, it uses 6-8 hours to get from 18C to 5C. I guess because it's brand new it is well insulated, so less cooling power is needed once it's cooled down, which translates to a weaker compressor with less power consumption.
As for the Raspberry Pi, it runs idle for 99% of the time, consuming _at most_ 1.15W. That amounts to 0.8 kWh in a month, which is about as much as a TV uses in 2.5 hours, or about a quarter the power consumption of a Sonos One Speaker sitting idle [1].
The total power consumption of the summerhouse uninhabited is ~1.8kWh/day, which includes the heat pump, fridge, internet modem and router, Raspberry Pi, Security Camera and a couple of hubs for IoT.
I started leaving a gallon jug of water (frozen) in mine as it helps give it enough thermal mass to even out the spikes.
My chest freezer maxes out at 180 watts TDP for instance, which is pretty anemic if you’re trying to freeze a pot of room temperature stew.
Owning a "summer" house?
Moving yourself to the summer house every two weeks probably has far more climate impact than the fridge.
Taking to a logical conclusion: you shouldn’t have a summer house at all. (That’s not a position that I hold, but one that seems more logical than “have a summer house; just empty out and unplug the fridge while you’re not there…”)
Another solution to this problem is to just empty the fridge and turn it off while not in the summer house for a few weeks.
> just empty the fridge and turn it off while not in the summer house for a few weeks
We tried that, but the fridge uses 6-8 hours to reach target temperature despite being brand new, and keeping food fresh without cooling for 10 hours is not easy.
It actually uses less power being powered on continiously for 2 weeks (12 days from departure to arrival), than it does being powered off and working overtime for 6 hours trying to cool down. It's not much less power, but still less.
I suppose you mean energy rather than power? (Cause the argument makes less sense with power.)
That's an interesting phenomenon then since it goes counter to my (rather naive) understanding of physics. There could be some anomaly with the fridges efficiency curve that causes this. Maybe it's really inefficient cooling down the contents from room temperature while being really efficient maintaining the cool temp? After all, the latter is its main purpose..
Normal power consumption of it (assuming 18C room temperature) is less than 0.4 kWh/day. That amounts to 4.8 kWh in a 12 day period. Cooling it down from 18C takes 5.6 kWh (measured).
Temperature varies of course, and the lower the temperature the lower the power consumption, and the heat pump keeps a minimum temperature of 12C. During long summer days, the indoor temperature can easily reach 26C, which will require more power, and i have not measured every possible scenario.
I'm not ruling out that actually cooling down the thing during winter will require less energy than keeping it running, but i'm fairly certain that doing the same from 26C will require more power than keeping it running.
This seems like a lot. It's 20 megajoules, so even at a COP of 1, it's enough energy to cool down 20 megajoules / (4182 J/kg/K) / 18 K = 265 kilos of water by 18 kelvin.
And a refrigerator is supposed to attain a COP better than 1, and I suspect you have less thermal mass around than 265 kilos of water.
Something is wrong-- either in measurement, or heat exchangers very badly occluded by dust, etc, or some other fault.
The fridge is in a cabinet, enclosed on 3 sides, so maybe it’s not letting enough hot air out ?
I could be wrong, but I think it was stated as using 184 kWh / year (at room temp 22C, optimal conditions), which is just over 0.5 kWh/day, and considering that the summerhouse is usually cooler than that (12-16C) when unoccupied, I would assume it used even less power.
As it is, for this entire week, the heat pump has (self) reported using 1.8 kWh/day, and according to my power company, my daily consumption has been 1.8 kWh/day. Considering that there are 5-7 IoT things plugged in, the fridge probably isn’t using much power.
Seems likely. Once it succeeds in warming its cabinet, a big share of that heat leaks back into the fridge. (And efficiency falls because of the bigger temperature difference).
* Every bit of warming that finds its way into a powered-off refrigerator is heat the refrigerator would need to take out anyways. And the heat flowing in declines as the inside warms up.
* There's standby power, etc.
* A refrigerator is expected to be more efficient removing heat when the temperature difference between inside and outside is less.
I'm not arguing that that's the case here. I'm arguing that you don't "always" expect things to be worse off letting it run. There is a threshold, when approaching the mentioned fantasy conditions, beyond which letting it run may be better.
Then it would never warm up, and at this limit the two cases are equivalent (powered and unpowered).
> when approaching the mentioned fantasy conditions, beyond which letting it run may be better.
Anywhere short of the limit, what I said holds. Anything better than infinite insulation, I don't know how to reason about.
Perfect insulation doesn't help if you open the door?
If you're saying 'well with perfect insulation it doesn't matter if it has power' - well yeah. It doesn't matter if it has power. So you aren't wasting any either leaving it on.
This continues all the way up to perfect insulation, though ultimately it's breakeven with perfect insulation.
> Perfect insulation doesn't help if you open the door?
This doesn't change the picture, but worse, it's just completely unrelated to the scenario. Who's opening the door when everyone has left for weeks?
Which seems unlikely in this case.
My chest freezer for instance turns on for less than 30 minutes a day even with people getting in and out. When I defrost it, it takes a day+ to cool down again.
The reason someone would leave the door open is to ensure no one accidentally left a jug of milk in there when they turned it off. Which would pretty much neutralize any possible gains in sheer grossness if nothing else.
I think you're a little confused. I'm a little lost at even how to minimally explain it to you. The rate at which heat leaks into the refrigerator declines as the refrigerator heats up, so this metric too improves with time.
> Which if the insulation is good, may take many weeks.
Implausible. "Many weeks"?
The best 20 cu ft refrigerators use about 375 kilowatt-hours per year according to energystar.gov. At a coefficient of performance of 10, that's an average of 400W flowing into the fridge and needing to be removed by the refrigerator.
At 20K of temperature difference, that's about 20W/degK. The contents will warm very quickly, absent phase change.
Having ice in there will slow things down until the ice melts; then it will warm very quickly. The typical numbers cited say you have 2-3 days with a full freezer without power before things have mostly thawed and become unsafe-- this is well over half of the way to room temperature in terms of energy budget (and in the best case of a full freezer where there's the most thermal mass inside).
(Specific heat capacity of water ~4200 J/(kg*k); heat of fusion of water ~334000 J/(kg*k)).
None of this changes the energy balance of the problem, though.
If you're writing a blog post about your service then chances are you're interested in driving people to the main website about that service?
So...make it easy for those potential users to get there. At the top of your blog have a prominent "what is [service]" link to the marketing home / about page or a call out box.
This guy does it as a CTA after the post which is a start but many, many people don't. That then leaves me and other potential users to have to figure out how to get to the homepage. I'll do it but you'll lose many others.
I almost didnt bother but I'm glad I did because it looks like a really useful service that i might have uses for. I'd assumed that it was a healthcare-related business tbh.
(Also: super-informative post.)
It is so cool to see one-person SaaS business running on old school bare metal servers without needing any fancy devOps / containerize tooling.
I follow a simple rule of thumb - the more abstractions you introduce the more complex it gets to maintain
Life is easier if you can get away with lesser abstractions.
You see this contextual disagreement a ton here on HN. Those who gave taken the time to learn these tools and have practical experience using them for real deployments see the alternative as primitive, error prone and fundamentally limiting. Those without the experience see the tools as overly-complex distractions. Both are true, depending on your situation. Like almost all tech decisions, there are no universally correct answers. The best decisions are those tailored to the specific circumstances.
In many ways it is a luxury to be able to use said abstractions and be able to open a ticket with someone else. Of course, a single-person startup could use the same tech, but it seriously impacts their costs and they don't have time to deal with the sprawl of abstractions.
Whereas when you are a one-man business, all you have is (free) time and prefer to spend as little money as possible while your business is not profitable. This means: walk away from dedicated databases, managed servers, huge aws bills, etc. You do everything "the old" way: rent a server (bare metal or cloud) and maintain it yourself.
This pattern plays out with many indie app developers as well, both now with iOS apps, and back in the day with Delphi shareware. The cost of taking a dependency on a 3rd party library turns out to be much higher for the solo developer than larger orgs. That's not to say they don't make bets, but they don't build the way that someone at a larger org would do because they have to maintain a higher degree of direct control.
Marco Arment (https://marco.org/) has elaborated a lot on this topic. You can also get a feel for the limits on the upper bounds of his developer experience. Many on HN are not willing to work at a low-level with such boring and dull technologies.
I know it's a bit silly, but anyone building out cloud services should look at this post and think about how easy it is to set up this. This is just slightly more complex than a static site, has a bit of heterogenity, but otherwise is a lot of known tools.
Can your stack provide all of this without breaking the bank for the people signing up? Can it be done without having to learn a bunch of bespoke systems? I think that Heroku's success in particular is totally down to managing this bespoke-ness balance (that is still kinda missing from container-based setups).
That's because their product is only slightly more complex than a static site.
I wish every project was as basic as a web server and database but it's pretty rare these days.
What you're seeing is survivorship bias where the SaaS companies that are doing well are the ones that can provide a simple, clean UI whilst adding the functionality that you need for the product to be useful.
Is it? What functionality are you thinking about that absolutely require more than that?
I find that most ideas even today can be solved with a web server + DB, while developers today like to over-engineer things and think about scaling a service for a million users while they haven't even figured out the value proposition of their product/service yet.
A lot of apps use an "AI" or "machine learning" component which usually just means a second heterogeneous internal service.
For a little bit larger companies it's pretty common to have a data analytics or data warehouse component.
In addition, a web server + database usually require some monitoring system. After you have more than a handful of servers, probably also log aggregation.
If you're just a single or handful of people working on something, you probably don't need any of that.
Once you get a little bigger (or downtime becomes expensive) those things get tacked on quickly.
I feel like the overwhelming majority of web based software projects can be catered for with exactly this setup.
If it's rare it's because everybody's drunk the Kool-aid
Having a 1-2 day outage because your architecture is not highly-available across data centres like this one just doesn't cut it. And multi-day outages are very much real because of the thundering herd effect as companies rapidly try to move between data centres exhausting capacity.
And much of the complexity in modern day architectures come from high availability and security.
i.e. they would likely be sharing the same power and network infrastructure.
So not redundant by any normal definition.
But you gotta pay more than 500$, way more.
If you're running a site like this, it's fine to be down for a bit, people will forgive you as you restore from backup (looks like a very solid foundation the site has).
on other bare metals, or on AWS, azure, ...
As long as you have database backup, git code, and build-scripts.
I run something similar (although smaller). My disaster recovery is: I have everything prepared, to go on AWS, if necessary. Would take less, than 1 hr to be back up (database size being the biggest time sink). If I wanted to minimize that time, I could have small DB replica running, so I would just have to run last day of wall files. But for my purposes everything less than a day is good enough.
And then you can take a few days, to find something cheaper, to migrate to.
At my last job we had one of those hyperfancy devops setups with all the fancy devops tools. Literally no one in the company knew how to spin up a new environment. I'm not exaggerating: no one knew how to run a dev environment and when they had to set up a new region for legal purposes it took weeks for the team. All of that was initially set up in an age of legends in the mythical past of a year and a half ago.
Theoretically it would all failover and scale to the moon and back. Emphasis on theoretically because no one seemed to understand how it all worked, so who knows how well it would behave when it failed. There was actually some major downtime in my time there which was rationalized away as "growing pains", but in in my opinion a major factor was just that no one really understood how it al worked.
If something goes wrong with a simple system then often the diagnosis and fix is simple (in this case: just deploy a new environment). If something goes wrong with a complex system then all of this is much harder.
Not saying you can't use these tools correctly or that there isn't an appropriate use for it, but it's a good case study on how the complexity can spin out of control if you're not careful in how you apply it.
He has daily backups for that reason - those can be changed to hourly or even every 15-mins if needed.
I'm surprised that the machines are that large, there must be a bunch of people using the service nowadays :) I remember the setup being even smaller in previous blog posts. But I really like it. Small, easy to reason about. Makes debugging when stuff goes wrong quite easy.
I'm a bit surprised that there's no config management listed though - seems like it really is just Fabric + a bunch of scripts. But hey, if it works, great!
This line here shows much thought has gone into "what if...." situations. Bravo!
I too have a similar setup - I have a contractual responsibility to my clients, and if anything goes wrong with my main equipment, I can resume ASAP. This includes a dedicated smartphone for hotspotting if my internet connection dies.
* the traffic from the monitored systems comes with spikes. Looking at netdata graphs, currently the baseline is 600 requests/s, but there is a 2000 requests/s spike every minute, and 4000 requests/s spike every 10 minutes.
* want to maintain redundancy and capacity even when a load balancer is removed from DNS rotation (due to network problems, or for upgrade)
There are spare resources on the servers, especially RAM, and I could pack more things on fewer hosts. But, with Hetzner prices, why bother? :-)
Personally I've had good experience with Braintree. Particularly their support has been impressively good – they take time to respond, but you can tell the support agents have deep knowledge of their system, they have access to tools to troubleshoot problems, and they don't hesitate to escalate to engineering.
> HAProxy
> NGINX
> SSLMate
with Caddy [0] which would reduce the number of dependencies and complexity involve in infra.
I've been using Caddy for a very long time now and it has been working out so well.
Whereas, with Caddy you would solve all 3 problems with single tool.
Here's a tiny PAAS that can be used if Caddy is too complex.
https://github.com/mardix/sailor
Sailor is a tiny PaaS to install on your servers/VPS that uses git push to deploy micro-apps, micro-services, sites with SSL, on your own servers or VPS.
For SSLMate I hardly understand what problem it solves (even if we forget about having Let's Encrypt free certs).
For SOPs - how it's integrated in your flow.
Thanks in advance.
I'm using both RSA and ECDSA certificates (RSA for compatibility with old clients, ECDSA for efficiency). I'm not sure but looks like ECDSA is not yet generally available from Lets Encrypt.
On sops: the secrets (passwords, API keys, access tokens) are sitting in an encrypted file ("vault"). When a Fabric task needs secrets to fill in a configuration file template, it calls sops to decrypt the vault. My Yubikey starts flashing, I tap the key, Fabric task receives the secrets and can continue.
* On the new host, I run a Fabric task which generates a key pair and spits out the public key. The private key never leaves the server.
* I paste the public key in a peer configuration template.
* On every host that must be able to contact the new host, I run another Fabric task which updates the peer configuration from the template ("wg syncconf").
One thing to watch out is any services that bind to the Wireguard network interface. I had to make sure on reboot they start after wg-quick.
Wonder why he is not using hetzner's load balancer [1]. At least cost-wise, the savings are huge (4xAX41-NVME are around 160€; LB31 - the most expensive one - is 36€)
I'm curious: doesn't Hetzner provide some sort of private virtual cloud? (Digitalocean calls that VPC)
> But, if you want to easily encrypt all traffic, Wireguard might still be a good idea.
But, if one is already using https I guess encrypting the traffic would be unnecessary, no?
I've seen on their custom solutions page [1] you can pay extra to have your own private interconnect or manged switches, but I've never used them nor heard much about them.
[0] https://www.hetzner.com/cloud - see "FEATURES" [1] https://www.hetzner.com/custom-solutions
If I'd run a one man show type of business, I'd love to have some kind of plan B in case I'd be incapacitated for more than half a day. THAT would be an interesting read to me.
A. What if another operator takes the source code, and starts a competing commercial service?
I've seen very few (I think 1 or 2) instances of somebody attempting a commercial product based on Healthchecks open-source code. I think that's because it's just a lot of work to run the service professionally, and then even more work to find users and get people to pay for it.
B. What if a potential customer decides to self-host instead?
I do see a good amount of enthusiasts and companies self-hosting their private Healthchecks instance. I'm fine with that. For one thing, the self-hosting users are all potential future clients of the hosted service. They are already familiar and happy with the product, I just need to sell the "as a service" part.
As a SaaS newbie, is it preferred to have two different email providers for transactional and support email? Why wouldn't one use Elastic Email for support email?
Also, transactional email services police their usage to ensure that they are trusted as sources of legitimate bulk email. It would be hard to get away with sending much transactional email via a normal email provider without having your account shut down.
In contrast. Support email is handled by people and is bidirectional and ad-hoc. Fastmail provides smtp/imap/webmail, spam filtering, mailboxes, etc. and is a good match for handling support.
10 send any due notifications
20 SLEEP 2
30 GOTO 10
The actual loop [1] is of course a little more complicated, and is being run concurrently on several machines.
[1] https://github.com/healthchecks/healthchecks/blob/09a99d3e9c...
For the healthcheck's integration solution (email, signal, discord....), how did you go about doing that? I would love to hear about all that. I did find a section about Signal's integration.
The Signal one took by far the most effort to get going. But, for ideological reasons, I really wanted to have it :-) Unlike most other services, Signal doesn't have public HTTP API for sending messages. Instead you have to run your own local Signal client and send messages through it. Healthchecks is using signal-cli: https://github.com/AsamK/signal-cli
I've automated the mechanics of the failover, but it still must be initiated manually.
I‘ve seen one developer using telegram for ha-failover which is great! Just a message and the scripts are executed for a failover. You can do it in seconds without even being on a computer.
> Any reason not to use a hosted postgres provider?
* Cost
* Schrems II
* From what I remember, both Google Cloud SQL and AWS RDS used to have mandatory maintenance windows. The fail-over was not instant, so there was some unavoidable downtime every month. This was a while ago – maybe it is different now.
Could you also share what is you strategy for marketing and acquiring new users ?
Are "app" servers the haproxy ones or the "web servers", or other ones not mentioned?
Off topic, any pgDash alternatives for MYSQL?
So by the time your browser requests an API call to a Hetzner server, there is a 100~170ms overhead. Compare that to a hosting provider that has datacenters in US West and Asia
The hardware used is consumer hardware too.
This is literally the opposite of what I would expect for a SaaS with uptime requirement... I feel like OP/the developers behind this might have made a really terrible choice and could be paying more in manpower than the savings they have by using a budget server discounter
Don't get me wrong, I have nothing against Hetzner, but this seems like an unfit usecase. This is literally what you use cloud platforms for..
really puzzling that a healthcheck provider itself is not accounting for this potential issue.
Is there some sort of edgy reverse proxy here? Doesn't seem like it. Even then there would be a big latency.
We haven't been using them for anything in production, so I am not 100% sure - they have a backbone outage right now however..
There are no edge / reverse proxies being described in the OP article, just a HAProxy configuration within the same DC.
They also do not have adequate protection systems in place
This caused major problems for Healthchecks, as I'm using Wireguard for private networking, and Wireguard uses UDP. Here's a postmortem about this incident: https://status.healthchecks.io/en/incidents/dcbcDNd89LptHoWi...
Hetzner also offers server-grade hardware btw. But if your setup is fully redundant anyway it doesn't really matter anymore if it's consumer grade or not (which the setup described in the article is).
1) runs single-DC 2) uses consumer hardware
I agree on the notion that with redundancy it is not as important what kind of hardware is used, but it still is something to consider (seeing how this is a one-person venture and server outages are still a huge time-sink), and I wouldn't even necessarily speak about true redundancy in this case
[0] https://docs.hetzner.com/general/others/data-centers-and-con...
Not going to argue about the shared fate, as I've seen more outages than not affecting whole locations..
even better if they offer Singapore & Tokyo.
I run dedicated servers at Hetzner and the consumer parts are CPU and motherboard. Memory is ECC, NVMe SSDs are data center grade. How often do you see CPU failure?
Over the same period, I think there's been two outages of AWS us-east-1 and a multi-region GCP outage. Plus weekly outages of a grownup unicorn SaaS like github or slack.
It's a tradeoff: the simplicity of such a setup reduces outage risk, and the costs saved can be spent on monitoring tools or extra redundancy.
His stack already uses multiple appservers (so the loss of one wouldn't even be noticeable) and has a manual way of failing over to a DB replica.
It's often cheaper to work around node failures (which you often need to do for other reasons anyway - DC-wide outage, etc) than have more reliable nodes.