We send all monitor metrics to DataDog. When a monitor fails, the appropriate teams will get a Slack notification with the full stack trace. A DataDog monitor will also be triggered, alerting the appropriate teams.
For browser monitors, we upload screenshots and Puppeteer tracing files to S3, then share links within each Slack hook. This allows people to figure out what's going on just by clicking links in Slack.
We were planning to improve this setup in the future, but it's good enough for us right now. For example, CircleCI goes degrades frequently so we sometimes get spotty coverage. We basically spend < $200/month with CircleCI to monitor about 300 APIs/pages every minute.
You can read more here:
- https://engineering.dollarshaveclub.com/monitor-all-the-thin...
- https://circleci.com/blog/how-dollar-shave-club-3x-d-velocit...
I'm not affiliated with them, just an happy customer
[0] E.g. status page: https://status.appdrag.com/
I facepalmed.
But once my setups get a bit more complex I was thinking I'd build webhooks into an analytics server, and ping those from each server with a json request which includes health data for each's databases and other servers.
I'm no where near that so I'll probably sign up for the free plan on uptimerobot later mentioned in other comments.
Step 2: After a few months, repeat step 1.
Then, I monitor just that one endpoint with Stackdriver (because it's easy). If any of the checks fail, it logs it, prints details, and sets a 500 header code. Adding new checks is just a code change.
I use a custom AWS Lambda function. It fires every four hours or so, and tries to make an https connection to each configured URL (the URLs are stored in a file in S3) and if the site is either down, or if there is an SSL error (which probably means an expired certificate) then it sends me a text message using SNS.
The whole thing is about 50 lines of code, and that's in Java. And it doesn't even come close to exceeding the free tier limit of Lambda calls, so it doesn't even cost anything so far.
To be fair, I could have used a 3rd party service, but writing this thing was my first foray into using Lambda, so I did it as much for the learning experience as anything. But it works really well, so I doubt I'll replace it anytime soon.
One of the sites is a SaaS offering, but it's not live yet, so I don't need to stay super-on-top-of-it. Once it's live we'll want more frequent monitoring and some other stuff, so we might either move to another approach, or supplement this with something else.
https://hackernoon.com/build-an-uptime-monitor-in-minutes-wi...
Another option if you don't feel in the mood for DIY is TJ Holowaychuk's Apex Ping: https://apex.sh/ping/. Great service, run by a solo developer, reasonable price.
Disclaimer: I founded Standard Library. :)
We haven’t historically had a problem with “stdlib”, we’re already the top Google result. “Standard Library” (full name) is new for us as we expand to a less technical cohort of customers. We’re working with some pretty great people and companies (Stripe, Slack) on our mission to build a, well, Standard Library — so if you can get over the name choice you should check out our online development environment! -> https://code.stdlib.com/
Still waiting to see if the CA Technologies acquisition [1] makes things worse or not.
[0] https://www.runscope.com/ [1] https://blog.runscope.com/posts/301
There are more resources working on Runscope than ever before. CA continues to invest in stability, new features and support. CA is also going through an acquisition and that introduces more variables, but as of today (1+ year after acquiring us), CA has been extremely supportive of Runscope.
Protectumus monitors the website uptime, speed, dns changes, scans the website for malware like a traditional antivirus, blocks bad bots, custom IP's and countries. Protectumos acts as a Firewall. We are the only security company specialized in SEO Security, we offer unique SEO services such as Search engines Cloaking monitoring, Google DMCA Complaints, Blacklist monitoring & removal and more.
Details here - https://news.ycombinator.com/item?id=18295381
Full disclosure: I'm the founder.
> Protectumus acts as a web application firewall (WAF) and scans the website for known malware. Once the malware is found it will be automatically removed.
How does a WAF remove malware? Is there an additional agent sitting on the server or something? And surely when a server is pwned, it needs to be reformatted, not just the malware 'removed'.
The firewall blocks bad bods, hack attempts like SQL injections, XSS, CSRF and more.
The antivirus scans for known malware (we have a big list of malware definitions), but we use AI and Machine Learning and the Antivirus learns from previous detections and is able to act alone. The script is able to automatically remove the malware once it is found.
It's not perfect but I hardly have issues with it.
Honestly, I don't do expiry checks...
For alerts we have a mix of grafana and AlerManager.
Also Prometheus only saves 15 days worth of data by default and since they aren't willing to develop down sampling, it's not good to keep any long term history of the metric but you can overcome this by using InfluxDB as the storage but I would rather use something like monit which is simpler and easier to setup and gets the actual job done well.
I use StatusCake for basic uptime monitoring for websites. I switched to them from Pingdom because they are cheaper. Only downsides I've had with StatusCake is if something is down it doesn't give you the cause. Pingdom would show you the trace route. That has made it hard to tell and I would get quite a few false positives saying sites were down. I haven't had false positive issues for months now though. Their paid plans have SSL/domain monitoring.
Monitis for cheap Linux CPU/RAM/Load monitoring.
Gives me all sorts of useful information that I use to make decisions.
SSL certificate expiry is easily checked with a nagios check (use the -C flag on check_http). If you use Letsencrypt with a client like acmetool (https://github.com/hlandau/acme) your certs will never expire. Of course the nagios check is still necessary to ensure acmetool keeps doing its job!
Domain name expiry checking could also be a nagios job, or alternatively you could write a small script that checks whois output and execute it regularly with cron.
Configuring your registrar for auto-renewal helps avoid a certain class of errors ("I forgot to renew!") but not others ("my credit card expired and the e-mail notifications from my registrar didn't reach me").
I wish the documentation was better but telegraf's documentation is light years ahead of collectd, which is similar software.
Kapacitor needs some more examples and the default Chronograf-generated TICKscript needed to be thoroughly modified to meet my needs. It took me way too long to figure out how to use stateChangesOnly() to prevent me from getting constant notifications once something went into an alarm state.
That said, the stack works well, even if it has a few rough edges. Thanks to Influxdata for the open source stuff. High quality open source software makes me want to endorse them and purchase the paid products.
I also created a very hacky, browser start page which has GIF's that are pulled from each of my servers. Sort of like how "game copy world" used to setup their mirror page. Super unsophisticated but it allows me to do an "uptime check" every time I open my browser.
If you drop me a line after you signed up I can flag you as a demo user that's free forever - or at least until you want to pay or cancel :)
Re SSL expiry: we've taken the route of running VirtualMin on our 'commodity' servers, sites setup in there have an auto-renew policy if they use SSL, so it requires no action and auto-updates whenever certs expire (IIRC it's every three months? but I might be wrong: it requires no thought nor action from me).
For domain names, we use Joker (Swiss company), we've found them to be good in all aspects, and they send renewal notices well in advance (including a link that anyone can use to renew the domain, even without having an account with them).
Uptime monitoring is a whole different kettle of fish, and we manage it on a per client basis depending on needs. For the sites we host, generally if the server's up, then the sites are up — but we also do more specific monitoring if it's a client requirement.
Feel free to let me know what you think about it, feedback is always greatly appreciated.
Disclaimer: I'm the developer.
For example, domain expiry - I have a script that Zabbix runs once a day that does a whois and grabs the expiry date for the domain in question. Convert that to a unix string and subtract the current date on it. Echo that from the script.
In Zabbix we can now alert if that item's value is <30d. Similar for SSL certificates and the web monitoring stuff is built in
Edit: Oh actually on the whois, I remember it was a huge pain in the butt getting the expiry for a variety of different domains - I now use https://jsonwhoisapi.com/ to get the whois info
Site monitoring: NewRelic
PagerDuty/OpsGenie: For alert routing if you have more than 2 people.
- Pingdom: polls custom endpoint every minute. Sends notification to PagerDuty if critical, or if warning for more than 5 mins.
- PagerDuty to notify team
- SSL Expiry: calendar notifications (whole team) and reminder emails from our SSL cert issuer
It ended up being a great tool for me because it allows for the most basic ping test and to content checks to be setup in 10 mins. It also has the ability to add reporting based on apps, servers, and databases. The aws add-ins helped me tune my usage, as i was paying way to much for a couple services that i could downgrade with out impacting my apps.
I feel it is priced right and a good value up to the 89/month plan.
Full disclosure: I'm the founder.
It can also alert if the SSL cert is within X days of expiring.
https://www.poweradmin.com/help/pa-server-monitor-7-1/monito...
It can use the certificate transparency (CT) logs to detect new certificates for your domain, so once set up you don't have to maintain it. Make sure to enable the weekly report email too.
For backend and frontend monitoring: https://atatus.com
What are free solutions?
One is Google Docs. What else?
SSL & domain name on auto-renew. SSL via lets encrypt with a renewal cron and domain is set to auto-renew in domain registrar dashboard.
A ton of different types of checks. A lot of value for not a lot of money. The public status page style is minimalastic but exactly what I want.
Example: https://status.regexplanet.com/
https://github.com/GoogleChrome/puppeteer/blob/v1.9.0/docs/a...
Best value of any tool I've ever used. It does literally everything you asked. I didn't even know it checked SSL expiry till it pinged me.