I can find out what’s wrong within a few seconds.
I can find out what’s wrong within a few seconds.
In my experience alerting often came down to some departments trying to push some shit to some other departments. I personally avoid working in monitoring/alerting for that reason, it's just human problems and dysfunctional organizations, nothing any tool can help with.
It's crazy how easy it was to setup and cover all our infra (AWS, ELB, postgresql, cassandra, kafka, haproxy, nginx, etc...). The tool paid for itself with the infra optimizations we could find in the first month of usage.
It makes me sad when I'm forced to work with graphite/prometheus/grafana in my newer company. These can't gather half the metrics and their charting capabilities are so bad in comparison.
It was also mind-blowing how things were integrated. For example. See a slow request? Click into the APM trace. Notice a service on that trace being slow? Click onto it, see what host it was running on. From there, another button pulls up all the Docker containers running on the host in that point in time. The CPU usage is visualized - and, aha! We forgot to set a CPU limit on one of those other jobs.
Debugging issues like that would've been nearly impossible otherwise, and we had more than a few cases of that.
As far as speed, I haven't had the issue with prometheus. We use recorded rules for things that benefit from being pre-computed.
I imagine the UX to be quite different by using a product.
If you're using clouds (AWS/Azure/Google). Datadog can capture all the AWS metadata automatically and merge with existing metrics, so you use instance tags and such for searching and filtering. It can also capture AWS metrics like ELB and S3 usage which are hard to get otherwise.
So you simply get all the metrics you need and get them easily (I appreciate that people who haven't worked with these probably can't fathom what they are missing out). There are defaults charts/dashboards that are quite good and available out of the box, whereas grafana is empty out of the box and you're once again forced to crawl for dashboard plugins.
Last but not least. The capabilities to search and visualize in datadog are incredible. To draw any metrics and combination of metrics in different ways and analyze usage. Prometheus can't chart shit. Grafana has limited charting and you're forced to create a dashboard to make one chart, which can't be done because don't have admin permissions.
By the way prometheus doesn't scale. It can reach 1000 or 2000 hosts top and that's the end of it. I've operated it at the limit, some operations get really slow and we had to cut down on tags and some metrics to avoid crashing.
Interesting, I haven't had this experience. I monitor the DBs and middleware you mention and the OSS plugins + OSS grafana boards worked quite out of the box. For what is worth we have around ~20 different technologies for DB and middleware.
We aren't using cloud since we have our own datacenters so there could be a big difference in usage.
As far as prometheus doesn't scale I don't know I agree. We have more than 5k hosts currently on it and is working fine. We do use some strategies like recorded queries and federation which are well documented.
The cloud does make a difference. Just seeing the daily S3 usage per bucket was life changing. Immediately found that backups were not expiring after a while as they should, costing more and more money. ^^
Do you know how many metrics you are ingesting in prometheus? storage size? and how many tags per host? We were reaching 1 TB of memory usage (mmap) on our server with 1500 hosts. Prometheus was literally grinding to a halt or crashing, was forced to cut down some metrics and stick to the absolute minimum tags.
Try prometheus_tsdb_head_series and prometheus_tsdb_storage_blocks_bytes or du command on the directory.
It give us a lot of value. We were using other solutions before (sensu/nagios) and it is night and day.
All the plugins that we use are listed in the official Prometheus website: https://prometheus.io/docs/instrumenting/exporters/ we haven't looked outside this page.
--
The cloud does make a difference. Just seeing the daily S3 usage per bucket was life changing. Immediately found that backups were not expiring after a while as they should, costing more and more money. ^^
Yeah this kind of visibility is really lacking on-prems. Storage is something that is hard, at scale, to see what is using what.
--
Do you know how many metrics you are ingesting in prometheus? storage size? and how many tags per host? We were reaching 1 TB of memory usage (mmap) on our server with 1500 hosts. Prometheus was literally grinding to a halt or crashing, was forced to cut down some metrics and stick to the absolute minimum tags.
We have around 500 GB of disk dedicated to prometheus server datacenter with a retention policy of 1 month. The VMs have around 16 GB of RAM.
I would say 95% of the hosts only have node exporter. Then 5% of the hosts will have redis/elasticsearch/postgres/mysql/etc exporters.
Also probably around 5% of the hosts have some sort of custom metrics. We have been leveraging the file exporter for writing custom service metrics.
Custom service metrics goes through a merge request process where we see the best way to structure them to avoid big cardinality. We also use recorded rules to pre-compute things that would be expensive to query (think CPU usage across a datacenter)
1 TB of RAM usage sounds insane. It looks like we can horizontally scale the prometheus servers and use some documented features for that.
For our use case it has been a very smooth ride so far!
I should probably say that my experience with datadog goes back as far as 5 years ago. Already had monitoring working perfectly back then, when prometheus didn't exist let alone the exporter plugins! So prometheus is really late and sub par to me. ^^
Looks like we got a similar amount of data in prometheus as you (1TB for 60 days) but with 40% of the hosts. Maybe you have many small VM? Got physical hosts with quad CPU (per CPU metrics) and network interfaces and stuff (couldn't tune the node exporter to ignore disabled interfaces and some useless devices). Check how many distinct timeseries you have, prometheus_tsdb_head_series.
Datadog had amazing support for custom metrics (but watch out for extra billing and cardinality!). Applications can just send metrics to localhost:1234 where the agent is listening, and they're enriched automatically with host information and environment. Magic.
This reminds me, prometheus is broken with its idea of pulling metrics, when metrics should be pushed instead. Applications and hosts have to push metrics when they come online, it's not the responsibility of the metrics storage to know about every goddamn thing running in the company and try to talk to them (can't cross firewall anyway). Prometheus worked okay enough for the last company that was on premise with fixed hosts (weeks or months to move anything physical), but it's de facto broken for the previous company that was on AWS with instances created intraday.
We disabled a lot of useless metrics in the node exporter. I think pulling works OK if you have a service discovery mechanism. We hook Prometheus to Consul.
We have a mix of very small VMs and very beefy bare metal.
That's all very interesting.
Our data lives on each datacenter only and then we query cross datacenter via grafana when we need.
We run a pair of servers both storing everything. It's the least that can be done to have any resiliency.
I would love to distribute the data, preferably per continent, but prometheus didn't have a good story on sharding. Running independent dataset is worthless in practice without the ability to aggregate. Also, the more servers the more expensive (and they're not easy to procure). Running 6 prometheus servers is in the same ballpark as paying for datadog, so might as well just pay for it.
There are a couple of articles about sharding and federation with prometheus, dunno if they existed when you tried it.
For us our problems are usually local to a datacenter. Having a dropdown where you can pick the datacenter has proven good enough. It is unlikely that we have a global issue in a service.
Sorry if unclear but we have our own datacenters, our prometheus VMs are essentially free in the grand scheme of things considering the number of compute we have.
Their pricing is pretty ridiculous at times and their sales people are often way over aggressive. You have to pay extra for containers on a host. They also make it impossible to keep users from consuming additional paid features.
I like having Datadog when I need to debug, but I'm pretty sick of the dark patterns and surprise bills. I'll probably go with Prometheus in my next greenfield.
I dread having to go back to Prometheus because the company is too cheap to pay for proper tooling (datadog) and developers would rather write their own time series database for their resume.
Disclaimer: I used to work there.
Datadog doesn't do that specific feature as well (it has alternatives), but it also has so many other features that all tie together very nicely: metrics, logging, events, very good dashboards, analysis notebooks, alerting, SLOs, performance monitoring, trace analysis, security monitoring. It's a really extensive product.
When it comes to general observability, I'm a strong believer that you need a wide range of different views – just logs aren't enough, just metrics aren't enough, etc. I've worked in a team trying to use Prometheus for everything and there was so much friction, whereas with Datadog there has always been a way to achieve something.
I think Honeycomb is a good feature that should be bought by a company like Datadog and integrated into a wider more mature feature set.
I’m somewhat familiar with monitoring/telemetry but haven’t heard that term before.
For example, if you have a log line to represent an HTTP request completing it might have a server hostname that processed it, a path, and a request duration – 3 dimensions. High dimensionality is just having lots of dimensions.
Most monitoring systems aren't great at correlating between lots of dimensions, or cost a lot if you want to high cardinality, which high dimensionality contributes to. The cardinality is how many possible options there are for all of the dimensions together (so if you have 3 paths and 2 servers, those two fields have a cardinality of 6, or 6 different places you need to store your request duration for).
When you start getting lots of dimensions and lots of values for each dimension, things start getting expensive and you have to be quite restrained about what you choose to monitor. Maybe you decide not to track duration per server because you hope your servers are roughly the same, whereas you know that URL paths perform differently much of the time.
Honeycomb's great advantage is that they support this high dimensionality/cardinality really well. Not only do they not cost a lot to do it, but they also have a really nice UI for exploring this data and how different bits correlate together, without you having to know up-front what you want to look for. Honeycomb is expensive, but not prohibitively so for some engineering teams.
I'd still recommend Datadog _first_ because it does a lot more for a little less money, but if you're pushing the boundaries of what Datadog is capable of with tracing through distributed systems then Honeycomb enables the next step.
iOS: https://apps.apple.com/us/app/datadog/id1391380318 Android (Play store): https://play.google.com/store/apps/details?id=com.datadog.ap...