Though you could get most things into Grafana with something like Prometheus. The problem with Grafana is understanding what the limitations are. If you're not careful with the number of panels and such it can get quite slow.
I've used Grafonnet before for doing Grafana at scale. Simply put, I hate it. Apparently an alternative is being worked on at Grafana so I'm waiting for that. But if you need to make hundreds of panels....it works well enough.
If you need to monitor some infrastructure you can just use Telegraf and output it to Grafana if needed. It kinda falls apart though because another great benefit of something like Datadog is not managing a time series db. That can get ugly real quick.
I guess it all just depends. If my bill was super high I wouldn't mind spending some resources on Prom/Grafana if you're in the Kube space or some Telegraf/InfluxDB if you're not.
I've also heard good things about Timescale but haven't used it.
Running it yourself is not too hard up if you are not having to do clustering ( say 1m metric series, 100GB/day logs). But different people have different comfort levels for that.
With any monitoring system most of the work is actually making use of the data. Tagging, Alerts, Dashboards and especially onboarding all the teams. You can spend a lot of time and money rolling something out and then barely anybody uses it.
Hi, I run the Grafana team at Grafana Labs. I'd love to learn more about your Grafonnet use to help us build something better. I'm david at grafana com
Best of all, can also be run entirely within your own AWS/GCP/Azure so you only pay OpsVerse for maintaining the stack based on your ingestion (and we also monitor the monitoring system for you ;))
You would need a team to configure this setup and make it right over time. It's worth the investment instead of paying a cloud market leader.
For others reading this - you can’t just switch back and forth a few times a week. A full platform user can be moved to a basic user only twice in a 12-month period.
The notion that every single log or metric across your entire technical architecture is worth keeping is one implanted by SaaS providers with a financial interest in naive engineers doing just that.
We have concepts of debug, info, warn and error… but I think we need apps to be developed with the concept of log concern.
For example take sshd. For infosec they are interested in IP of failed attempts, operations might want to know connection failures, etc.
Note: in qryn s3/r2 are as close to /dev/null as it gets!
Unfortunately, avoiding insanely costly SaaS solutions requires engineers to plan ahead and design the entire stack on top of certain open source solutions. I suspect that many engineers today receive kickbacks from SaaS providers to lock-in their employers. Employers are none the wiser and rarely push back when an engineer suggests a big-name SaaS solution with insane lock-in factor. Nobody seems to care about lock-in these days, it's only when your costs reach almost 100 million and interest rates are going up that you start thinking "Damn, I could have had all that for free if I had planned ahead and resisted all these platform lock-ins and unnecessary proprietary tools..."
Cmon man, really? Drop the conspiracy theories. I’ve personally been the guy advocating for datadog at 4 startups. Mainly because of opportunity cost - we have 10-100 engineers, I want them building product not figuring out how to deploy a whole ecosystem of observability tools. IF we get big let’s reevaluate… but in the meantime. am I doing it wrong? If others are getting kickbacks I want in
Imagine being the guy who convinced Coinbase to use DataDog... That person will probably end up working at DataDog sooner or later if not already there... You can bet they will be getting a very cushy salary.
I could probably make a living out of extorting corrupt engineers. It's so predictable.
Why don’t you walk the walk instead of merely talking the talk.
They wouldn't hire someone dumb enough to spend $5M a month on their product.
The difference between datadog and doing it yourself is that datadog is a well thought through product rather than a cobbled together set of various tools
Having a single interface for everything makes life so much easier across a number of different teams
Search is fast and easy to use for logs and traces
Being able to see what a user actually clicked on in their session is absolutely game changing for support teams
I’m not a huge fan of the bill but it’s so much better than anything we could do ourselves without a team of engineers dedicated to observability (which would cost far more than datadog)
You should have sorted it out before the integration, not after. Now you have no leverage.
If you change your mindset from ignoring problems to looking for problems, you will find that there are problems everywhere. I'd rather be biased in that way than in the former. In my position, I can't afford to ignore even the tiniest problems.
Moving away from those SaaS tools can be extremely painful and a lot more costly due to vendor lock in. In practice, typically, this "let's reevaluate" time never happens.
On the other hand, I don't really care. I normally suggest open source tools, but if people want to throw money at some vendor, fine by me.
I love how good DataDog is. It's a great product. Too expensive though. I love most of the people I've worked with at Grafana Cloud but it's a painful product. The price makes up for it though, so we use Grafana Cloud.
We may end up with something like signoz, when we have the cycles but the ROI is bad when I already have twice as much work as people and that barely more than KTLO.
Hi, I run the Grafana team at Grafana Labs. If you could fix one thing, what would it be?