A cryptocurrency company had a $65M bill, per Datadog’s Q1 earnings call
twitter.com
twitter.com
Then, when my invoice isn’t paid, I threaten collections on them personally and the company. Usually that solves it. Then I’m “such a dick” but highly effective in recovering my time.
"Hi this is Arnie from CHewemup'n'Spitemout Staffing, is this Bob?"
"Hey Arnie, what perfect timing! I just started looking for a new opportunity, and I'm really excited to— CLICK."
Ring ring.
"This is Arnie, we seem to have been cut off."
"Oh Arnie, right, thought you hung up on me."
"No, not me, must be a bad connection. You were saying?"
"Yes, I was saying this is a great time to talk about opportunities. I just finished a major Java Enterprise JavaBeans project, and I'm— CLICK."
Lather, rinse, repeat, as the meme used to go before we called them memes.
https://jollyrogertelephone.com/
I also will give out the number to local strip clubs for sales people which seems to be confusing to them.
Honestly if you want to complain at the cost of hosting (which, we don't know if they would) then licensing the software and allowing people to self-host would be the solution.
$65M is enough that I could fund a team running google's monarch system for 7-8 years.
1. use Datadog, because it gets you a bunch of stuff without having to really set it up, like anomaly detection, which is poor man's monitoring & alerting
2. once you start getting product-market fit and the number of instances you run grows you notice your monthly bill going crazier and crazier and you now have something that starts to resemble an ops team -> migrate to a different product, set up proper monitoring
I've seen that at 3 companies I worked at.
> and you now have something that starts to resemble an ops team
Like what if Datadog just replaces your ops team completely? What if we start to see AI tools that do cost a lot of money but they can replace a team? Just curious.
PS: I am one of the maintainers at SigNoz
it currently supports logs, traces, session replay/RUM, alerts and dashboards - would love to hear what you think :) https://www.hyperdx.io
I asked them how them can justify that.
They recommended I use "modern infrastructure" which means Docker.
It looks to me they recommended to never consider Datadog for anything whatsoever, ever again.
Also: it’s 2023. Every company needs to be getting compatible with open standards like OpenTelemetry, Prometheus, etc.
I was never sure of what exactly Datadog did, so I looked at their pricing. At first, I thought "$23/month/host isn't THAT bad...", then notice that was only for one product.
If you used their full suite, those costs could REALLY add up.
Every year, I get a barrage of phone calls and emails from them, and I actively choose to not engage.
Pricing has remained pretty much flat aside from the expected growth. They work hard to get you to overcommit and overpay.
Also, you don’t need all your logs indexed. Saved a company I’m under contract with a massive amount of money (10s of thousands/month) by pointing out that you should just index (sample) a percentage of them to identify if there is a trend, and you can rehydrate later.
We reached out to DD to try to work something out, and they agreed to cut the bill in half, but only IF we signed up for additional services.
I will never use their services again, and will always share this story when their name is brought up.
Harassing engineers is a hard no.
What are some of the best ways to ensure the folks developing business leads don't sell out your future for their present?
What are the best incentive structues you can provide so they don't want or need to?
The idea behind this are a few-fold but essentially:
As an engineer (now CEO/CTO), I've hated having to wait the full year for my bonus. It's just a way to lock me in for the year when my incentive to stay should be to love the work & team. I don't want to create a place to work where you're forced to stay because of some guaranteed bonus - if you want to leave, leave & then let's hire someone who finds the work engaging + we all know performance slips as you wait for the bonus.
For the sales team, it means they're incentivized to work with the engineering & product teams to make sure they get the engineers the proper feedback in order to build a better product that they can sell more easily.
We've found this has generally built a better more team-oriented & results-oriented culture. Happy to expand but overall I think a quarterly revshare for everyone is a much better end-result (other than the fact I'm now forced to care more about making sure engineers are happy but that should be a huge focus regardless...).
Edit - also worth noting that we give everyone equity so there's still a long-term focus of building a company, not just cashing out quickly.
I think their revenue says it is, unfortunately.
We decided to never work with them after that. That and the horror stories of people that experienced the product itself in their former companies.
Because that's what happened.
I am a user of New Relic. Not because I'm happy. But because OpenTelemetry doesn't come close to the same features. Fortunately, at least OT is about 10x harder to set up with worse documentation.
Wait a minute...
New Relic is really really good too, so it was even more painful.
There are a lot of subtle issues, but we've been able to work through them to get usable traces. (Metrics and log ingestion already have pretty good existing open-source tooling, like statsd).
Incumbents in this space are in for a rough time as more applications provide meaningful telemetry, beyond just logs. Fortunately for them that timeline is 'fuzzy' at best.
I agree it was a bit rapidly evolving in early days, but now its much more mature.
You can check out our docs for distributed tracing here - https://signoz.io/docs/instrumentation/
I see cloud costs like this a lot and it really puzzles me. It seems like people would rather pay 10X+ more to just not have to think about it than even to hire other people to think about it, because then you have to think about hiring and HR.
"Here's a blank check. Just make it go away."
Of course corporate consultants run on that, so I guess it's not without abundant precedent elsewhere. I guess if you work for a big company with budget and it's not your money you really have little incentive not to take the easy path.
Does seem like a pretty wild bill, thou.
Take for instance GDPR - in AWS it was a company wide effort to get all the services GDPR compliant and that was basically a non-existent pricing change to consumers.
Also the fact that I can call up AWS support and have them look into a bug immediately with real devs on the other end is invaluable when my business needs rely on a certain feature working.
That takes calendar time.
There's a multiple to what people are willing to pay for SaaS solutions precisely because HR for knowledge work is such a pain.
OTel _does_ prevent lock-in on the agent side (making it easier to switch vendors) with open source components and consistent schemas, but OTel doesn't enable you to do anything that you couldn't do before with a specific vendor. It just empowers you to take ownership of your observability data, should you want to. Many don't, though. They want to throw money at someone else who can do it relatively well, hence the ridiculous DD pricing.
Regarding:
> Not sure how companies rationalize these types of services at scale when there are so many open source options to run for a fraction of the cost
At Coinbase's scale (which I assume is a lot due to the DD bill, but I haven't looked closely at it) the open source options simply won't cut it. Plus there are no scalable open source options for many of the things that DD does (synthetics, SIEM come to mind - not to mention onerous regulatory requirements). 65M seems like a lot, but it also means their cloud costs are insane - so maybe it lets them put focus elsewhere?
This is exactly the reason why we are moving away from NewRelic's SDK to an OpenTelemetry SDK (even though we are still using NewRelic to ingest everything). If (more like when) we decide to switch vendors, it will be much easier to do so.
Datadog specifically, I don't know if they care. They had an army of junior engineers that existed to hack Datadog into every open source project imaginable. If something had monitoring, they would just add their vendor-specific stuff and upstream it. That was probably expensive but probably accounts for a vast amount of their early marketshare. The other vendors wanted in on that racket without having to do too much work; OTel was born.
In theory you could use the Otel Collector (or any other Otel-compatible agent) instead of the DD Agent to collect metrics/logs/traces. This would then make it easier for you to switch from DD to another Otel-compatible provider (Grafana, for example)... but 99.9% of what DD provides is _not_ the agent, it's dashboarding, alerting, RUM, synthetics, etc.
Basically Otel has made _agent_ switching costs effectively drop to zero, but that is a very small part of the whole picture. Like I said above, this primarily hurts vendors with proprietary agents that can't/won't adopt Otel for ingesting data.
As a vendor building in this space [1] - it definitely is. We're able to onboard teams faster to do side-by-side comparisons because they can simply point their existing Otel telemetry to both us and their existing provider with just a few lines of config. That wasn't possible before otel, and levels the playing field more than before. As otel matures, it'll continue to erode against DD's position.
It also allows us as a company to focus on what users care about (as you mention that's things like dashboarding, search performance, RUM, etc.) as opposed to spending all our time building basic integrations into every platform (though we still do plenty of work to polish places where Otel hasn't). Again, levels the playing field.
As mentioned in some other places in the thread, DataDog pricing is very unpredictable and high - and I think more open standards based solutions are the way forward which provides users more predictability and flexibility
They paid upfront for 3 years of usage, and yes they were burning > $20m/year on datadog
The headline seems misleading, like it’s for a single month.
20 fucking million
getting splunk with 1pb a month ingest isnt that expensive.
or something with chatgpt
Earnings call said it was an upfront cost
My favorite "overpriced support contract" was for an Oracle product. The cost of support was seven figures, and in the entire year a single phone call had been placed to support.
1. Can we use MySQL for <new product>?
No, use Oracle, we have a support contract
2. <oracle related issue occurs> Can we call that support contract in now?
No, let's try the inhouse expertise first.
3. <inhouse expertise comes up with barely passable hack> Can we check if they have any better solutions?
No, it's "solved" now
Like, what were we paying for? I have to assume there's per-engagement costs as well as the ongoing costs, given how hestitant our contract owning team were to let us anywhere near Oracle.
Okay, it seemed like it was 1 year.
So, only insane, not insane^2.
If you're considering using them keep that in mind and tbh I would strongly recommend considering if CloudWatch or some other cheaper alternative is suitable for your needs.
At a previous company, we had DD, and I got asked to find some problem. I was able to sift through the data in it and zero in on the instance and find the bad data that came in.
Then it got expensive and so they turned on 'sampling' and the next time I was asked to look for a particular problem, we had no idea if it was even logged.
Our company helps avoid these kinds of observability bills and issues like scaling for fast-growing cloud deployments. Generally speaking, many vendors let you fall into the cardinality trap b/c they have an economic incentive to let you do so. One of our biggest selling points is that we provide an observability control plane that helps drill down into wasted queries, shows how metrics can be aggregated, and other ways of avoiding wasted cost/effort. tbh no one should have to pay more to observe a service than to operate it. Where’s the ROI in that? Another plus is that we're all in on open source instrumentation with OpenTelemetry & Prometheus so none of that annoying vendor lock-in.
“We’ll show you how to make sure you don’t have even one crystal fall off the plate.”
My personal pet peeve is Azure Application Insights which uses Log Analytics under the hood… at a rate of $2.75 per ingested GB of logs stored for one month. That’s highway robbery.
Let that sink in: They charge $2,800 to store a TB of text that takes a few hundred dollars of overpriced cloud disk and maybe $10 of CPU time for the actual processing. That’s the cost of a serviceable used car or a brand new gaming PC!
But wait! There’s more.
In reality that 1 TB is column compressed down to maybe 100 MB, making it about $30K charged per terabyte stored on disk.
It doesn’t stop there! Thanks to misaligned incentives, the ingested data format is fantastically inefficient JSON that re-sends static values for every metric sample collected. Why would anyone ever bother to optimise their only revenue?!
They won’t.
The reality is that a numeric metric collected once a second (not minute!) is just 21 MB if stored as a simple array. Most metrics are highly compressible, and that would easily pack to 100 KB per metric per month.
A typical Windows server has about 15,000 performance metrics. We could be collecting these once a second and use a grand total of… 1.5 GB per month. That’s every metric for every process, every device, every error counter, everything.
Modern server monitoring is inefficient and overpriced by 5 orders of magnitude. It’s that simple.
That fact that your company can exist at all is a testament to that.
Totally agree about the compressibility of metrics and toying with the scraping interval. I started out working for an enterprise monitoring vendor that had a proprietary agent that already decided sane intervals to emit metrics, when I learned that Prometheus let users configure that to me...just sounds like an expensive foot gun.
My real beef with metrics is at least for app layer insights is the waste. I'd so much rather have a span/event configured with tail sampling so you can derive metrics from traces and tie them to logs in a native contextualized way vs having to do that correlation on the backend and within different systems and query langs. Seems much more efficient and cost-effective that way, I'm scarred from seeing a zillion "service_name.http_response.p95.average" metrics that are imo useless
I’m starting to come to the same conclusion, but the point I’m making is a general one: efficient formats would allow finer grained telemetry to be collected without having to be tuned and carefully monitored.
What’s the point of a monitoring system that itself needs baby sitting?!
I am one of the maintainers at SigNoz. We have come across many more horror stories around Datadog billing while interacting with our users.
We recently did a deep dive on pricing, and found some interesting insights on how it is priced compared to other products.
Datadog's billing has two key issues:
- Very complex SKU based pricing which makes it impossible to predict how much it would cost
- Custom metrics billing ($0.05 per custom metric) - we found that custom metrics can account for up to 52% of the total billing which just does not make sense
More details in the blog here with a complete spreadsheet for detailed calculation https://signoz.io/blog/pricing-comparison-signoz-vs-datadog-...
This may come as a surprise but when giving money to a for-profit company, not only are you paying for corporate bloat, but you're also paying for the CEO's lavish compensation package, free lunches, and very costly health insurance for their employees. You're even paying for employee salaries while they're not doing work while on vacation!
If that's a big problem for you, Graphana may be the better product for you.
Corporate bloat is like the mini empires people build, headcount for the sake of headcount, that guy who's been here forever passion project that doesn't make money. Process because it helped someone's resume. Those kinds of inefficiencies. This stuff is different then treating employees nicely.
The notion that every single log or metric across your entire technical architecture is worth keeping is one implanted by SaaS providers with a financial interest in naive engineers doing just that.
We have concepts of debug, info, warn and error… but I think we need apps to be developed with the concept of log concern.
For example take sshd. For infosec they are interested in IP of failed attempts, operations might want to know connection failures, etc.
Note: in qryn s3/r2 are as close to /dev/null as it gets!
Unfortunately, avoiding insanely costly SaaS solutions requires engineers to plan ahead and design the entire stack on top of certain open source solutions. I suspect that many engineers today receive kickbacks from SaaS providers to lock-in their employers. Employers are none the wiser and rarely push back when an engineer suggests a big-name SaaS solution with insane lock-in factor. Nobody seems to care about lock-in these days, it's only when your costs reach almost 100 million and interest rates are going up that you start thinking "Damn, I could have had all that for free if I had planned ahead and resisted all these platform lock-ins and unnecessary proprietary tools..."
Cmon man, really? Drop the conspiracy theories. I’ve personally been the guy advocating for datadog at 4 startups. Mainly because of opportunity cost - we have 10-100 engineers, I want them building product not figuring out how to deploy a whole ecosystem of observability tools. IF we get big let’s reevaluate… but in the meantime. am I doing it wrong? If others are getting kickbacks I want in
Imagine being the guy who convinced Coinbase to use DataDog... That person will probably end up working at DataDog sooner or later if not already there... You can bet they will be getting a very cushy salary.
I could probably make a living out of extorting corrupt engineers. It's so predictable.
Why don’t you walk the walk instead of merely talking the talk.
They wouldn't hire someone dumb enough to spend $5M a month on their product.
The difference between datadog and doing it yourself is that datadog is a well thought through product rather than a cobbled together set of various tools
Having a single interface for everything makes life so much easier across a number of different teams
Search is fast and easy to use for logs and traces
Being able to see what a user actually clicked on in their session is absolutely game changing for support teams
I’m not a huge fan of the bill but it’s so much better than anything we could do ourselves without a team of engineers dedicated to observability (which would cost far more than datadog)
You should have sorted it out before the integration, not after. Now you have no leverage.
If you change your mindset from ignoring problems to looking for problems, you will find that there are problems everywhere. I'd rather be biased in that way than in the former. In my position, I can't afford to ignore even the tiniest problems.
Moving away from those SaaS tools can be extremely painful and a lot more costly due to vendor lock in. In practice, typically, this "let's reevaluate" time never happens.
On the other hand, I don't really care. I normally suggest open source tools, but if people want to throw money at some vendor, fine by me.
I love how good DataDog is. It's a great product. Too expensive though. I love most of the people I've worked with at Grafana Cloud but it's a painful product. The price makes up for it though, so we use Grafana Cloud.
We may end up with something like signoz, when we have the cycles but the ROI is bad when I already have twice as much work as people and that barely more than KTLO.
Hi, I run the Grafana team at Grafana Labs. If you could fix one thing, what would it be?
Though you could get most things into Grafana with something like Prometheus. The problem with Grafana is understanding what the limitations are. If you're not careful with the number of panels and such it can get quite slow.
I've used Grafonnet before for doing Grafana at scale. Simply put, I hate it. Apparently an alternative is being worked on at Grafana so I'm waiting for that. But if you need to make hundreds of panels....it works well enough.
If you need to monitor some infrastructure you can just use Telegraf and output it to Grafana if needed. It kinda falls apart though because another great benefit of something like Datadog is not managing a time series db. That can get ugly real quick.
I guess it all just depends. If my bill was super high I wouldn't mind spending some resources on Prom/Grafana if you're in the Kube space or some Telegraf/InfluxDB if you're not.
I've also heard good things about Timescale but haven't used it.
Running it yourself is not too hard up if you are not having to do clustering ( say 1m metric series, 100GB/day logs). But different people have different comfort levels for that.
With any monitoring system most of the work is actually making use of the data. Tagging, Alerts, Dashboards and especially onboarding all the teams. You can spend a lot of time and money rolling something out and then barely anybody uses it.
Hi, I run the Grafana team at Grafana Labs. I'd love to learn more about your Grafonnet use to help us build something better. I'm david at grafana com
For others reading this - you can’t just switch back and forth a few times a week. A full platform user can be moved to a basic user only twice in a 12-month period.
Best of all, can also be run entirely within your own AWS/GCP/Azure so you only pay OpsVerse for maintaining the stack based on your ingestion (and we also monitor the monitoring system for you ;))
You would need a team to configure this setup and make it right over time. It's worth the investment instead of paying a cloud market leader.
A lot of this discussion reminds me of this talk:"Netflix built its own monitoring system - and why you probably shouldn't" (https://www.infoq.com/presentations/netflix-monitoring-syste...) where Roy Rappport describes Netflix as a "monitoring system that happens to stream movies"
As someone who spent a few years at New Relic and Lacework, I can also say that pricing observability fairly is crazy hard when you account for different architectures, usage pricing, and the humans experience the value.
We attempted to migrate from datadog to prometheus at GitHub and that stack did not cover our use case at all. So much tooling had to be recreated. I took a lot of flak when I pointed out numbers made sense to stay on DataDog and migrate to a Microsoft product instead, but the cost savings spoke for itself
It depends what you're looking for: metrics, logs, APM, tracing, synthetics, web analytics, etc.
That said, managing observability yourself should result in <5% of cloud spend. So I'm figuring someone at Coinbase said "WTF" to this bill and migrated to Grafana/Loki or Kibana/OpenSearch or Kibana/Elastic. Well, that, and Coinbase's business also dropped off a cliff. Combined, I could easily see a one-time influx of $65M from one customer, gone the next quarter.
Had a whole team of 10+ engineers working on it for 2 quarters, then scrapped it because it performed terribly
The only thing that came of it was negotiation leverage with datadog ("give us X% off or we go self hosted")
How much logging are we talking here?
DD for letting an obviously huge and important client run up a bill so crazy they have to quit DD out of shame and governance.
CB for demonstrating a complete failure of financial management, vendor management and any sort of ability to track expenses.
Doesn’t help either side to get to a point where you have to quit to demonstrate you’re not totally incompetent and crazy.
> Datadog is an observability service for cloud-scale applications, providing monitoring of servers, databases, tools, and services, through a SaaS-based data analytics platform.
So it checks if your servers have crashed or slowed down with a nice dashboard?
Any better summaries or descriptions of what it does and how coinbase would have used it?
They also have quite an unethical sales operation: https://news.ycombinator.com/item?id=35837965
And with a healthy dose of blue cross bolted on for the surprise bills and difficult bureaucracy.
But it was super tied to VMs at that point, and we were running a bunch of lambdas, herokus and docker instances, along with a shit tonne of AWS services, and java lumps from the 90s
My team use Grafana’s open source LGTM stack. We use Prometheus metrics to track anything from JVM/Go runtime stats, K8S metrics, saturation of CPU/memory, scalability issues, crashes/OOMs, custom metrics for business insights, debugging. We use USE/RED metrics (see: Google’s SRE handbook) to track our production services performance in an objective way. We track SLAs and SLOs so we know when it’s time to focus on features and business impact, and when it’s time to put that aside to focus on stability and maintenance before our customers notice reduced reliability.
As a developer it’s really helpful for testing changes. For example, I added a new database index in dev, then run some load tests and check our dashboards before and after. I look at Q95 latency of APIs and database load to see if it has the desired effect, then when I roll out to production I can monitor those same dashboards and make sure the same desired improvement can be seen for real-word usage.
I used traces recently to discover that something that should have been happening in parallel was instead happening sequentially leading to very long/timing out requests. Adding visualisations via traces helps get your head around how something is working.
I added annotations to our dashboards that shows when our K8S pods restart alongside the metrics. This made me realise that some requests were failing exactly around deployments because we were not cleanly handling SIGTERM in some services.
We have started adding horizontal auto scaling based on metrics for the number of queued messages on a specific Kafka queue. If a large number of messages are waiting we spin up more K8S replicas, and then once this reduces, we reduce the replicas to keep costs down.
I optimise the resource allocations on our services by looking at historical CPU/memory usage so we make the best use of our K8S cluster and avoid OOMs as we scale.
We use Loki for log querying and parsing, you can create really advanced/domain-specialised log querying dashboards and provide that to your support team, and integrate those logs with traces to debug different stages of a request as it traverses your microservices or different processing stages.
You can even build dashboards from logs, which is helpful when debugging a particular type of error over time that you were not specifically monitoring with metrics, or determine which customer(s) are affected by this error. Alternatively if you have a legacy system that does not have effective metrics, you can build metrics from its logs.
We use our metrics for alerting and paging in a way that provides a better signal-to-noise ratio than old-school alerts like “high memory usage” so people don’t get woken up as much (we’ve had zero pages since my product launched 6 months ago!). It’s better to alert only when we have a measurable impact on customer experience, like when a smoke test has failed more than 80% of the time, or HTTP requests 5xx rate is elevated to abnormal levels.
It’s also really reassuring when you do a prod rollout to easily see that stuff is still working without digging into logs, so you can spend more time coding and less time babying prod.
Overall I think having good observability is definitely a worthwhile investment. There are cheaper ways to do it than datadog. I expect much of the trouble is that switching providers is a huge job, we have invested so much time building our observability stack, the challenge of moving seems massive. Thankfully we picked Grafana’s open source LGTM stack and self-hosted it. Even if you picked their SaaS offering, switching to open-source self-hosted is an option so you are less tied in.
I'm not an expert on monitoring/observability/telemetry etc, nor an expert on Datadog pricing/billing, but paying a lot of money for major infrastructure components doesn't surprise me.
You give them 1m each.
You never see them again.
I know Datadog probably has a hell of a lot more add-on features, I'm more interested in head to head of comparable produts
The Grafana pricing is more cleanly volume based that is disconnected from number of “hosts” which is where datadog really squeezes you in a kubernetes setup with many pods.
Perhaps it unlocked insurance?
We use Datadog for VM and database monitoring.
When I worked at a place that was all in on Azure, application insights was so we needed because we had no dedicated VMs just all built in Azure services (Cosmos, queues, blob/table storage and functions etc)
It’s the simple mathematics of perpetual motion, as observed in all stable systems.
Since 2022 Q1, the allowance for doubtful accounts has remained under $6 million. Bankruptcy generally triggers a writedown of any associated receivables, so Datadog appears to not have had any material exposure to FTX or any other bankrupt customer.