DataDog asked OpenTelemetry contributor to kill pull request
github.com
github.com
> Speaking as a Collector contrib maintainer, I just wanted to say that I am not going to continue reviewing this PR or start reviewing other future PRs related to the Datadog APM receiver to avoid any semblance of conflict of interest given my role both as a maintainer on this repository and as a member of Datadog's OpenTelemetry team.
And from their GitHub profile:
> Open Source Software Engineer at Datadog , focusing on OpenTelemetry
What's the conflict of interest exactly? You work at Datadog, supposedly to work on OSS, with a focus on OpenTelemetry and you don't want to review Datadog related code for OpenTelemetry? Sounds weird, that kind of profile is exactly the type of person who should be reviewing the code, they have knowledge of both sides of it.
Rather, it sounds like Datadog is walking back and don't want to support OpenTelemetry if it means it'll support their own tooling, instead of just others.
It lowers the switching cost to get off of DD.
DD_API_KEY= DD_SITE="datadoghq.com" bash -c "$(curl -L https://s3.amazonaws.com/dd-agent/scripts/install_script_agent7.sh)"
They tell you to sign in because installing without a key leads to non-working agents and support tickets.As for the libraries themselves, they're all on the regular package manager for that language, eg. pypi.
Edit: I missed the environment variable before "curl". The .sh is downloaded without the API key but the rest could be done using the API key, since it is passed to the script.
Secret Agent Man...
From the perspective of a customer, I can tell you that DD already has quite a bit of a moat. Their main competitive advantage, and what got us into using it, is being able to correlate data across APM, custom metrics, and logging through the use of tagging. They then densely link data together across the platform. There is also a built-in Jupyter-style notebooks. By correlating data like that, you get more value out of ingesting as much data into DD as you can. There are some additional services we're not using, such as auto-correlation with ML (and alerting for anomaly detection), and security monitoring that also looks across the entire platform using their ML tech.
Like AWS/GCP/Azure, it can get expensive, quite fast, using on-demand pricing, so there are negotiated annual contracts. Right now, our team is small, and to replicate the functionality we do use, using self-hosted open-source tooling, we might as well hire another engineer for just setting up and maintaining such a platform.
I get it that, you want to defend the moat and that eroding the little things can lead to eroding the big things. As I see it though, if you need those correlations, you'll need a certain scale and team size before it makes sense to build out something like that for yourself.
Maybe we are too small but Datadog is one of the few vendors which we haven't been able to negotiate down in years. The price has always been whats on the website. I honestly don't even mind, with some vendors it feels like you are on a basar and they always tell you that their final discountns had to get approval by the CEO.
Compared to most Enterprise vendors it is a lot harder to get a discount from Datadog. Most vendors will give you 1/3 off just for signing a contract and committing to a spend, Datadog is not like that.
We spend a few thousand a month with Datadog and our account manager reaches out every quarter to adjust our monthly commit up/down which provides a 20% discount (I think) or so off from the website prices.
I am one of the maintainers. We are building a DataDog alternative with native support for opentelemetry.
Hopefully someone else will contribute the notebooks feature. Those are very useful.
Something that DD is not careful about, is being able to consistently use UTC for all time labels in all graphs (and maybe a quick way to convert to a local time if we need to communicate with stakeholders).
(I don't know why your comment was downvoted).
Thanks for the point about Notebooks, we have not thought in detail on how people use that. Is it primarily to collaborate between team members when an incident happens or even when there is no incident, and you are analysing stuff
- Incidents, collecting different metrics and showing them next to each other, with comments
- Longer-term reliability debugging. They can form a kind of ad-hoc dashboard. These are usually issues that degrade performance, don't have immediate or wide-spread customer impact, and are things we are not immediately able to detect
- Related, performance tuning. Sometimes, the key metric is unknown. We want to explore it, and then make changes to infra, and then see if that moved the needle
- Sometimes, the ad-hoc widgets are useful enough to export to a dashboard
- I can take any widget anywhere else and import it into a notebook, or start a new notebook out of it.
The notebooks are similar to the dashboard, just that, the layout engine only allows a linear notebook layout instead of a grid. There are already text widgets, though the button to add that is easier to access. Other than the comments, it's basically a dashboard with the UI changed so that it feels like a notebook.
Keep in mind too, all dashboard and notebooks modify timestamps and other states in the browser URL, so it is easy for me to copy-paste those into Slack so that other people can see what I am seeing.
As-is we go through a song and dance whenever we look at logs and metrics “oh, this happened at X time which is Y time for most people.
When we talk to stakeholders and customer-facing folks though, tend to convert it to local time.
Got any examples?
I tried running my own "stack" for a project I wanted alerting on. I landed on Jaeger all-in-one (wasted time on Zipkin, the UI just was nowhere near as good as it ought to be) Docker container in docker-compose with COLLECTOR_OTLP_ENABLED.
We offer a free trial and don't charge per a seat.
Meh. The best way to keep somebody on your product is to make it easy for them to get off your product.
GitHub's own answer to this is to force engineers to use a /slash command to post a summary of the week's updates. Clunky, but it works.
It seems like right now, data can flow IN datadog libraries/agents but not out. This PR would sort of allow data to flow OUT of datadog's libs/agents?
And DD doesn't want that because it removes their lock-in power?
Is this correct? This would be extremely crappy of Datadog.
Yes.
This allows you to expose a Telemetry collector on the datadog agent port 8126[0], allowing you to collect Application Performance Monitoring (APM) traces from any APM-enabled datadog library[1].
If I had to guess, DataDog's argument is that they don't want you using the engineering hours they invest into their libraries to have DD do the heavy lifting of collecting APM traces and send the messages off to another service.
<removed OSS comment>
0: https://github.com/boostchicken/opentelemetry-collector-cont...
1: https://docs.datadoghq.com/tracing/#send-traces-to-datadog
Is that not a verbatim the 3-clause BSD license (which is an FSF/OSI approved OSS license)?
In practice, DD has a lot going for it that I don't see in New Relic. There are also some key features in DD that is not in OTEl -- for example, we can't use DD's APM ingestion controls for controlling sampling rates for OTEL spans, and DD has no incentive to add such a feature. I'm actually working on adding in Otel sampling into our project right now. (In our case, we have to use Otel because DD does not have SDKs for Elixir)
Even if we were self-hosting, there's a cost to ingesting and storing every single span.
And even if we are able to pay for ingesting 100%, not everything is practical to be ingested 100%. Our most common request type (heartbeat) generate a span payload size that is a multiple of the original request. We're using Elixir in production, and those can absorb a tremendous amount of traffic, saturating the entire CPU capacity of the hardware if we let it. The agents are not capable of keeping up.
Firstly, not all spans are interesting. When 99.99% of your traffic is just going to serve up an HTTP 200 within your acceptable latency threshold, you don't need every one of those. You probably do want to keep 100% of error spans, or those where the root has a duration beyond a configured threshold. There's tools to be able to sample that way.
Secondly, there's ways to also attach your effective sample rate as metadata to spans, and if there's a backend that supports re-weighting counts based on that, you can still get accurate all-up counts of overall traffic.
Admittedly, OTel and many other backends don't have the best story for this yet. But it's getting better.
The only reason they support "Open" Telemetry is because they're worried about lock-in at the data sources.
For example, App Insights supports rich/structured telemetry via its proprietary SDK and various APIs. No open-source developer in their right mind would ever hard code such a proprietary dependency into something published under a truly open license.
Now that rich telemetry instead of simple text logging is starting to become an increasingly popular approach, the proprietary APM vendors got nervous that they would get "locked out" of the entire open source ecosystem, to be replaced by a data source that is open and not compatible with their proprietary sinks.
Hence Open Telemetry.
It was always about making the source open, not the sink.
From my perspective (maintainer, employed by a vendor), all of us who work for these different companies collaborate very well together. We all recognize that it's both technically tractable and fundamentally user-friendly to make instrumentation be a common standard that anyone can use to point at any of the OSS and commercial tools in this space. There's plenty to differentiate on with telemetry backends, querying experiences, API capabilities, data analysis tools and UX, etc. We have a long way to go to see this vision fully realized, but it's quite far along and I have no doubt we'll arrive at the right outcome here.
The vendors are concerned about being left out at the "source", and are seeking to differentiate with their proprietary "sink", usually closed-source SaaS solutions.
I'm not even arguing that this is bad, it's just how markets work, and it's currently beneficial to developers in general, including both open-source developers and the type working in a cubicle farm somewhere.
Data dog has always been a proprietary POS. I don’t know why people use it, APM traces? How long before Grafana has these capabilities in OSS?
So annoying seeing a company like DD who cannot innovate at all, trying to lock in the average company.
They've already started down that path with Grafana Tempo[0]. It's functional, but their UX and discoverability need a lot of work.
You see the same thing around integrations, everyone used to have to roll their own proprietary chunks of code that in the end were all querying mostly the same data points back from servers and API's. Now everyone just prefers to wait for the Prometheus exporter and they adopt that instead.
As it stands, this was/is trying to use DataDog's existing APM libraries to turn them into OTel for ingest into other providers.
PS: I am one of the maintainers
> Hello, I'm using this receiver in production for about one year, and left some comments that may be helpful.
>> Are you serious? That is intense. Did it scale? Any memory issues? If I remember correctly I tuned all that away.
Always astonishing to see broken stuff doing well in prod lol.
Switching APM providers isn't hard at all. Maybe the Open Telemetry ecosystem lacks a good agent but my understanding is that their APM agent is pretty much already a copy of the Datadog agent.
I think New Relic would be a more interesting agent to use because their approach to APM isn't primarily to do sampled traces.
https://github.com/open-telemetry/opentelemetry-collector-co...
They're very hot on licensing though.
I could be totally wrong, but maybe combined we can make sense of it. From my read it feels like DD misguided this third party which then advised the author to spike.