Why and how GitHub is adopting OpenTelemetry
github.blog
github.blog
It's nice to see all the traces through all the different services of one API request and all its log items it generated. Using a custom span exporter to ensure PPI is stripped away as much as possible for span attributes.
Using W3C trace context and baggage to send some more data along, e.g. sending the trace id, span id etc into pub sub events to ensure to connect them with the trace.
Loving it! For now, I think it's worth the costs of adding tracing to services.
This article is about a fairly large sized tech company adopting a fairly recent but increasingly mainstream & popular tool that helps them understand their operations. It'll give them a standard way to see what their computers are doing, across their various systems.
OpenTelemetry is one of the key emerging cloud standards, and I expect many many many many more articles like this going forwards, from all kinds of companies.
As for concerns about user tracking & privacy, that's not generally what these operational tools are used for. I haven't heard of a single case of them being used for user tracking or behavioral analytics. Thusfar these are purely operational tools, to understand the health of systems, to debug & typically to understand what happens to an incoming request as it works through dozens of systems & services to get processed. That said, I tend to think over time the importance of this distributed after-the-fact log we are building is going to become inverted. That we will start to see the potential to harvest these records for analytics, and more generally, to forward-feed them into other processes to automatically build & advance Event Sourcing systems. Right now these systems are relatively pure & good, but what's really at stake here is that we've been doing computing blind, with no record of what's happened, and OpenTelemetry is a key first step in lifting that veil of ignorance as to what computing has happened. We are beginning to capture the data of what compute occurred. Many things will emerge as we open this box.
I'm not sure that the intended use matters. Unless the telemetry systems are carefully designed to not capture PII at all they become yet another channel that must be secured as if they are collecting PII. For example, see Windows 10 telemetry vs. HIPPA: https://hipaaone.com/2015/09/22/windows-10-and-hipaa/
Ultimately the intended use matters more than the tech. And in this case it seems purposely built to measure performance.
It's just a protocol to see how systems are performing, but a better one than statsd or logs.
These aren't scary things.
From reading the OpenTelemetry site, that doesn't seem to be their main, stated purpose, but thefounder's post in another part of the thread leads me to think it may, in fact, see that kind of use, too.
[EDIT] damn, sorry, downvoters, for explaining the reason this is getting knee-jerk negative reactions from people, when someone expressed confusion about it. Again, I think the main reason is the word used, and what it's mostly associated with now, among some folks.
[EDIT] and, especially, that's why they're jumping to the conclusion that github is gearing up to do more spying on its users, which is the part that I think people are bothered by, not the existence of this software package.
It's like being afraid your browser profiler is being used to spy on you. Could it do that? Sure. but there are so many easier ways to accomplish the same task.
People here are blindly rage-triggered by the word "telemetry" without even taking 2 seconds to glance through to see what it actually is.
> I haven't heard of a single case of them being used for user tracking or behavioral analytics.
and then you said:
> That we will start to see the potential to harvest these records for analytics…
So this can be used to gather analytics of any sort of data, such as spying then? This is still worrisome.
It is technically true that such a tool could be used on any sort of data, in the same way that, say, Perl or SQL can be used on any sort of data.
I'd really hope for better among the HNews readership.
Isn't it a better approach to just use logs?
e.g. defer ctx.WithField("path", path).Trace("opening").Stop(&err)
from https://medium.com/@tjholowaychuk/apex-log-e8d9627f4a9a
No code changes needed.
I am effectively getting trace level insights from my logs. I must be doing something wrong!
What you're probably looking at is all the boilerplate set up, e.g. configuring the provider and backend to point to the right stuff. It reminds me of SL4J, which isn't actually an insult. It's just boilerplate, because people want a lot out of their logging and tracing systems.
Demo applications like this are often easily misleading because there are only like 50 lines of "business logic" and 50 lines of tracing setup, so it makes it seem like the tracing is excessively hard. But those two things don't scale the same way. In a large application where tracing is really valuable, you'll have 100,000 lines of business logic, and still only 50 (or maybe like 100) lines of tracing setup, per application.[2] Actual usage at the call sites remains only a line or two in most cases, and easy to add as you need, where you need it, just like a logger.
It is also worth noting in other ecosystems like when I played with tokio_trace, I found integrating tracing easy, even at the very start. So some of this definitely involves the "philosophy" of the client library.
[1] https://github.com/michaelperel/otel-demo/blob/master/cmd/cl...
[2] I guess if the 100,000 LoC running your business is split into 2000 microservices with 50 lines each, then yes, it may be excessive.
Refer to: https://github.com/michaelperel/otel-demo/blob/master/cmd/se...
51: Start a span. Equivalent to one line of log 60: Add a event. Equivalent to one log 69: Set attribute. This retroactively add attribute to whole span since the start. While log don't have exact same effect, one line of log can be used here. 74: RecordError. Equivalent to one log
I haven't compare amount of code to setup a proper logger which connected to correct infra with amount of code to setup otel yet. Still, I don't think it gonna make much difference.
In general, I won't mind either approach if I get a great visualizer. The main reason I would choose OpenTelemetry is I get trace visualizers for free and I can switch to better visualizer anytime that I want.
Yes technically you could use OT to exfiltrate data; guess what, you can do it with http headers, too.
Been using OpenTelemetry for python and golang. It's great. I really like using Jaeger for tracing across small ML webapps. We have users on the scale of dozens so I just trace everything without sampling. Way, way easier than going through logs.
Also used python + jaeger to narrow down some wifi issues on my home network. Basically run a server on each device and round-robin a request through each and look at the flame graph. Turns out, some Macs' location preferences can cause 500-800ms outage multiple times an hour, as it tries to scan know AP ssids or something. Yeah, dumb.
This is compounded by the fact that Github is now a Microsoft company and Microsoft has been caught in the past hiding "telemetry" in programs compiled with Visual Studio.
https://old.reddit.com/r/cpp/comments/4ibauu/visual_studio_a...
The VS telemetry injection is scary as hell though.
Telemetry is like a knife. Context and location matters. Great in the kitchen, bad when lodged in your back.
Sharing this from myself downthread:
Telemetry is the in situ collection of measurements or other data at remote points and their automatic transmission to receiving equipment (telecommunication) for monitoring. The word is derived from the Greek roots tele, "remote", and metron, "measure".
All spyware is, by definition, telemetry. Heck, the whole field of spycraft ( espionage ) is telemetry of sorts. Not all telemetry is spyware.
Exactly!
I hadn't really found an acceptable solution that would work across Java, Node.js, browser, and so on. We'd invented our own formats and then we owned all the integration problems with various monitoring tools. I left the team before we started to adopt, but they've started doing it and it looks like it's help with reducing integration burden. I also think using someone else's opinionated library can help avoid bikeshedding on concepts not related to your core value.
Curious what you mean about the design of OpenTelemetry precluding efficient implementation?
Its important to know what audience you're building for. I believe the audience for otel consists largely of companies that don't look anything like Google. So its fine to sacrifice that last bit of perf gain if it means the code is easier to use and maintain.
FWIW, it also appears that companies like Google would fork or reimplement such systems anyway.
Looks like Opentelemetry (at least its precedessor, OpenCensus) is originated from Google. From OpenCensus website (https://opencensus.io/):
> OpenCensus and OpenTracing have merged to form OpenTelemetry
> OpenCensus originates from Google, where a set of libraries called Census are used to automatically capture traces and metrics from services.
Original internal Google tracing system was probably designed for scale. And opentelemetry's design is probably based on that internal system.
So, maybe poor performance is just an implementations issue.
I haven't used it myself, I've used census, and now looking into OpenTelemetry (though from the least finished version - C++). Had mixed success with it in the past, but trying again. Also not looking at all into side-cars, etc. - We compile all our internal tools, so adding this inside is where I'm getting into.
I've had several times (while at Google), being asked by an SRE that I would call on issues, and they would request to bump the sample tracing from minimal defaults (was it 1 in 100,000 or million - forget) sometimes to 1:1 - for say 30 seconds. This way they'll receive on their end (in their systems that we use) flags to sample too, and at the end get full logs.
Usually the whole trace is visible in few minutes. There were few UI's (nothign like zipkin/jaeged/others outside) - some of them with very "imgui"-like hacky (in good sense) view - like programmer art all over (which I loved - it was much more condensed than standard zipkin/jaeger).
You could've marked something as important, and it'll retain for longer period - otherwise - poof - soon gone. Also it would collect info sometimes directly from the machine it was in (rather than wait to populate).
Obviously, I don't know the details - I was just an user, or more like - allowing (when oncall) trace sampling to be bumped by the SRE - so they would get more info. It's what hooked me actually, because how else would one get everything end-to-end.
Surprisingly it's also useful for single apps, where you have threads (or concurrency tasks, like with TBB/ConCRT) doing nested parallel_for's or spawning jobs, and you want to get idea what's going on. The only tricky bit is how to get your "context" propagated from one thread to another (also not readily done).
It's one thing that the "golan" got right with their context for example.
So it's really awesome, but probably really hard to get right the first few times.
These clients will usually buffer the stats in memory and push them out asynchronously. Performance is definitely affected but I'm pretty sure it's negligible for most cases.
Best practice would be to reduce tracing ratio in production too. So most requests are literally just a timing.
https://davidgildeh.com/2021/03/08/running-python-openteleme...
I need to update the code as I didn't use the BatchSpanProcessor and as a result all my calls are taking 2-3 seconds vs. a more reasonable 300ms.
My only issue so far is the stability and maturity of the project. Now they're at V1 stability of the main SDK looks better, but the instrumentation libraries are still all over the place and need some cleaning up/maturing.
I like the vision though and we're keeping an eye on the project as it matures.
I really would love to see a simpler set of tools that works better for small / medium companies to self-host if they really can't send data to a vendor for some reason.
On the website it even says: "OpenTelemetry is in beta across several languages and is suitable for use. We anticipate general availability soon."
Fyi: that quote has been there for a while.
On a Github scale where you have more engineers available, it's often worth it to adopt cutting-edge alpha/beta software for their novel features, as you generally also have more expertise/mechanisms to ensure stability of the overall product and run less risk of being suffocated in tech debt.
Finally they released a 1.0 and the interfaces are more stable, but the docs are still incomplete or stale in some places. I currently use it in a minor way.
Frankly I can't wait because I love the tools and ecosystem, but I personally would be hesitant to put it into full blown production at such a big company.
It’s obviously not our only goal, but we are in a position to help make the system better for everyone and sincerely wish to do so.
> It's like being afraid your browser profiler is being used to spy on you. Could it do that? Sure. but there are so many easier ways to accomplish the same task.
So take a second to straighten out your panties and then actually look at the thing first. It's just an open-source (!) APM protocol that competes with (edit: more like, adjacent to) DataDog, DynaTrace, New Relic etc etc. It's not even a suitable tool for spying.
You'd write your opentelemitry traces throughout your code and at some top level point in your app you configure and say "Hey, OpenTelemetry, report to New Relic".
By itself, OpenTelemetry does nothing.
It's usefulness is that someone writing a lib can add OpenTelemetry calls throughout and anyone using that lib can then collect metrics/traces into whatever metrics solution they are currently usings (be it DynaTrace, DataDog, or New Relic).
That's not in the cards today. But longer term, I think "knowing what computers are doing" is big business. And, I am very sad to say, eventually it will the obvious & logical way to spy on folks too. Because it will be the obvious & logical way to do many many many things, not because it's tech that's built or intended for spying. But right now computing is ephemeral, we don't persist any of the stack traces we compute through, and fundamentally, tracing really is about distilling out & keeping higher level stack traces. It's something computing needs to have been doing, that will bring us to a radically higher level of understanding (but it also does have some scary uses).
It's been a couple years, but tracing is a powerful new basis with which to start re-engaging the "Turning the datqbase inside out"[1] / "I <3 Logs" view of data-storage (& computing) that had some buzz for a bit. We're not at all here yet, but I think it's coming.
[1] https://martin.kleppmann.com/2015/11/05/database-inside-out-...
It certainly wouldn't be opentelemetry itself as that's just a interface you add adapters to.
Are you thinking a man in the middle would spy? How would that work? This information is pushed over secure connections on the backend likely in a VPN. On the front end, it'd be transmitted over HTTPS. Shouldn't we be more fearful of information collected from DNS than encrypted data sent over HTTPS?
Or are you thinking the metrics aggregators are going to do the spying? How would that impact their business model? Do you think a company would continue to pay the likes of new relic if they were caught giving access to metrics data to outside groups? Do you worry about postgres or prometheus sharing your data with 3rd parties? What about SQL server?
What is the risk model and how would it be different from say the risk model of making an REST call or using a 3rd party library?
Or is it just that "because this is well integrated throughout, it could be used to spy"? Because, generally speaking, these traces don't have enough information to identify what individual users are doing to the system. Even if they did, that wouldn't be a great way to track a user, you'd simply put that tracking information right at the front end of the system. Plumbing it from one end of the system to the other gives little value for a spy. It's adding a bunch of noise to the question you'd want to ask "what are the user's behaviors with our product?"
And even if the demand is there, why would you do this through tracing an not a purpose built spy tool. Wouldn't it be easier for a nefarious lib writer to make a plugin purpose built to collect evil data? If that sells, why wouldn't a tech company buy that instead of buying a solution which combs 3 layers of separation to get worse answer? Why wouldn't google analytics still exist?
You have a bunch of weird straw men that I don't get. I tried to de-emphasize the role of behavioral analytics & user-tracking, because I think it's just one small part of what this will be used for. But I am fairly confident we will eventually start to do more user-tracking via these systems. I've used half a dozen different user-tracking products at various companies, and they all read like ultra-low-fi versions of the ops tools. Ops tools have been evolving, at a far faster rate, far more in the public domain, and at some point, it just wont make sense to instrument your product twice.
I want to re-iterate that I see this as one of the smallest, least interesting aspects of a coming Event-Sourcing-powered-by-tracing world. There's far more profound implications for what could happen to computing in general here (a de-wiring of the request/response microservice world & a shift towards async, reactive systems). But even today, it feels to me like folks work very hard to draw a distinction between their ops tools & their behavioral analytics tool. As a developer, they often work, I often interface with them in very very similar fashion. The desire to draw the distinction has felt illogical, and felt unsupported. Especially as the ops tools advance, I think it will be harder to reconcile the idea that there are & ought be separate systems of tracking/viewing/understanding.
I look at myself and I find that many opinions on technology that I hold have clearly been shaped by how much time I've spent on HN. Data-exfiltration telemetry is one example. But I now think about a lot of things I might not have given much thought to before, like how being backed by a VC can shape the direction of a service for the worse. I also find that I'm unnecessarily hardline on topics like paid services versus open source, and shut out anything positive people have to say about advertising as a revenue source.
I sometimes imagine what would happen if I was born inside the borders of a different country. Some countries have very nationalistic citizens. How much of my thought processes are a result of the people and signals I surround myself with? How many of those signals will reach me whether or not I consent to seeing them (advertisements being one example)?
I think I could have been a very different person if I was raised even a hundred miles from where I actually grew up, despite having a similar chemical makeup. That makes me be more conscious about where I'm likely to seek information, and to try not to look for opinions that I already agree with entirely.
Other relevant euphemisms that have appeared within the past few years: "analytics", "diagnostics", "usage reporting", "connected experiences", "experience improvement".
Take a look at the Jaeger and Zipkin websites and it should be pretty clear what it is used for.
The point of OpenTelemitry is to make it so you could write the places you get your traces/metrics in one part of code and configure which backend system it is collected into in another part.
So, for example, you'd add a `trace("my slow thing"){ be slow }` into your code and later add `report to zipkin` or `report to Jaeger` in another part. The place where you trace "my slow thing" doesn't care about how to interact with the backend system or which backend vendor is ultimately used. You can start using prometheus, zipkin, new relic, Jaeger, whatever, just so long as they have an OpenTelemetry adapter you are golden.
The analog is SLF4J in Java.
It's baffling to me, what do people think happens when they pull/push from a GitHub repo?
GitHub ≠ Git. Far from it.
>> The answer to not getting pressure from your bosses, your stakeholders, your investors or your members, to do the wrong thing later, when times are hard, is to take options off the table right now.
These promises need to be written down in the articles of association in a company, and I have never seen this.
Why is this useful? Because it lets people writing libraries provide the ability for users of their libraries to track performance information without dictating to them what performance tracking tools they should use.
See: https://github.com/open-telemetry/opentelemetry-js-contrib/t...
Anything that leaves your computer and enters hardware owned by MS should make you think twice about using a certain service/product.
Same goes for almost every other company out there.Also FOMO is not an argument, there are plenty of alternatives for almost everything out there.If it's comfort people are worried about then they should stick to iPads and ditch software-dev, because trends like these are why the quality of software has been declining drastically in the last decade.
These streamlined processes of using telemetry for the "product's good" are just a way to compensate that.