I am glad that the observability sector has standardized on a common protocol but my god are the reference implementations lacking.
I am glad that the observability sector has standardized on a common protocol but my god are the reference implementations lacking.
https://learn.microsoft.com/en-us/dotnet/core/diagnostics/ob...
Can you share some of your experience, what do you mean by that? Are there edge cases causing problems, or major missing features? Easy or difficult to use?
[1] https://opentelemetry.io/docs/specs/otel/metrics/data-model/...
[2] https://github.com/open-telemetry/opentelemetry-python
[3] https://github.com/open-telemetry/opentelemetry-python/issue...
I've mentally decided to just go Prometheus and ignore OpenTelemetry for the foreseeable future.
It's one of those things big players are hyping to preemptively lock you in their solution, but it's actually just alpha-quality new tech and "boring" "old" tech like Prometheus or statsd are simply more functional and better supported in the wild.
Btw, metric generation is not enabled in Tempo by default.
# tempo.yml
overrides:
defaults:
metrics_generator:
processors: [service-graphs, span-metrics]
# Prometheus
--web.enable-remote-write-receiver
# Grafana.yml
[feature_toggles]
enable = tempoSearch tempoBackendSearch traceToMetrics
[1] https://grafana.com/docs/tempo/latest/metrics-generator/
[2] https://hexdocs.pm/opentelemetry_process_propagator/Opentele...> What we ended up implementing was a little tee inside the o11y library. As well as sending events to Honeycomb, we also converted them to JSON, and wrote them to stdout. That way, after sending to stdout, we then pumped off to our standard log aggregation system. This way, we've got a fallback. If Honeycomb is not working, we can just see our logs normally. We could also send these off to S3 or some other long term storage system if we wanted.
I'd like to go a step further, and say that in addition to being worried about honeycomb being down, sometimes you just want to check with kubectl to get an idea what is going on.
Our current projects are very log light because of the heavy tracing instrumentation, but it'd be nice to integrate this with the otel paradigms as they were originally intended
So. Much.
Flames. Flames! On the sides of my face.
Breath… Heaving breaths.
Just the general problem you get with big, slow moving OSS projects like this. Mostly just docs not current and a massive delta between certain languages; a feature is `stable` for some languages but not others which makes it hard to push for consistent otel roll out in a mixed-language environment.
Some other "misc" points:
- Google how to do $thing and you might find the proposed spec which gives example code ... that isn't what actually got implemented. That's a different link further down on your google results.
- Python auto-instrumentation is ... fragile at best. It's not super clear if instrumentation is supported only with well known frameworks or just ... in general. I'd sure love some docs that explain how it works, too.
- certain things require the collector use GRPC, others work with grpc or http... and I only found this out after googling an obscure error and reading through a _very_ long GH issue thread.
One other key area is resources which can help get engineers/implementors to get organizational buy-in
I think it's a mistake for Otel to do its own thing instead of just building on top of Prometheus.
https://grafana.com/docs/grafana-cloud/send-data/otlp/send-d...