HNHacker News
TopNewBestAskShowJobs

nikolay_sivko

217 karma · joined April 27, 2017

submissionscomments
nikolay_sivko··on OpenTelemetry for Go: Measuring overhead costs
As suggested, I measured the overhead at various sampling rates:

No instrumentation (otel is not initialized): CPU=2.0 cores

SAMPLING 0% (otel initialized): CPU=2.2 cores

SAMPLING 10%: CPU=2.5 cores

SAMPLING 50%: CPU=2.6 cores

SAMPLING 100%: CPU=2.9 cores

Even with 0% sampling, OpenTelemetry still adds overhead due to context propagation, span creation, and instrumentation hooks

nikolay_sivko··on OpenTelemetry for Go: Measuring overhead costs
I'm the author. I wouldn’t say the post is critical of OTEL. I just wanted to measure the overhead, that’s all. Benchmarks shouldn’t be seen as critique. Quite the opposite, we can only improve things if we’ve measured them first.
nikolay_sivko··on Show HN: Coroot – eBPF-based, open source observability with actionable insights
We could totally add that, but no one's asked for it so far
nikolay_sivko··on Show HN: Coroot – eBPF-based, open source observability with actionable insights
1. Regarding overhead — we ran a benchmark focused on performance impact rather than raw overhead [1]. TL;DR: we didn’t observe any noticeable impact at 10K RPS. CPU usage stayed around 200 millicores (about 20% of a single core).

2. Coroot’s agent captures pseudo-traces (individual spans) and sends them to a collector via OTLP. This stream can be sampled at the collector level. In high-load environments, you can disable span capturing entirely and rely solely on eBPF-based metrics for analysis.

3. We’ve built automated root cause analysis to help users explain even the slightest anomalies, whether or not SLOs are violated. Under the hood, it traverses the service dependency graph and correlates metrics — for example, linking increased service latency to CPU delay or network latency to a database. [2]

4. Currently, Coroot doesn’t support off-CPU profiling. The profiler we use under the hood is based on Grafana Pyroscope’s eBPF implementation, which focuses on CPU time.

[1]: https://docs.coroot.com/installation/performance-impact [2]: https://demo.coroot.com/p/tbuzvelk/anomalies/default:Deploym...

nikolay_sivko··on Show HN: Coroot – eBPF-based, open source observability with actionable insights
From a user’s perspective, it doesn’t really matter how the data is collected. What actually matters is whether the tool helps you answer questions about your system and figure out what’s going wrong.

At Coroot, we use eBPF for a couple of reasons:

1. To get the data we actually need, not just whatever happens to be exposed by the app or OS.

2. To make integration fast and automatic for users.

And let’s be real, if all the right data were already available, we wouldn’t be writing all this complicated eBPF code in the first place:)

nikolay_sivko··on Show HN: Coroot – eBPF-based, open source observability with actionable insights
Enterprise Edition = Community Edition + Support + AI-based Root Cause Analysis + SSO + RBAC
nikolay_sivko··on Show HN: Coroot – eBPF-based, open source observability with actionable insights
Yes, it captures traffic before encryption and after decryption using eBPF uprobes on OpenSSL and Go’s TLS library calls.
nikolay_sivko··on Show HN: Coroot – eBPF-based, open source observability with actionable insights
Currently, you can define custom SLIs (Service Level Indicators, such as service latency or error rate) for each service using PromQL queries. In the future, you'll be able to define custom metrics for each application, including explanations of their meaning, so they can be leveraged in Root Cause Analysis
nikolay_sivko··on Show HN: Coroot – eBPF-based, open source observability with actionable insights
Initially, we relied on the ClickHouse OTEL exporter and its schema, but for performance optimization, we decided to modify our ClickHouse schema, and they are no longer compatible :(
nikolay_sivko··on How Netflix Accurately Attributes eBPF Flow Logs
At Coroot, we solve the same problem, but in a slightly different way. The traffic source is always a container (Kubernetes pod, systemd slice, etc.). The destination is initially identified as an IP:PORT pair, which, in the case of Kubernetes services, is often not the final destination. To address this, our agent also determines the actual destination by accessing the conntrack table at the eBPF level. Then, at the UI level, we match the actual destination with metadata about TCP listening sockets, effectively converting raw connections into container-to-container communications.

The agent repo: https://github.com/coroot/coroot-node-agent

nikolay_sivko··on Show HN: Coroot – eBPF-based, open source observability with actionable insights
Coroot builds a model of each system, allowing it to traverse the dependency graph and identify correlations between metrics. On top of that, we're experimenting with LLMs for summarization — here are a few examples: https://oopsdb.coroot.com/failures/cpu-noisy-neighbor/
nikolay_sivko··on Show HN: Coroot – eBPF-based, open source observability with actionable insights
It only requires a modern Linux kernel. Note: The agent does not support Docker-in-Docker environments, such as KinD or Minikube (D-in-D plugin).
nikolay_sivko··on Show HN: Coroot – eBPF-based, open source observability with actionable insights
(I'm a co-founder). At Coroot, we're strong believers in open source, especially when it comes to observability. Agents often require significant privileges, and the cost of switching solutions is high, so being open source is the only way to provide real guarantees for businesses.
nikolay_sivko··on Show HN: Coroot – eBPF-based, open source observability with actionable insights
In addition to raw logs, Coroot can extract recurring patterns to generate log-based metrics [1].

We also plan to convert structured logs into OpenTelemetry attributes [2].

[1] https://demo.coroot.com/p/tbuzvelk/applications/default:Depl... [2] https://github.com/coroot/coroot/issues/490

nikolay_sivko··on Any cost effective newrelic alternative?
Take a look at Coroot: https://github.com/coroot/coroot. It supports OpenTelemetry and stores traces in ClickHouse.

Here you can find a demo of the Distributed Tracing UI: https://demo.coroot.com/p/tbuzvelk/traces

Disclaimer: I'm a co-founder

nikolay_sivko··on Using AI for Troubleshooting: OpenAI vs. DeepSeek
Sure, but the scope is quite broad. For example, we might need to explain logs from a relatively rare app or framework. In such cases, a large model could unintentionally have knowledge of it.

Disclaimer: I'm the author of the post

nikolay_sivko··on Ask HN: How do you handle observability and alerting at tiny companies?
Give https://github.com/coroot/coroot a try for full visibility in minutes with eBPF

disclaimer: I'm a co-founder

nikolay_sivko··on Ask HN: What's your preferred logging stack in Kubernetes
Take a look at Coroot [0], which stores logs in ClickHouse with configurable TTL. Its agent can discover container logs and extract repeated patterns from logs [1].

[0] https://github.com/coroot/coroot

[1] demo: https://community-demo.coroot.com/p/qcih204s/app/default:Dep...

nikolay_sivko··on Grafana Labs Observability Survey 2024
Take a look at https://github.com/coroot/coroot (Apache 2.0). It offers plenty of ready-to-use dashboards and inspections
nikolay_sivko··on eBPF-based auto-instrumentation outperforms manual instrumentation
Disclaimer: I'm a co-founder of Coroot. We're currently benchmarking our eBPF-based agent to measure its performance impact.

Could you please elaborate on a few more details about your benchmark?

- Did you measure the CPU usage of the eBPF agent?

- How does Odigos handle eBPF's perfmap overflow, and did you measure any lost events between the kernel and the agent?

nikolay_sivko··on Show HN: Coroot – Copilot for Application Performance Troubleshooting
Coroot leverages eBPF to collect metrics and traces, and its agent parses container logs to extract recurring patterns from them.

Here is the list of collected metrics with detailed descriptions: https://coroot.com/docs/metric-exporters/node-agent/metrics

nikolay_sivko··on Show HN: Coroot – Copilot for Application Performance Troubleshooting
Nik and Anton here - we're building an open-source observability tool that transforms telemetry data into actionable insights. While there are many excellent observability tools on the market, pinpointing the root cause of an incident often requires manual searching through a sea of metrics, logs, and traces.

Coroot acts as a virtual assistant, conducting system audits just like an experienced engineer would:

- It utilizes telemetry data collected through eBPF to construct a model of the distributed system and understand its topology.

- It traverses the dependency graph and audits every relevant service to identify the root cause.

Our journey has been long, but we've made great strides in learning how to build better models of distributed systems, and improving our agents to gather the right metrics. Coroot is an open-source product (Apache 2.0), so you can self-host it for free. We charge for a cloud version with AI-based Root Cause Analysis, RBAC, and premium support.

nikolay_sivko··on A cryptocurrency company had a $65M bill, per Datadog’s Q1 earnings call
Take a look at Coroot: it's open-source, ebpf-powered, and can be integrated in minutes: https://github.com/coroot/coroot
nikolay_sivko··on Using eBPF and predefined inspections to minimize “observability tax”
Grafana offers one of the best solutions for storing and navigating through telemetry data, but there remains a challenge in using this data to generate insights. Our goal is to address this issue, even if it means occasionally reinventing the wheel. This is a necessary step at this stage.
nikolay_sivko··on Using eBPF and predefined inspections to minimize “observability tax”
Coroot's agent collects data from various sources to cover all aspects of container behavior. Ebpf-exporter perfectly solves the problem of running custom ebpf programs and turning their output into metrics, but using it as a foundation for more specific solutions doesn't seem reasonable
nikolay_sivko··on Using eBPF and predefined inspections to minimize “observability tax”
With ebpf-exporter it is not possible to implement complex logic, such as converting the PID of each TCP connection into a container name and the destination IP into a real IP according to the conntrack table.
nikolay_sivko··on [dead]
Take a look at https://github.com/coroot/coroot (Apache 2.0), a zero-instrumentation observability tool for microservice architectures. Thanks to eBPF, it can be integrated in minutes.

Live demo: https://community-demo.coroot.com

nikolay_sivko··on Show HN: Failurepedia - a playground with recorded real-world failure scenarios
When we started working on Coroot, we weren't sure if it would be possible to create such a product or not. In order to verify the initial hypotheses, we reproduced real-world failure scenarios in our staging environment and recorded their telemetry. Now we are able to replay this telemetry data, applying Coroot's inspections, to detect particular failures.

Along the way we discovered two things:

* Real-life telemetry data is noisy and inaccurate, so using this allows us to develop more accurate inspections.

* Replaying these scenarios is actually very entertaining!

At Coroot we call our library of recorded failure scenarios Failurepedia, and we want to share it with you.