HNHacker News
TopNewBestAskShowJobs

hagen1778

55 karma · joined February 9, 2017

Works at VictoriaMetrics

Github: https://github.com/hagen1778

submissionscomments
hagen1778··on Traceway: MIT-licensed observability stack you can self-host in ~90s
I work at VictoriaMetrics.

Just to clarify: VictoriaMetrics doesn't use bots for HN or for any other media for promotion.

I don't know the person who you responded to. Most of the activity you see is coming from community members who genuinely use the project or from the core engineering team trying to answer user's questions or address misunderstandings.

> never responding to any comments

Could you please share examples like this? I can't say for community members, but our internal policy for engineers is very much focused on great support. You can check our slack/github to see that every question is answered and well explained.

hagen1778··on OpenData Timeseries: Prometheus-compatible metrics on object storage
I am curious to see more tests on the reading path. The article mentions matching 500 series over 6h window with 1m step - and it takes 2s for warmed caches. That doesn't sound good at all.

Especially nowadays, when metrics from k8s ramping up churn rate to hundreds of thousands and millions series.

hagen1778··on Moving a large-scale metrics pipeline from StatsD to OpenTelemetry / Prometheus
I was under impression that problem of zero injection was solved with Start Timestamp from OpenMetrics 2.0 spec - see https://prometheus.io/docs/specs/om/open_metrics_spec_2_0/#s...
hagen1778··on OpenData Timeseries: Prometheus-compatible metrics on object storage
Comparing self-hosted prices with managed solutions isn't exactly apples to apples.

But if you do compare, VictoriaMetrics cloud for 3Mil active series and twice higher ingestion rate (100K samples/s or 30s scrape interval) will cost you ~$1k/month + storage costs.

See https://victoriametrics.cloud/#estimate-cost

hagen1778··on Moving a large-scale metrics pipeline from StatsD to OpenTelemetry / Prometheus
What do you use instead of Prometheus?
hagen1778··on Benchmarking Kubernetes Log Collectors: Vector, Fluent Bit, OpenTelemetry
Disclaimer: I am affiliated with VictoriaMetrics.

> The test setup is divergent from the real world (single node k8s cluster)

The setup was chosen to simplify the suite, so it can be easily run anywhere. In real world, log collectors are mostly deployed as deamonsets and this is what was tested during the benchmark. The vlagent was initially developed to run as a deployment, though. So I don't think changing setup will affect its performance.

> Their product is #1 in every metric they measured, but is also missing features from those compared against. Will those features change the results?

Depending on what features will be involved into the testing. In the benchmark, all collectors are doing the same job: collecting logs, parsing JSONs, shipping log records. So they are even in used features. Of course, the #1 product is missing features for log transformations, but these features aren't used during testing and shouldn't affect performance of other log collectors.

The bottom line of the post has "Should I switch" section explaining what's missing yet in product #1 for transparency.

hagen1778··on Benchmarking Kubernetes Log Collectors: Vector, Fluent Bit, OpenTelemetry
The benchmark suite is available here: https://github.com/VictoriaMetrics/log-collectors-benchmark
hagen1778··on I can't recommend Grafana anymore
Using "period" triggers me :)

If Mimir is the only one, why Roblox, GrafanaLabs's customer, isn't using Mimir for monitoring? They're using VictoriaMetrics on approx scale of 5 Billion active time series. See https://docs.victoriametrics.com/victoriametrics/casestudies....

None solution is perfect. Each one has its own trade-offs. That is why it triggers me when I see statements like this one.

hagen1778··on I can't recommend Grafana anymore
Just to add, VictoriaMetrics covers all 3 signals:

- VictoriaMetrics for metrics. With Prometheus API support, so it integrates with Grafana using Prometheus datasource. It has its own Grafana datasource with extra functionality too.

- VictoriaLogs for logs. Integrates natively with Grafana using VictoriaLogs datasource.

- VictoriaTraces for traces. With Jaeger API support, so it intergrates with Grafana using Jaeger datasource.

All 3 solutions support alerting, managed by same team, are Apache2 licensed, are focused on resource efficiency and simiplicity.

hagen1778··on I can't recommend Grafana anymore
There are plenty of ways to scale Prometheus:

- Thanos

- Mimir

- VictoriaMetrics

All of them provide a way to scale monitoring to insane numbers. The difference is in architecture, maintainability and performance. But make your own choices here.

Before, I remember there was m3db from Uber. But the project seems pretty dead now.

And there was a Cortex project, mostly maintaned by GrafanaLabs. But at some point they forked Cortex and named it Mimir. And Cortex is now maintained by Amazon and, as I undersand, is powering Amazon Managed Prometheus. However, I would avoid using Cortex ecaxctly because it is now maintained by Amazon.

hagen1778··on I can't recommend Grafana anymore
I think OTEL has made things worse for metrics. Prometheus was so simple and clean before the long journey toward OTEL support began. Now Prometheus is much more complicated:

- all the delta-vs-cumulative counter confusion

- push support for Prometheus, and the resulting out-of-order errors

- the {"metric_name"} syntax changes in PromQL

- resource attributes and the new info() function needed to join them

I just don’t see how any of these OTEL requirements make my day-to-day monitoring tasks easier. Everything has only become more complicated.

And I haven’t even mentioned the cognitive and resource cost everyone pays just to ship metrics in the OTEL format - see https://promlabs.com/blog/2025/07/17/why-i-recommend-native-...

hagen1778··on Datadog's $65M/year customer mystery solved
My understanding is that with Prometheus+Grafana, and the rest of their stack, you can achieve the same functionality as Datadog (or even more) at much lower costs. But, it requires engineering time to set up these tools, monitor them, build dashboards and alerts. Build an observability platform at home, in other words.

But what about other open source solutions that already trying very hard to become an out-of-box solution for observability? Things like Netdata, Hyperdx, Coroot, etc. are already platforms for all telemetry signals, with fancy UIs and a lot of presets. Why people don't use them instead of Datadog?

hagen1778··on Datadog's $65M/year customer mystery solved
Ofc you need to monitor your monitoring, because you run it. Datadog runs their own systems and monitors them, that's why they charge you so much. I barely can imagine a criticial piece of software that I need to run and not monitor it in the same time.
hagen1778··on Netdata vs. Prometheus: A 2025 Performance Analysis
> Our tests revealed that Prometheus v3.1 requires 500 GiB of RAM to handle this workload, despite claims from its developers that memory efficiency has improved in v3.

AFAIK, starting from v3 Prometheus has `auto-gomemlimit` set by default. It means "Prometheus v3 will automatically set GOMEMLIMIT to match the Linux container memory limit.", which effectively prevents garbage collection until process reaches the specified limit. This is why, I think, Prometheus has increased flatlined mem usage in the article.

> Query the average over the last 2 hours, of system.ram of all nodes, grouped by dimension (4 dimensions in total), providing 120 (per-minute) points over time.

The query used for Prometheus here is an Instant query `query=avg_over_time(netdata_system_ram{dimension=~".+"}[2h:60s])`. This is rather a very weird subquery, that probably was never used by any of Prometheus users. Effectively, it instructs Prometheus to execute `netdata_system_ram{dimension=~".+"}` on 2h interval `2h/60s=120` 120 times, reading `120 * 7200 * 4k series = 3.5Bil` data samples. Normally, Prometheus users don't do this. They'd rather run a /query_range query `avg_over_time(netdata_system_ram{dimension=~".+"}[5m])` with step=5m on 2h time interval, reading `7200*4k=29Mil` samples.

Another weird thing with this query is that in return Prometheus will send 4k time series with all the labels in JSON format. I wonder, how much time from 1.8s it took just to transfer data over the network.

hagen1778··on Prometheus 3.0
It usually comes with increase of active series and churn rate. Of course, you can scale Prometheus horizontally by adding more replicas and by sharding scrape targets. But at some point you'd like to achieve the following:

1. Global query view. Ability to get metrics from all Prometheis with one request. Or just simply not thinking which Prometheus has data you're looking for.

2. Resource usage management. No matter how you try, scrape targets can't be sharded perfectly. So you'll end up with some Prometheis using more resources than others. This could backfire in future in weird ways, reducing stability of the whole system.

hagen1778··on Prometheus 3.0
What makes you think that about docs? Of course, it was written by developers, not tech writers. But anyway, what do you think can be improved?
hagen1778··on Build a serverless ACID database with this one neat trick (atomic PutIfAbsent)
We use the same approach in time series database I'm working on. While file creation and fsync aren't atomic, rename [1] syscall is. So we create a temporary file, write the data, call fsync and if all is good - rename it atomically to be visible for other users. I had a talk about this [2] a few month ago.

[1] https://man7.org/linux/man-pages/man2/rename.2.html

[2] https://www.youtube.com/watch?v=1gkfmzTdPPI

hagen1778··on Show HN: Oodle – serverless, fully-managed, drop-in replacement for Prometheus
ClickHouse recently got the support of TimeSeries table Engine [1]. It is marked as experimental, so yes - early stage. This engine is quite interesting, the data can be ingested via Prometheus remote write protocol. And read back via Prometheus remote read protocol. But reading back is the weakest part here, because Prometheus remote read requires sending blocks of data back to Prometheus, where Prometheus will unpack those blocks and do the filtering&transformations on its own. As you see, this doesn't allow leveraging the true power of ClickHouse - query performance.

Yes, you can use SQL to read metrics directly from ClickHouse tables. However, many people prefer the simplicity of PromQL compared to the flexibility of SQL. So until ClickHouse gets native PromQL support, it is in the early stages.

[1] https://clickhouse.com/docs/en/engines/table-engines/special...

hagen1778··on The Rise of Open Source Time Series Databases
Here you go https://victoriametrics.com/blog/mimir-benchmark/ It is from Sep 2022, it would be great to get newer results.
hagen1778··on All you need is Wide Events, not "Metrics, Logs and Traces"
> And yet people use ClickHouse quite effectively for this very problem

There is no doubt that ClickHouse is a super-fast database. No one stops you from using it for this very problem. My point is that specialized time series databases will outperform ClickHouse.

> There are also time-series databases out there that are OK with high cardinality

So does this blog say that tolerance to cardinality means that QuestDB indexes only one of the columns in the data generated by this benchmark?

TSDBs like Prometheus, VictoriaMetrics or InfluxDB will perform filtering by any of the labels with equal speed, because this is how their index works. Their users don't need to think about the schema or about which column should be present in the filter.

But in ClickHouse and, apparently, in QuestDB, you need to specify a column or list of columns for indexing (the fewer columns, the better). If the user's query doesn't contain the indexed column in the filter - the query performance will be poor (full scan).

See like this happened in another benchmarketing blogpost from QuestDB - https://telegra.ph/No-QuestDB-is-not-Faster-than-ClickHouse-...

hagen1778··on All you need is Wide Events, not "Metrics, Logs and Traces"
Storing telemetry efficiently is only part of what Monitoring is supposed to do. The other part is querying: ad-hoc queries, dashboards, alerting queries executed each 15s or so. For querying to work fast, there has to be an efficient index or multiple indexes depending on the query. Since you referred ClickHouse as efficient columnar storage, please see what makes it different from a time series database - https://altinity.com/wp-content/uploads/2021/11/How-ClickHou...
hagen1778··on PostgreSQL is enough
> Would it be possible to do this in Postgres as well?

Of course! The question is only in your requirements. Keeping a simple counter with limited cardinality should work just great. But nowadays monitoring is much more serious than that. For monitoring k8s clusters the average ingestion rate of metrics per second varies from 100K to 2Mil. I don't know if, resource-wise, it would be a right decision to use Postgres for storing this.

So when requirements are high, and they are for real-time infrastructure and applications monitoring, it is better to consider something like ClickHouse (for people familiar with Postgres) or VictoriaMetrics (for people familiar with Prometheus).

hagen1778··on Migrating to OpenTelemetry
> mimir because of scale & self-host options

Have you looked at VictoriaMetrics [0] before opting for Mimir?

[0] https://victoriametrics.com/blog/mimir-benchmark/

hagen1778··on Influxdb made the switch from Go to Rust
> What is it that you think prometheus offers over other solutions?

I like Prometheus and think this is a great piece of software. But even if we won't go into actual details, Prometheus is baked in into Kubernetes monitoring [0]. That's the first monitoring system young engineers will meet with when learning k8s. Although, k8s and Prometheus are both CNCF projects which means both of them will be promoted in synergy with each other.

> It is more likely that the younger engineer is going to learn that companies don't care about what is popular on HN.

This is not what I think younger engineers do :)

[0] https://kubernetes.io/docs/tasks/debug/debug-cluster/resourc...

hagen1778··on Influxdb made the switch from Go to Rust
> Prometheus only handle aggregated data, though.

That's not true. You're referring to pull-based approach for metrics collection. It has its tradeoffs (like fixed interval scraping), but has a lot of benefits too (like higher reliability). Check the following link [0] from VictoriaMetrics docs, which supports both push and pull approaches. Prometheus also gained push support this year, though.

However, the main difference between Prometheus-like systems (Thanos, Mimir, VictoriaMetrics) and more traditional DBs for time series like InfluxDB or TimescaleDB is that first are designed to reflect system's state, and last are designed to reflect system's events. That's the main difference in paradigm, data model, and query languages. There is a reason why PromQL is so easy in 99% of cases, and so complex and annoying when users want to express what they get used to in traditional databases.

I'm saying this because I went through creating a Grafana datasource for ClickHouse [1] and I felt how complicated it is to express a most straightforward PromQL query in SQL, and vice versa.

If you'd like to learn more about differences between common queries for plotting time series in PromQL and SQL see my talk here [2].

[0] https://docs.victoriametrics.com/keyConcepts.html#write-data

[1] https://grafana.com/grafana/plugins/vertamedia-clickhouse-da...

[2] https://youtu.be/_zORxrgLtec?t=835

hagen1778··on Influxdb made the switch from Go to Rust
That's only a matter of time when they hire younger engineers who are familiar with modern monitoring systems and eager to apply their knowledge in practice.
hagen1778··on Influxdb made the switch from Go to Rust
If I'm reading this [0] right, there will be no a standalone OS influxdb 3.0 version. So there's no point in comparing. I also wonder if it would be allowed to publish benchmarks of ENT version by 3rd-parties.

[0] https://www.influxdata.com/blog/the-plan-for-influxdb-3-0-op...

hagen1778··on CERN swaps out databases to feed its petabyte-a-day habit
Yes! Alerting and recording rules are supported by vmalert [0]. vmalert then integrates with alertmanager for sending alerts, and alertmanager then dispatches notifications. Besides this, vmalert has features of retro-active rules evaluation [1] (both alerting and recording) and detection of useless alerting rules [2].

[0] https://docs.victoriametrics.com/vmalert.html [1] https://docs.victoriametrics.com/vmalert.html#rules-backfill... [2] https://docs.victoriametrics.com/vmalert.html#never-firing-a...

hagen1778··on How Grammarly Improved Monitoring by over 10x with VictoriaMetrics
> Our proof-of-concept trial showed dramatically reduced compute and storage costs, translating into a 10x lower AWS bill
hagen1778··on OpenTelemetry in 2023
You shouldn't unless you want to use the new open source standard for telemetry. You won't benefit from simplicity or performance improvements. It would be quite the opposite. You can check what is the actual cost of open telemetry adoption here [0]

But if you ever decide to go this path - VictoriaMetrics supports OpenTelemetry protocol for metrics [1]

[0] https://github.com/VictoriaMetrics/VictoriaMetrics/pull/2570

[1] https://docs.victoriametrics.com/Single-server-VictoriaMetri...

Page 1 of 3Next →