904 karma · joined July 25, 2017
My own wish is that I hope something like this https://github.com/knightss27/grafana-network-weathermap is adopted too. I know it maybe possible using Canvas but Canvas is too much of a blank slate and needs a lot of work to set it up the same way.
Clickhouse has proven to also be a very capable database for logs and there are stacks that use it for log storage.
It's evident that some of those trusting people are willing to make or save a buck while putting you at risk.
At the end of it, people should probably make their own assessment as to whether they should put themselves at risk. And I don't mind drying a handful of dishes if the alternative is to lace them all with surfactant.
Our rinse aid is disabled on our dishwasher.
The architecture is quite different between Thanos and the others you've listed as unlike the others, Thanos queries fan out to remote Prometheus instances for hot data and then ship out data (typically older than 2 hours) via a sidecar to s3 storage. As the routing of the query depends on setting Prometheus external labels, our developer queries would often fan out unnecessarily to multiple prometheus instances. This is because our developers often search for metrics via a service name or some service related label rather than use an external label which describes the location of the workload which is used by Thanos.
Upon identifying this, I migrated to Mimir and we saw immediate drops in query response times for developer queries which now don't have to wait for the slowest promethues instances before displaying the data.
We've also since adopted OpenTelemetry in our workloads and directly ingest otlp in to Mimir (Which VictoriaMetrics also support).
I think Elasticsearch had its day when it's used to derive metrics from logs and performing aggregate searches. But now as logging is often paired with metrics from Prometheus or similar tdb, we don't run such complex log queries anymore, and so we find ourselves questioning whether it's worth running such a intensive and complex Elasticsearch installation.
Capture method can have a big effect on latency and technologies like the Desktop Duplication API or NV FBC are also necessary for achieving lower latencies by minimising data transfer between cpu and the gpu when capturing and encoding.