Do you have a roadmap of the changes you are looking to bring? Any changes to Otel you’re looking to add as well?
Does it still apply that if my Prometheus goes down or network glitches then metrics for that period is lost forever?
Unfortunately, yes. The OTel Collector has plans to implement a WAL for the OTLP exporter and when it does, you should be resilient to upstream temporarily having issues.
WAL = Write-Ahead Logging ?
Yes sorry
It does today. You have a retry queue and you can use a persistent storage for it.
How will that affect timestamps, won't all the metrics have the timestamp as the time when Prometheus finally receives them?
What are the merits of the prometheus approach versus one where events/metrics and their original timestamps can be preserved, stored (during temporary outages), then forwarded to backends when connectivity is reestablished?