Phlare: open-source database for continuous profiling at scale
grafana.com
grafana.com
Edit:
Just tried to run it in Grafana, but it's not easy. Datasource for Phlare is not in a stable grafana image:
--set image.repository=aocenas/grafana \
--set image.tag=profiling-ds-2 \
flamegraph plugin is in beta behind feature flaghttps://grafana.com/docs/grafana/next/panels-visualizations/...
>Note: This panel is currently in beta & behind the flameGraph feature toggle.
With these two issues in mind, announcement of this product feels a bit rushed just to show it during ObservabilityCON, when I can't run it locally with stable images and plugins. I hope to see it release in mainstream repos soon!
I have been working on a single datastore that can effectively manage all the different kinds of data. So far it can manage hundreds of millions of files better than file systems. It can form relational tables that query faster than other databases (https://www.youtube.com/watch?v=Va5ZqfwQXWI) and it has a schema flexible enough to handle stuff normally stored within Json files. It still needs a lot of work, but it is currently in open beta.
I'd guess this rules out AWS as well as containers, too, right?
[update] hmm.. per docs wants pprof format and an agent. https://grafana.com/docs/phlare/latest/operators-guide/confi...
language support https://grafana.com/docs/phlare/latest/operators-guide/confi...
re multi-lang support via pprof - afaics python hasn't been touched in years, java is shiny new albeit also first party, in golang pprof is native, and rust pprof seems active.
It takes a lot of time and effort to bake a cross-vendor cross-language standard.
Sounds like it, with the addition of a Grafana panel. It seems like there is a bit of overlap between this and the other products like Tempo, Loki, and Mimir. This graphic seems to indicate it stands independently though, aside from Grafana visualizations. https://grafana.com/static/assets/img/diagrams/grafana-diagr...
But I really, really like the idea. Often when I want to test the performance of a change I'll launch test and control canary instances with a small percentage of live traffic, run perf against each, collect the data, load it into a local https://profiler.firefox.com/, and try to compare the differences. It would be awesome to automate that process. Beyond that, I often keep notes about the tests but the profiles themselves are a real pain to store and catalogue.
We need to pull these profiles in at a regular interval by hitting the HTTP endpoint and we call this “scraping profiles”.
This is very similar to how Prometheus.io works.
It's impossible to keep track of all the moving parts of the modern observability stack.
And what's with those names? Are they picking the names so they can work toward a nice acronym, like they did with Loki+Grafana+Tempo+Mimir (LGTM)?
Profiling is a layer below tracing and metrics in that it requires very minimal and generic instrumentation (of the runtime), while the others require specific instrumentation (of the application).
Profiles like CPU, heap allocs, goroutines, etc are collected on a continual basis, allowing you to see at any time how your application was making use of its runtime resources.
Grafana Faro: An open source project for front end application observability - https://news.ycombinator.com/item?id=33439799 - Nov 2022 (2 comments)
Basically they are churning out all these different projects that just need to be "good enough" from a performance perspective
Grafana is obviously the main one, their original product and the most popular one. It's to aggregate data from various sources and make dashboards.
Loki is a logs collector that I personally didn't try but I think it's popular. They released it after Grafana.
I don't know the third one.
The last one seems to be a wrapper around Prometheus (a metrics collector/database).
Fair to assume you should start with Grafana. For the source, if you don't have a Prometheus instance, you can test it with any SQL database.
Second, when you go to download OSS version, they will first nag you with the cloud version (I'm downloading that thing, not signing up!), then will, by default, link to the enterprise version. Something similar Elastic has done for years.
Also, their cloud offering advertises "Free Forever" (whatever) - we all know how these things end ;)
Grafana: 2.6k issues, 275 PRs
Loki: 531 issues, 113 PRs
Mimir: 305 issues, 40 PRs
Tempo: 159 issues, 19 PRs
For example, I created an issue requesting a flamegraph visualization in grafana[1] and now it makes sense that they didn't initially respond because they were building it internally in secret and didn't want to spoil the big reveal (when they did respond they did mention that it was a secret).
They're also less incentivized now to tend to issues and PRs that help others outside of their ecosystem (i.e. competing logs, metrics, tracing, profiling, etc products).
Secrets kinda conflict with the whole Open part of OSS.
By the OSI definition, most of android is in fact open source. However, much of it is thrown over the wall after a release. Don’t conflate the two.
That's in direct contract with Grafana's own messaging. The following is from their OSS marketing page.
>Open Source is at the heart of what we do at Grafana Labs. We believe building software in the open, with thriving communities, helps leave the world a little better than we found it. https://grafana.com/oss/
Secrets kinda conflict with the whole Open part of OSS.
Please note that the open source definition has no mention of community:As I said, please stop conflating open source software with open communities. They’re part of a thriving and healthy project, but are not the same thing.
The thing with the core Grafana product being open source is that there's not that much dissimilarity between paying/enterprise Grafana users, and open-source users. Feedback from one set will almost always work in favor for the other.
Also it's a bit hard to judge what the number of open issues means right now, because we use github issues also to track internal tasks (for better visibility). So as we have more engineers and more users there will inevitably be more open issues/PRs at any given moment.