https://clickhouse.com/blog/how-we-used-clickhouse-to-store-...
https://clickhouse.com/blog/how-we-used-clickhouse-to-store-...
w/ 1 month retention for traces:
┌─parts.table─────────────────┬──────rows─┬─disk_size──┬─engine────┬─compressed_size─┬─uncompressed_size─┬────ratio─┐
│ signoz_index_v2 │ 26902115 │ 17.06 GiB │ MergeTree │ 6.21 GiB │ 66.74 GiB │ 0.0930 │
│ durationSort │ 26901998 │ 5.44 GiB │ MergeTree │ 5.40 GiB │ 53.02 GiB │ 0.10190 │
│ trace_log │ 123185362 │ 2.64 GiB │ MergeTree │ 2.64 GiB │ 37.96 GiB │ 0.0695 │
│ trace_log_0 │ 120052084 │ 2.46 GiB │ MergeTree │ 2.45 GiB │ 37.60 GiB │ 0.06528 │
│ signoz_spans │ 26902115 │ 2.21 GiB │ MergeTree │ 2.21 GiB │ 76.73 GiB │ 0.028784 │
│ query_log │ 16384865 │ 1.91 GiB │ MergeTree │ 1.90 GiB │ 18.31 GiB │ 0.10398 │
│ part_log │ 17906105 │ 846.73 MiB │ MergeTree │ 845.39 MiB │ 3.84 GiB │ 0.21521 │
│ metric_log │ 4713151 │ 820.92 MiB │ MergeTree │ 806.13 MiB │ 14.56 GiB │ 0.05405 │
│ part_log_0 │ 15632289 │ 702.82 MiB │ MergeTree │ 701.70 MiB │ 3.34 GiB │ 0.20490 │
│ asynchronous_metric_log │ 795170674 │ 576.24 MiB │ MergeTree │ 562.50 MiB │ 11.11 GiB │ 0.049429 │
│ query_views_log │ 6597156 │ 461.35 MiB │ MergeTree │ 459.75 MiB │ 6.36 GiB │ 0.07060 │
│ logs │ 6448259 │ 408.59 MiB │ MergeTree │ 406.65 MiB │ 5.99 GiB │ 0.06627 │
│ samples_v2 │ 949110122 │ 345.01 MiB │ MergeTree │ 325.31 MiB │ 22.09 GiB │ 0.014382 │
If I was less stupid I'd get a machine with the recommended Clickhouse specs and save myself a few hours of tuning, but this works great.Downsides:
- clickhouse takes about 5 minute to start up because my tiny sc1 drive has like 4 IOPS allowed
- signoz's UI isn't amazing. It's totally functional, and they've been improving very quickly, but don't expect datadog-level polish
If anyone wants to check our project, here’s our GitHub repo - https://github.com/SigNoz/signoz
- I really like your new Logs & Traces Explorers. I spend a lot of time coming up with queries, and having a focused place for that is great. Especially since there's now a way to quickly turn my query into an alert or a dashboard item.
- You've also recently (6mo?) improved the autocomplete dramatically! This is awesome, and one of my annoyances with Datadog
Other feedback, and honestly this is all very minor. I'd be perfectly happy if nothing ever changed.
- where do I go see the metrics? There's no "Metrics" tab the way there's a "Logs" and "Traces" tab. A "Metrics Explorer" would be great.
- when I add a new plot, having to start out with a blank slate is not great. Datadog defaults to a generic system.cpu query just to fill something in, I find this helpful.
- when I have a plot in a dashboard and I see it is trending in the wrong direction, it would be nice to be able to create an alert directly from the chart rather than have to copy the query over.
- the exceptions tab is very helpful, but I've only recently discovered the LOW_CARDINAL_EXCEPTION_GROUPING flag. It'd be super nice if the variable part of exceptions was automatically detected and they were grouped
- once nice thing in DD is being able to preview a span from a log or logs from a span without opening a new page. Or previewing a span from the global page. Temporary popping this stuff up in a sidebar would be great.
- I'm not sure if there's a way to view only root spans in the trace viewer.
- This might be a problem with the spring boot instrumentation, but I can't see how to figure out what kind of span it is. Is it a `http.request`, `db.query`, etc?
> - where do I go see the metrics? There's no "Metrics" tab the way there's a "Logs" and "Traces" tab. A "Metrics Explorer" would be great.
Great, idea. This is some thing which few users have asked for and we will be shipping this in few releases
> - when I have a plot in a dashboard and I see it is trending in the wrong direction, it would be nice to be able to create an alert directly from the chart rather than have to copy the query over.
Fair point, this is something which is also in the pipeline.
> - I'm not sure if there's a way to view only root spans in the trace viewer.
We launched a tab in the new traces explorer for this, does it not serve your use case?
> - when I add a new plot, having to start out with a blank slate is not great. Datadog defaults to a generic system.cpu query just to fill something in, I find this helpful.
We can do something like this, but we don't necessarily know the name of metrics users are sending us, wrt. DataDog which has some default. metrics which their agents generate.
Will also look into other feedbacks you have given
At a former place, we were doing 5% of non-error traces.
Imagine you're sampling successful traces at, say, 1%, but sending all error traces. If your error rate is low, maybe also 1%, your trace volume will be about 2% of your overall request volume.
Then you push an update that introduces a bug and now all requests fail with an error, and all those traces get sampled. Your trace volume just increased 50x, and your infrastructure may not be prepared for that.
In general you don't want your system to do more work in a bad state. In fact as the AWS well architected guide say when you're overloaded or you're in a heavy air State you should be doing as little work as possible. So that you can recover
With head sampling, the first service in the request chain can make the decision about whether to trace, which can reduce tracing overhead on services further down.
With tail-based sampling, the tracing backend can make a determination about whether to persist the trace after the trace has been collected. This has tracing overheads, but allows you to make decisions like “always keep errors”.
I don’t mean for this to sound insulting but I honestly do not think this is an acceptable take to have as a developer.
Not knowing SQL is like refusing to learn any language that has classes in it, simply because you don’t like it.
I’ve heard stories of huge corporations failing product launches because some code was written to SELECT * from a database and filtering it in-app instead of doing the queries correctly, and what’s so fun with these types of issues is that they usually don’t appear until weeks later when the table has grown to a size where it becomes a problem.
When you’re saying that you’d rather find the data in-app than in-database, you’re putting the work on an inferior party in the transaction simply because you can’t be bothered.
The code will never* find the correct data faster than the database.
* there may be exceptions, but they’re far enough between to still say “never”.
Of course not all engineers can operate with SQL as efficiently as code -- that's the whole point. Otherwise why would we be writing code? Learning SQL intimately doesn't change that fact.
I'll admit I'm a little curious about what exactly you mean here.
If you've got a successful data hungry web service with a reasonably normalized schema and moderately complex access patterns though, you're not going to be looping over the whole thing on every page load.
We’re not talking about Assembly here, “dropping down” to SQL is something that anyone should be expected to do as soon as you’re grabbing or modifying any data from a database in any scenario where performance or integrity matters. The errors you can see in situations like this are extremely complex and databases literally exist to solve them for us.
Also, if we just completely disregard the performance for a second and focus on data security instead, how do you ensure sensitive data isn’t passed to the wrong party if you don’t care about what queries are being sent?
I mean, it doesn’t matter if it’s not “in the end” displayed to an end user in the application you’re writing, or if its not stored in the intermediary node where your code is running, that data is now unnecessarily on the wire in a situation where it never should have been in the first place. If you end up mixing one customers data with another’s and sending all of it in such a way that it could even theoretically be accessed by a third party, that’s a lawsuit waiting to happen regardless of whether it was “displayed” or “forwarded” or not.
Imagine if you sniffed the packets going to some logistics app you use on your phone and you saw meta-data for all packages in your zip code in the response, or if some widget showing you your carbon footprint actually was based on a response containing the carbon footprint of every customer in the database. Even if it’s just [user_id,co2] it’s still completely unacceptable.
Never mind scenarios where you’re modifying, adding or deleting data, those are even worse and no explanation should be necessary for why.
This opens some interesting options if you want to join the result with data from your database.
To be fair, we wanted to get experience on ClickHouse and it's a special database need special attention to details on both ops and schema design.
A decent sized server will host a hugely capable instance that you may not have to think about for years. The scoffing down at DIY has made sense to some degree, but it just works brilliantly keeps getting to be a stronger & stronger case & most just assume reality can't actually work that well, that it'll be bad, and those folks won't always be right.
SaaS, especially in this space, can be *extremely* costly and its cost will scale up quickly as you send more traffic (either willingly or by mistake). Yes, Datadog, NewRelic etc will give you many pre-built and well-thought dashboards and some fancy AI-powered auto-detection thing but they will charge many $$$ for it. Consider that now cost management/analysis tools that were historically focused only on cloud, are now adding the same tooling for costly SaaS solutions!
I understand that many HN readers are skewed towards SaaS solutions, usually because they work at a SaaS shop, but depending on the size of the company, the overhead for managing it internally can totally be worth. There is overhead with SaaS as well...