61 karma · joined May 24, 2012
We published a public demo of the Agentic Data Stack, I'd love to hear your feedback https://clickhouse.com/blog/agenthouse-demo-clickhouse-llm-m...
Keep in mind that it's not fully "fair", since these public dataset are often documented in the internet so already present in pre-training of the models underneath (Claude Sonnet 4.5 in this case)
Our own experience running internal agents taught us that the best remediation comes from providing the LLMs with the maximum and most accurate context possible. Robust evaluations are also critical to measure accuracy, detect regressions, and improve. But there is no silver bullet.
SOTA LLMs are increasingly better at generating SQL and notoriously bad with math and numbers in general. Combining them with powerful querying capabilities bridges that gap and makes the overall experience an useful one.
IMO, we'll always have to deal with the stochastic nature of these models and hallucinations, which calls for caution and requires raising awareness within the user base. What I found watching our users internally is that, while it's not magical, it allows users to request data more often, and compounds in data-driven decision-making, assuming the users are trained to interpret the interactions
It's a fair concern, and I understand where you are coming from. What I can say is that it's not our first rodeo incorporating another OSS product in our family. I tried to summarize it in the post:
> "This proven playbook is the same one that we applied when joining forces with PeerDB to provide our ClickPipes CDC capabilities, and HyperDX, which became the UX of our observability product, ClickStack."
If you research both instances above, the result is that these projects got more traction and adoption overall.
I hope this helps! and thank you for using LibreChat
My favourite use-case: our sales and support folks systematically ask DWAINE (our dwh agent) to produce a report before important meetings with customers, something along the lines of: "I'm meeting with <customer_name> for a QBR, what do I need to know?". This will pull usage data, support interactions, billing, and many other dimensions, and you can guess that the quality of the conversation is greatly improved.
My colleague Dmitry wrote about it when we first deployed it: https://www.linkedin.com/pulse/bi-dead-change-my-mind-dmitry...
So, why this move ?
Basically, we noticed that the existing "agentic" open-source ecosystem is primarily focused on developer tools and SDKs, as developers are the early adopters who build the foundation for emerging technologies. Current projects provide frameworks, orchestration, and integrations The idea behind the Agentic Data Stack is a higher-level integration to provide a composable software stack for agentic analytics that users can setup quicky, with room for customization.
ps. I work for ClickHouse and happy to help
Ps. I work for ClickHouse
I'm curious to hear your take about the tradeoffs that the HTAP model introduces? any impact on ingest times or query throughput for example?
Feature differentiation is actually a pretty interesting topic for o11y. You can do many things with an OLAP store but you need to be aware of the differences with of the shelf solutions, I try to summarize it here: https://clickhouse.com/blog/the-state-of-sql-based-observabi...
I hope this helps! I'd love to hear your opinion about it
It's fairly technical and goes into details but overall it displayed how to achieve increased compression and query performance
Another important dimension I'd consider as well is the broader ecosystem of integrations. In OSS, it's a byproduct of the success of the main project but an often overlooked aspect when choosing a solution.
Eg. Here are some of the Kafka integrations options for CH https://clickhouse.com/docs/en/integrations/kafka
https://clickhouse.com/docs/en/operations/utilities/clickhou... https://clickhouse.com/docs/en/getting-started/quick-start
Disclaimer: I work at ClickHouse
https://clickhouse.com/cloud/clickpipes https://www.youtube.com/watch?v=rSUHqyqdRuk https://clickhouse.com/docs/en/integrations/clickpipes
Don't hesitate to reach out if you have any questions or feedback!
(I work at ClickHouse)
https://clickhouse.com/docs/en/sql-reference/functions/dista...
Fyi, we recently improved the Parquet support in 23.2 https://github.com/ClickHouse/ClickHouse/pull/45878
Also, we still have optimizations for reading Parquet from S3 coming so that might improve
Here's a example of a query against a Parquet file you can run on your laptop: ``` ./clickhouse local -q "SELECT town, avg(price) AS avg_price FROM s3('https://datasets-documentation.s3.eu-west-3.amazonaws.com/ho...') GROUP BY town ORDER BY avg_price DESC LIMIT 10" ``` From: https://clickhouse.com/docs/knowledgebase/parquet-to-csv-jso...
From the blog post: "the fastest baseline here is ClickHouse server running on an AWS m5d.24xlarge instance that uses 48 threads for query execution. As you can see, an equivalent cloud service with 48 threads performs very well in comparison for a variety of simple and complex queries represented in the benchmark" so there can be a small difference depending on the sizing but it's something to consider on a case per case basis and often other dimensions need to be taken into account (operational cost, bottomless storage, linear scalability, reliability etc.).