HNHacker News
TopNewBestAskShowJobs

ryadh

61 karma · joined May 24, 2012

submissionscomments
ryadh··on [dead]
An agentic analytics exploration of some tech hype cycles using HackerNews, GitHub, and Stack Overflow data
ryadh··on ClickHouse acquires LibreChat, open-source AI chat platform
Thanks for your detailed reply. It is great to see that you have been experimenting with this approach.

We published a public demo of the Agentic Data Stack, I'd love to hear your feedback https://clickhouse.com/blog/agenthouse-demo-clickhouse-llm-m...

Keep in mind that it's not fully "fair", since these public dataset are often documented in the internet so already present in pre-training of the models underneath (Claude Sonnet 4.5 in this case)

ryadh··on ClickHouse acquires LibreChat, open-source AI chat platform
(Ryadh from ClickHouse here) Your comment is spot-on. This the main challenge with Agentic Analytics and there are known limitations. It is also where we are orienting our investments atm.

Our own experience running internal agents taught us that the best remediation comes from providing the LLMs with the maximum and most accurate context possible. Robust evaluations are also critical to measure accuracy, detect regressions, and improve. But there is no silver bullet.

SOTA LLMs are increasingly better at generating SQL and notoriously bad with math and numbers in general. Combining them with powerful querying capabilities bridges that gap and makes the overall experience an useful one.

IMO, we'll always have to deal with the stochastic nature of these models and hallucinations, which calls for caution and requires raising awareness within the user base. What I found watching our users internally is that, while it's not magical, it allows users to request data more often, and compounds in data-driven decision-making, assuming the users are trained to interpret the interactions

ryadh··on ClickHouse acquires LibreChat, open-source AI chat platform
Interestingly, LibreChat has a broad range of applications already and we'll continue to support them. The investment area we want to tackle in priority is around the analytics use-case specifically.In that space, I don't see an SSO-tax scheme unfolding tbh, it's really about better visualizations, semantic layers and anything that can improve the quality of the insights produced on top of analytics data
ryadh··on ClickHouse acquires LibreChat, open-source AI chat platform
Ryadh from ClickHouse here, I commented below about the overall intent. Let me know if anything needs clarifying!
ryadh··on ClickHouse acquires LibreChat, open-source AI chat platform
The LibreChat folks are now my colleagues, and it's exciting
ryadh··on ClickHouse acquires LibreChat, open-source AI chat platform
(Ryadh from ClickHouse here)

It's a fair concern, and I understand where you are coming from. What I can say is that it's not our first rodeo incorporating another OSS product in our family. I tried to summarize it in the post:

> "This proven playbook is the same one that we applied when joining forces with PeerDB to provide our ClickPipes CDC capabilities, and HyperDX, which became the UX of our observability product, ClickStack."

If you research both instances above, the result is that these projects got more traction and adoption overall.

I hope this helps! and thank you for using LibreChat

ryadh··on ClickHouse acquires LibreChat, open-source AI chat platform
It actually comes from our own experience at ClickHouse. We deployed this stack internally 8 months ago, and since very few people here have touched our legacy BI systems :) I have never seen an adoption curve like this one tbh. It's obviously not perfect and can hallucinate sometimes, which can be tricky, but with the right approach and awareness in place, the value it delivers is massive. What really happens is that more users get access to data instantly, and as a result, we make better, data-driven, decisions overall.

My favourite use-case: our sales and support folks systematically ask DWAINE (our dwh agent) to produce a report before important meetings with customers, something along the lines of: "I'm meeting with <customer_name> for a QBR, what do I need to know?". This will pull usage data, support interactions, billing, and many other dimensions, and you can guess that the quality of the conversation is greatly improved.

My colleague Dmitry wrote about it when we first deployed it: https://www.linkedin.com/pulse/bi-dead-change-my-mind-dmitry...

ryadh··on ClickHouse acquires LibreChat, open-source AI chat platform
Ryadh from ClickHouse here, happy to answers questions if folks have any.

So, why this move ?

Basically, we noticed that the existing "agentic" open-source ecosystem is primarily focused on developer tools and SDKs, as developers are the early adopters who build the foundation for emerging technologies. Current projects provide frameworks, orchestration, and integrations The idea behind the Agentic Data Stack is a higher-level integration to provide a composable software stack for agentic analytics that users can setup quicky, with room for customization.

ryadh··on Launch HN: Inconvo (YC S23) – AI agents for customer-facing analytics
Congrats on the launch, any plans to support ClickHouse?

ps. I work for ClickHouse and happy to help

ryadh··on ClickHouse gets lazier and faster: Introducing lazy materialization
That's great feedback, thank you! I just added your comment to the GH issue: https://github.com/chdb-io/chdb/issues/101#issuecomment-2824...

Ps. I work for ClickHouse

ryadh··on Agent-Facing Analytics
Author here, happy to take any questions!
ryadh··on ClickHouse Acquires PeerDB
Congrats on Tablespace! it's good to see innovation in that space and I always love to see ClickBench mentioned in the wild :) (I work for ClickHouse).

I'm curious to hear your take about the tradeoffs that the HTAP model introduces? any impact on ingest times or query throughput for example?

ryadh··on Show HN: ADS-B visualizer
The readme contains a lot of the implementation details: "We use three different tables with different levels of detail: planes_mercator contains 100% of the data, planes_mercator_sample10 contains 10% of the data, and planes_mercator_sample100 contains 1% of the data. The loading starts with a 1% sample to provide instant response even while rendering the whole world. After loading the first level of detail, it continues to the next level of 10%, and then it continues with 100% of the data. This gives a nice effect of progressive loading."
ryadh··on We Built a 19 PiB Logging Platform with ClickHouse and Saved Millions
We explored a scale similar to what you described in another blog: https://clickhouse.com/blog/cost-predictable-logging-with-cl...

Feature differentiation is actually a pretty interesting topic for o11y. You can do many things with an OLAP store but you need to be aware of the differences with of the shelf solutions, I try to summarize it here: https://clickhouse.com/blog/the-state-of-sql-based-observabi...

I hope this helps! I'd love to hear your opinion about it

ryadh··on Ask HN: How to properly build a multi-terabyte DuckDB database?
You are very welcome
ryadh··on Ask HN: How to properly build a multi-terabyte DuckDB database?
Have you considered ClickHouse?
ryadh··on Running Redshift at Scale
We did a deep-dive comparison between Redshift and ClickHouse Cloud for analytics workloads at scale: https://clickhouse.com/blog/redshift-vs-clickhouse-compariso...

It's fairly technical and goes into details but overall it displayed how to achieve increased compression and query performance

ryadh··on Thoughts on OLAPs ClickHouse vs. Apache Druid vs. Starrocks in 2023/2024
(disclaimer: I work at ClickHouse)

Another important dimension I'd consider as well is the broader ecosystem of integrations. In OSS, it's a byproduct of the success of the main project but an often overlooked aspect when choosing a solution.

Eg. Here are some of the Kafka integrations options for CH https://clickhouse.com/docs/en/integrations/kafka

ryadh··on ChDB: Embedded OLAP SQL Engine Powered by ClickHouse
ChDB is the in-process version for Python. You can try clickhouse-local if you want a CLI experience, or clickhouse-client against a clickhouse-server for the server experience

https://clickhouse.com/docs/en/operations/utilities/clickhou... https://clickhouse.com/docs/en/getting-started/quick-start

ryadh··on ChDB: Embedded OLAP SQL Engine Powered by ClickHouse
Here is an example, DB Pilot recently switched from duckdb to ChDB, mostly for faster queries and broader data formats support: https://dbpilot.io/changelog#embedded-clickhouse-and-standal...

Disclaimer: I work at ClickHouse

ryadh··on Turning an OLAP database into a fully-fledged data hub (ClickHouse PolyVLDB 23) [video]
ClickHouse is an established open-source columnar OLAP database that has gained significant popularity due to its ability to handle large-scale real-time analytical workloads without compromising speed and efficiency. As users and organizations increasingly adopted ClickHouse for their data processing needs, they also shaped the open-source project with an ever growing list of integration capabilities turning it into a powerful data hub. In this presentation, we'll display how ClickHouse's integration engines (a concept to provide virtual tables to remote data stores like OLTP systems and datalakes) empower real-world use-cases from simply importing data or ad-hoc querying to leveraging powerful algorithms to perform data grouping locally. We will also explore how external dictionaries can be leveraged to augment ClickHouse's querying capabilities by integrating with external data sources and look-up tables. Finally, we will cover ClickHouse’s materialized views, focusing on their role in precomputing and aggregating data to optimize query performance and provide continuous transformations on top of heterogeneous datasets.
ryadh··on ClickPipes: Seamlessly Connect Kafka to ClickHouse
You can follow the instructions in the docs to create your first ClickPipes. More details are available in:

https://clickhouse.com/cloud/clickpipes https://www.youtube.com/watch?v=rSUHqyqdRuk https://clickhouse.com/docs/en/integrations/clickpipes

Don't hesitate to reach out if you have any questions or feedback!

ryadh··on Can you even trust benchmarks these days? ClickHouse vs. Druid vs. Rockset
One way to add some trust is to make benchmarks open-source and reproducible: https://github.com/ClickHouse/ClickBench/

(I work at ClickHouse)

ryadh··on [dead]
Discover the powerful capabilities of the ClickHouse Foreign Data Wrapper (FDW) and Postgres Table Engine. In this webinar, we’ll display how we built a user facing real-estate application that combines using both Supabase and ClickHouse Cloud as backends, using the best of both worlds
ryadh··on Ask HN: Seeking a Vector Database for ClickHouse Users – Suggestions Appreciated
ClickHouse can actually store vectors as tuples or arrays. It also comes with some handy distance functions

https://clickhouse.com/docs/en/sql-reference/functions/dista...

ryadh··on Parquet: An efficient, binary file format for table data
Interesting, thanks for sharing this feedback! I didn't realise intially that it was about running clickhouse-local in Lambdas.

Fyi, we recently improved the Parquet support in 23.2 https://github.com/ClickHouse/ClickHouse/pull/45878

Also, we still have optimizations for reading Parquet from S3 coming so that might improve

ryadh··on Parquet: An efficient, binary file format for table data
Interesting, why it would take so long ? I'm genuinely curious so that we can make it better (Disclaimer:I work at ClickHouse)

Here's a example of a query against a Parquet file you can run on your laptop: ``` ./clickhouse local -q "SELECT town, avg(price) AS avg_price FROM s3('https://datasets-documentation.s3.eu-west-3.amazonaws.com/ho...') GROUP BY town ORDER BY avg_price DESC LIMIT 10" ``` From: https://clickhouse.com/docs/knowledgebase/parquet-to-csv-jso...

ryadh··on Building ClickHouse Cloud from scratch in a year
That's a fair question and something we obsess about at ClickHouse! It's mentioned in the post but we maintain continuously updated benchmarks for ClickHouse Cloud and it's on-prem shared nothing counterpart (as well as other databases!). You'll find in [1] the comparison for every ClickHouse deployment option.

From the blog post: "the fastest baseline here is ClickHouse server running on an AWS m5d.24xlarge instance that uses 48 threads for query execution. As you can see, an equivalent cloud service with 48 threads performs very well in comparison for a variety of simple and complex queries represented in the benchmark" so there can be a small difference depending on the sizing but it's something to consider on a case per case basis and often other dimensions need to be taken into account (operational cost, bottomless storage, linear scalability, reliability etc.).

[1] https://tinyurl.com/chbench

ryadh··on Building ClickHouse Cloud from scratch in a year
Good catch! thanks for letting us know
Page 1 of 2Next →