HNHacker News
TopNewBestAskShowJobs

BohuTANG

26 karma · joined October 3, 2011

CEO at DatabendLabs | Data Warehouse, Data Cloud
submissionscomments
BohuTANG··on Better Models: Worse Tools
Suggestion for Pi: capitalize tool names for the Sonnet/Opus models (edit -> Edit, bash -> Bash, ...).

The rationale: Anthropic's own harness (Claude Code) uses PascalCase tool names — Bash, Edit, Read, Write, Glob, Grep. Since the models are post-trained/aligned against that harness, those naming conventions are effectively baked into the model. Matching your harness's tool names to the same casing puts your inputs closer to the training distribution, which lines up with the more reliable tool use I've seen in evaling.

A related pattern that fits the same distribution: for long outputs, have the model reserve placeholders first and complete the work across multiple steps.

Reference: https://github.com/evotai/evot/commit/765151796c43965964a9da...

BohuTANG··on Snowtree: Review-Driven Safe AI Coding
Hey everyone, I'm excited to open-source Snowtree: my tool for safe AI coding!

Team up with AI like Claude or Codex in your IDE—stay in control with isolated sessions, easy reviews, and snapshot saves. No chaos, just smarter workflows.

BohuTANG··on Databricks acquires Neon
Yes, Databend's metadata service (https://github.com/databendlabs/databend/tree/main/src/meta) also uses a Raft implementation written in Rust: https://github.com/databendlabs/openraft.

We've been running OpenRaft in production for several years now and have found it to be quite stable. It's designed as a generic, feature-complete Raft library that handles the complexities of distributed consensus well. If you're looking for a mature Rust Raft implementation, it's definitely worth considering.

BohuTANG··on Databricks acquires Neon
Neon (open-source alternative to Aurora) is 73.6% Rust. Databend (open-source Snowflake alternative) is even more Rust-heavy at 97.2%.

Interesting trend - modern serverless databases choosing Rust for its memory safety, performance predictability. Makes sense for systems where reliability and efficiency are non-negotiable.

BohuTANG··on Show HN: Query Fast - Lightweight data analytics with AI
BendSQL is a cli tool.

I suggest to use the drivers list here: https://docs.databend.com/guides/sql-clients/developers/

BohuTANG··on Show HN: Query Fast - Lightweight data analytics with AI
Thanks, sverg! From my experience, query-fast is a better product than many other AI-to-SQL solutions. I'm really looking forward to this product.
BohuTANG··on Show HN: Query Fast - Lightweight data analytics with AI
Interesting, I tried query-fast and it's impressive! It would be great to add Databend as a data source—it's a cost-efficient alternative to Snowflake: https://www.databend.com/
BohuTANG··on Why Choose Rust as Your Development Language
Rust is great for building databases because it ensures memory safety, avoids garbage collection, and prevents memory leaks and data races. Databend is a prime example—it's one of the largest database projects written in Rust, with over 476,000 lines of code developed from scratch over the past three years. While Rust isn't as widely adopted as some other languages, its use in projects like Databend demonstrates its potential for creating high-performance, reliable systems.
BohuTANG··on A new JSON data type for ClickHouse
Yes, reliable data ingestion often involves Kafka, which can feel complex. An alternative is the transactional COPY INTO approach used by platforms like Snowflake and Databend. This command supports "exactly-once" ingestion, ensuring data is fully loaded or not at all, without requiring message queues or extra infrastructure.

https://docs.databend.com/sql/sql-commands/dml/dml-copy-into...

BohuTANG··on DuckDB vs. Snowflake vs. Databricks
Looking forward to adding Databend, the open-source alternative to Snowflake: https://github.com/databendlabs/databend/issues/13059
BohuTANG··on How Do People Use Snowflake and Redshift?
The snowset[1] in the post is quite dated—it’s based on data from February 21 to March 7, 2018, which doesn't reflect current usage trends. Many users find Snowflake too expensive now, especially with the added costs of ETL. With the rise of open-source alternatives like Databend[2], users have more options when considering a move away from Snowflake.

[1]snowset: https://github.com/resource-disaggregation/snowset

[2]Databend: https://github.com/datafuselabs/databend

BohuTANG··on Top 5 Snowflake Alternatives for Cost-Effective Data Warehousing
Check out Databend, an open-source alternative to Snowflake with 90%+ product and syntax similarity (almost no migration cost). It’s cost-effective and powerful.

GitHub: https://github.com/datafuselabs/databend/issues/13059

Comparison: https://www.databend.com/comparison

BohuTANG··on TIL: Versions of UUID and when to use them
UUID v7's timestamp is a game-changer for Databend. We're using it to quickly locate metadata files on AWS S3 by timestamp, making operations like vacuuming much faster.

PR: https://github.com/datafuselabs/databend/pull/16049

BohuTANG··on FoundationDB at Snowflake: Architecture and Internals (2021) [video]
Databend is a cutting-edge alternative to Snowflake, offering cost-effective and straightforward solutions for large-scale analytics. It utilizes a lightweight meta service instead of the resource-intensive FoundationDB and employs OpenRaft for replication and high availability.

- Databend: https://github.com/datafuselabs/databend

- OpenRaft: https://github.com/datafuselabs/openraft

BohuTANG··on Building a dynamic lib plugin system for Rust
Great post! The sandboxing and pluggability of UDFs are key. Databend now supports Python, JavaScript, and WebAssembly to simplify data processing.

[1] Embedded UDFs: https://docs.databend.com/guides/query/udf#embedded-udfs

[2] https://x.com/DatabendLabs/status/1796830863344463884

BohuTANG··on Show HN: Turn CSV Files into SQL Statements for Quick Database Transfers
Yeah. The insert statement fails when dealing with larger datasets, and the parser incurs significant overhead. The most efficient method is to use the COPY INTO command like in Databend[1] or Snowflake[2] for bulk data loading.

[1] Databend: https://docs.databend.com/guides/load-data/load-semistructur...

[2] Snowflake: https://docs.snowflake.com/en/sql-reference/sql/copy-into-ta...

[3] Data Ingestion Benchmark: https://docs.databend.com/guides/benchmark/data-ingest

BohuTANG··on Show HN: Hacker Search – A semantic search engine for Hacker News
Nice work. It's recommended to start with a keyword search (including titles, comments, and the keyword itself) before using an LLM-based search. For example, if I search for 'databend' on https://hackersearch.net/search?q=databend&period=last-3-yea..., it defaults to an LLM-based search. However, I prefer using a simple keyword search instead.
BohuTANG··on Show HN: My AI writing assistant for Chinese
Interesting. I use your promot and create a GPT[1], give an example: `I服了you`:

服了u: Correctly translates to "I give up on you." Unscramble: Incorrectly interpreted as "You impress me."

- [1] https://chat.openai.com/g/g-tI0XLZxuR-fu-liao-u

BohuTANG··on Are you sure you want to use MMAP in your DBMS?
Please don't misunderstand me; I've also implemented tokudb_buffer_pool_size for TokuDB because their control over memory is excellent, not falling into the black hole of the page cache. This allows my DBMS to be stable enough. This is also my point in this post, and my discussion ends here.
BohuTANG··on Are you sure you want to use MMAP in your DBMS?
Thank you for the insight.

We manage memory directly within the DBMS, for instance, by setting MySQL's innodb_buffer_pool_size to 1GB. This precise control method ensures that each cloud-based DBMS instance on our large EC2 instances remains within its memory limit without adversely affecting others.

When the internal memory pool limit of the DBMS is reached, we understand that it will spill dirty pages to disk and utilize the write-ahead log to guarantee ACID properties in the event of a crash. This process is predictable and manageable. In contrast, the effects on the DBMS when cgroup memory limits are reached are less clear, introducing uncertainty into system behavior under memory pressure. This unpredictability strongly supports our preference for managing memory within the DBMS itself.

Our experience over a decade with cloud DBMS products has shown that mmap doesn't suit our needs. We prioritize direct memory management methods within the DBMS to ensure system reliability and optimal performance.

BohuTANG··on Are you sure you want to use MMAP in your DBMS?
Imagine you're using a EC2 to run several DBMS instances at the same time, each for a different customer. How do you make sure that each instance only uses a certain amount of memory and doesn't affect the others?
BohuTANG··on Are you sure you want to use MMAP in your DBMS?
These articles all miss the point.

If you are a DBMS kernel developer, the most urgent need is how to control memory more precisely to prevent out-of-memory (OOM) situations. Clearly, MMAP cannot achieve this, which is also why most databases still use their own memory pool + direct-IO.

BohuTANG··on Solutions to manage runaway Snowflake costs?
TPC-H Benchmark: Databend Cloud vs. Snowflake: https://docs.databend.com/guides/benchmark/tpch
BohuTANG··on Solutions to manage runaway Snowflake costs?
Managing costs with Snowflake's virtual warehouse is more straightforward, allowing you to adjust the warehouse's size and operational status to control expenses. However, optimizing costs for Snowflake's cloud service is more complex. Each SQL query incurs costs for parsing and planning. These costs, less transparent than warehouse expenses, increase as your business and request volume grow, making budgeting and cost control challenging.

To address this, consider DatabendCloud – a cost-effective alternative to Snowflake, built entirely in Rust.

Key Benefits of DatabendCloud:

*Significant cost savings, over 50% compared to Snowflake.

*Similar user experience to Snowflake, with comparable syntax and product functionalities.

*Minimal to no cloud service fees; costs are primarily based on warehouse usage.

*Transparent pricing model, allowing you to easily estimate expenses at: https://www.databend.com/plan/

Links:

[1] DatabendCloud: https://databend.com/

[2] Databend vs. Snowflake: https://github.com/datafuselabs/databend/issues/13059

[3] TPC-H SF100: https://twitter.com/DatabendLabs/status/1714456893941526761

BohuTANG··on The Cost of the 1B Row Challenge with Snowflake and Databend
With Snowflake and #Databendtackling the #1BRC: - How fast can we process? - What's the cost?
BohuTANG··on Moving from DynamoDB to tiered storage with MySQL+S3
Interesting! However, it's important to note that S3 Select isn't ideal for handling large-scale data scans, as this can lead to exorbitant costs. In contrast, Databend(https://github.com/datafuselabs/databend) utilizes S3 primarily for storage purposes. Key functionalities like bloom filter indexing, min-max indexing, aggregating indexing, join reordering, and filter pushdown are all managed within the query layer. This approach optimizes performance and cost-efficiency, particularly when dealing with substantial amounts of data.
BohuTANG··on What's the Future of Unified Cloud Storage?
Yes, we need a unified API layer like OpenDAL: https://github.com/apache/incubator-opendal
BohuTANG··on Summing columns in remote Parquet files using DuckDB
For http(s) remote file, we can support glob pattern, for example: https://<thehfurl>/resolve/main/data/0000[00-55].parquet'

Databend supports this pattern: https://databend.rs/doc/load-data/load/http#loading-with-glo...

BohuTANG··on Summing columns in remote Parquet files using DuckDB
Cool, thanks for the insight. Analyzing HuggingFace datasets is now easier: https://twitter.com/DatabendLabs/status/1725361248509067434
BohuTANG··on Show HN: Lexeme – open-source ChatGPT text editor
Great, I'm using it now. If Lexeme is similar to Typora (https://typora.io), it could be fantastic and might even surpass Typora in terms of quality. On the other hand, if Typora already has these features, it's quite powerful.
Page 1 of 2Next →