HNHacker News
TopNewBestAskShowJobs

tison

431 karma · joined December 17, 2021

Co-Founder @ ScopeDB

Board Member & Incubator Mentor @ The Apache Software Foundation

I develop and maintain open-source software: https://github.com/tisonkun

submissionscomments
tison··on A cloud-native database should be as elastic as the cloud itself
We have discussed these in previous blogs:

1. Insight In No Time: https://www.scopedb.io/blog/insight-in-no-time

2. Manage Data in Petabytes for an Observability Platform: https://www.scopedb.io/blog/manage-observability-data-in-pet...

That is, Snowflake follows a traditional data warehouse workflow that requires extenral ETL process to load data into the warehouse. Some of our customers did researching of Snowflake and noticed that their event streaming ingestion can not fit in Snowflake's stage-based loading model - they need real-time insights end-to-end.

Apart from this major downside, about leveraging S3 as a primary storage, Snowflake doesn't have adaptive indexes, and its performance would be significantly degraded as data grows and queries involve a large range of data + multi-condition filters when the simple minmax index can't help.

tison··on Stop Forwarding Errors, Start Designing Them
FWIW, here is a general discussion about error handling in Rust and my comment to compare it with Go's/Java's flavor: https://github.com/apache/datasketches-rust/issues/27#issuec...

That said, I can live with "if err != nil", but every type has a zero value is quite a headache to handle: you would fight with nil, typed nil, and zero value.

For example, you need something like:

  type NullString struct {
   String string
   Valid  bool // Valid is true if String is not NULL
  }
.. to handle a nullable value while `Valid = false && String = something` is by defined invalid but .. quite hard to explain. (Go has no sum type in this aspect)
tison··on Stop Forwarding Errors, Start Designing Them
I think they are almost compatible.

`thiserror` helps you define the error type. That error type can then be used with `anyhow` or `exn`. Actually, we have been using thiserror + exn for a long time, and it works well. While later we realize that `struct ModuleError(String)` can easily implement Error without thiserror, we remove thiserror dependency for conciseness.

`exn` can use `anyhow::Error` as its inner Error. However, one may use `Exn::as_error` to retrieve the outermost error layer to populate anyhow.

I ever consider `impl std::error::Error` for `exn::Exn,` but it would lose some information, especially if the error has multiple children.

`error-stack` did that at the cost of no more source:

* https://docs.rs/error-stack/0.6.0/src/error_stack/report.rs....

* https://docs.rs/error-stack/0.6.0/src/error_stack/error.rs.h...

tison··on Stop Forwarding Errors, Start Designing Them
This is the pull request of this post: https://github.com/fast/fast.github.io/pull/12

See comments like https://github.com/fast/fast.github.io/pull/12#discussion_r2...

Quote my comment in the other thread:

> That said, exn benefits something from anyhow: https://github.com/fast/exn/pull/18, and we feed back our practices to error-stack where we come from: https://github.com/hashintel/hash/issues/667#issuecomment-33...

> While I have my opinions on existing crates, I believe we can share experiences and finally converge on a common good solution, no matter who made it.

tison··on Cancellations in async Rust
Rust's Future is somehow like move semantics in C++, where you may leave a Future in an invalid state after it finishes. Besides, Rust adopts a stackless coroutine design, so you need to maintain the state in your struct if you would like to implement a poll-based async structure manually.

These are all common traps. And now cancellations in async Rust are a new complement to state management in async Rust (Futures).

When I'm developing the mea (Make Easy Async) [1] library, I document the cancel safety attribute when it's non-trivial.

Additionally, I recall [2] an instance where a thoughtless async cancellation can disrupt the IO stack.

[1] https://github.com/fast/mea

[2] https://www.reddit.com/r/rust/comments/1gfi5r1/comment/luido...

tison··on Show HN: ScopeQL, a new query language based on relational algebra
Not quite. As described in the FAQ:

What about the interoperability with SQL?

Some libraries and tools enable developers to write queries in a new syntax and translate them to SQL (e.g., PRQL, SaneQL, etc.). The existing SQL ecosystem provides solid database implementations and a rich set of data tools. People always tend to think you must speak SQL; otherwise, you lose the whole ecosystem.

But wait a minute, those libraries translate their new language to SQL because they don't implement the query engine (i.e., the database) themselves, so they have to talk to SQL databases in SQL. However, ScopeQL is the query language of ScopeDB, and ScopeDB is already a database built directly on top of S3.

Thus, what we can leverage from the SQL ecosystem are data tools, such as BI tools, that generate SQL queries to implement business logic. For this purpose, one should write a translator that converts SQL queries to ScopeQL queries. Since both ScopeQL and SQL are based on relational algebra, the translation must be doable.

tison··on Why Not SQL: The Origin of ScopeQL
The syntax is still changable and welcomes any comments for improvements.

To try my best to avoid divergent discussion, I'd include two most significant FAQs:

*What about the interoperability with SQL?*

Some libraries and tools enable developers to write queries in a new syntax and translate them to SQL (e.g., [PRQL](https://prql-lang.org/), [SaneQL](https://www.cidrdb.org/cidr2024/papers/p48-neumann.pdf), etc.). The existing SQL ecosystem provides solid database implementations and a rich set of data tools. People always tend to think you must speak SQL; otherwise, you lose the whole ecosystem.

But wait a minute, those libraries translate their new language to SQL because they don't implement the query engine (i.e., the database) themselves, so they have to talk to SQL databases in SQL. However, ScopeQL is the query language of ScopeDB, and ScopeDB is already a database built directly on top of S3.

Thus, what we can leverage from the SQL ecosystem are data tools, such as BI tools, that generate SQL queries to implement business logic. For this purpose, one should write a translator that converts SQL queries to ScopeQL queries. Since both ScopeQL and SQL are based on relational algebra, the translation must be doable.

*Project Foo has already implemented similar features. Why not follow them?*

ScopeQL was developed from scratch but was not invented in isolation. We learn a lot from existing solutions, research, and discussions with their adopters. It includes the syntax of PRQL, SaneQL, and SQL extensions provided by other analytical databases. We also deeply empathize with the challenges outlined in the [GoogleSQL](https://research.google/pubs/sql-has-problems-we-can-fix-the...) paper.

However, as answered in the previous question, we first developed ScopeDB as a relational database. Then, we learned users' scenarios where an enhanced syntax helps maintain their business logic and increases their productivity. So, directly implementing the enhanced syntax is the most efficient way.

tison··on Show HN: Make Easy Async Rust (Mea), runtime-agnostic primitives
Welcome to create an issue on GitHub for sharing and discussion :D
tison··on Show HN: Make Easy Async Rust (Mea), runtime-agnostic primitives
I don't have a dedicated benchmark for these primitives, but we use them in a database that processes petabytes of data [1] and we don't find specific bottlenecks.

[1] https://www.scopedb.io/blog/manage-observability-data-in-pet...

Most of the performance factors would be the sync Mutex in used. I can imagine that by switching between the std Mutex, parking_lot's Mutex, and perhaps spin lock in some scenarios, one can gain better performance. Mea has an abstraction (src/internal/mutex.rs) for this switch, but I don't implement the feature flag for the switch since the current performance is acceptable in our use case.

The internal semaphore's implementation may be improved also. Currently, to keep code safe, I implement the linked list with `Slab<Node>` (you can check src/internal/waitlist.rs for details). Using a link like [2] may help, but that's not always a net win and needs much more time to do it right.

[2] https://github.com/Amanieu/intrusive-rs

tison··on Build a Database in Four Months with Rust and 647 Open-Source Dependencies
> they were probably just trying to be humble about their accomplishment

Thanks for your reply. To be honest, I simply recognize that depending on open-source software a trivial choice. Any non-trivial Rust project can pull in hundreds of dependencies and even when you audit distributed system written in C++/Java, it's a common case.

For example, Cloudflare's pingora has more than 400 dependencies. Other databases written in Rust, e.g., Databend and Materialize, have more than 1000 dependencies in the lockfile. TiKV has more than 700 dependencies.

People seem to jump in the debt of the number of dependencies or blame why you close the source code, ignoring the purpose that I'd like to show how you can organically contribute to the open-source ecosystem during your DAYJOB, and this is a way to write open-source code sustainable.

tison··on Build a Database in Four Months with Rust and 647 Open-Source Dependencies
And contributing back is one of the approaches to maintaining open-source dependencies. I have described how to deal with OSS dependencies in [1] (yet to translate it :P).

[1] https://www.tisonkun.org/2024/11/17/open-source-supply-chain...

tison··on Build a Database in Four Months with Rust and 647 Open-Source Dependencies
This article is actually a translated one. In the original article[1], I talked about commercial open-source and how one can collaborate with the open-source community when running a software business.

This section is moved to the second-to-last section in the posted blog, including:

[QUOTE]

When you read The Cathedral & the Bazaar, for its Chapter 4, The Magic Cauldron, it writes:

> … the only rational reasons you might want them to be closed is if you want to sell the package to other people, or deny its use to competitors. [“Reasons for Closing Source”]

> Open source makes it rather difficult to capture direct sale value from software. [“Why Sale Value is Problematic”]

While the article focuses on when open-source is a good choice, these sentences imply that it’s reasonable to keep your commercial software private and proprietary.

We follow it and run a business to sustain the engineering effort. We keep ScopeDB private and proprietary, while we actively get involved and contribute back to the open-source dependencies, open source common libraries when it’s suitable, and maintain the open-source twin to share the engineering experience.

[QUOTE END]

I wrote other blogs to analyze open-source factors within commercial software[2][3][4][5], and I have practiced them in several companies as well as earned merits in open-source projects.

When you think about it, there are many developers working for their employers, and using open-source software in their $DAYJOB is a good motivation to contribute more (especially for distributed systems; individuals can seldomly need one). I know there is open-source developers who develop software that has nothing to do with their $DAYJOB. I'm maintaining projects that has nothing to do with my $DAYJOB also (check Apache Curator, the Java binding of Apache OpenDAL, and more).

[1] https://www.tisonkun.org/2025/01/15/open-source-twin/

(Need a translator) [2] https://www.tisonkun.org/2022/10/04/bait-and-switch-fauxpen-...

[3] https://www.tisonkun.org/2023/08/12/bsl/

[4] https://www.tisonkun.org/2022/12/17/enterprise-choose-a-soft...

[5] https://www.tisonkun.org/2023/02/15/business-source-license/

tison··on Build a Database in Four Months with Rust and 647 Open-Source Dependencies
I've updated the Gist with a full Cargo.lock file that can be audited - https://gist.github.com/tisonkun/06550d2dcd9cf6551887ee6305e...

Running cargo audit -n --json | jq -r '.vulnerabilities.list[] | (.advisory.id + " - " + .package.name)' gives:

RUSTSEC-2023-0071 - rsa

which is transitively introduced by sqlx-mysql while we don't use the MySQL driver in production.

tison··on Build a Database in Four Months with Rust and 647 Open-Source Dependencies
Datadog always builds their own event store: https://www.datadoghq.com/blog/engineering/introducing-husky...

It may not be named "database" but actually take the place of a database.

Observability vendors will try to store logs with ElasticSearch and later find it over expensive and has weak support for archiving cold data. Data Warehouse solution requires a complex ETL pipeline and can be awkward when handling log data (semi-structured data).

That said, if you're building an observability solution for a single company, I'd totally agree to start with single node PG with backup, and only consider other solution when data and query workload grow.

tison··on Build a Database in Four Months with Rust and 647 Open-Source Dependencies
In the linked article below, we talked about "If RDS has already been used, why is another database needed?" and "Why RDS?"

Briefly, you need to manage metadata for the database. You can write your own raft based solution or leverage existing software like etcd or zookeeper that may not "a relational database". Now you need to deploy them with EBS and reimplement data replication + multi AZ fault tolerance, and it's likely still worse performance than RDS because first-class RDS can typically use internal storage API and advanced hardware. Such a scenario is not software driven.

https://flex-ninja.medium.com/from-shared-nothing-to-shared-...

tison··on Show HN: SPath is a Rust lib for query JSONPath over any semi-structured data
Here are several points I have in mind:

1. JSONPath/SPath supports multiple selector, e.g., $["a", "b"] or $[1, 3:10:2, 101]. This may be a bit more tidy than dot or subscript.

2. JSONPath/SPath supports descendant query: a descendant segment produces zero or more descendants of an input value. For example, $..[0] selects all the first element of arbitrary successors that is an array.

3. JSONPath defines filter selectors. SPath doesn't support it now, while it's on the Roadmap. This can be more powerful as a query language.

4. The original reason I wrote such a library is, however, to use JSONPath/SPath syntax beyond a JSON value. That is, I've written a database for processing semi-structured data [1], and my clients told me that's like to extract inner value with JSONPath syntax. All the existing JSONPath libraries, whether mature or not, are, of course, assuming they are handling JSON values. But in ScopeDB, we define our own variant value.

[1] https://www.scopedb.io/reference/datatypes-variant

tison··on Show HN: SPath is a Rust lib for query JSONPath over any semi-structured data
That's a good point. Tracked at https://github.com/cratesland/spath/issues/8.

I think it would be done before I call a 1.0 release.

tison··on ElasticSearch and many other repos are gone
Found https://status.elastic.co/incidents/9mmlp98klxm1

"Some public repositories are temporarily unavailable"

tison··on Show HN: Cronexpr, a Rust library to parse and iter crontab expression
I knew it. Please see also the comment at https://news.ycombinator.com/item?id=41666886

... and this PR [1][2]

[1] https://github.com/tisonkun/cronexpr/pull/12

[2] https://docs.rs/cronexpr/latest/cronexpr/struct.Crontab.html...

tison··on Show HN: Cronexpr, a Rust library to parse and iter crontab expression
Croner has a table for the syntax supported by it and saffron [1], you can compare it with the one cronexpr supported.

[1] https://github.com/hexagon/croner-rust?tab=readme-ov-file#wh...

As supported in [2],

> There are several good candidates like croner and saffron, but they are not suitable for my use case. Both of them do not support defining timezone in the expression which is essential to my use case. Although croner support specific timezone later when matching, the user experience is quite different. Also, the syntax that croner or saffron supports is subtly different from my demand. > > Other libraries are unmaintained or immature to use. > > Last, most candidates using chrono to processing datetime, while I’d prefer to extend the jiff ecosystem.

[2] https://docs.rs/cronexpr/latest/cronexpr/#why-do-you-create-...

tison··on Show HN: Cronexpr, a Rust library to parse and iter crontab expression
This can be optionally supported. To be clear, does it mean almost "RANDOM at construction" and the parser will assign a value by a hash/random function and then the crontab struct has an immutable value of that field?
tison··on Show HN: Cronexpr, a Rust library to parse and iter crontab expression
Yeah. Actually, this is possible to extend the interface with options to accept

1. Optional Timezone.

2. Second-level precision (perhaps feature flags are more suitable here)

It just falls out my first requirements so I don't support it. Being too generic is a common source of failure in my experience.

As a library developer I have my opinion on how things should be done and provide the default fits that mind :D

tison··on Show HN: Cronexpr, a Rust library to parse and iter crontab expression
If you mean timezone, I wrote an FAQ [1]: "Why does the crate require the timezone to be specified in the crontab expression?"

[1] https://docs.rs/cronexpr/latest/cronexpr/#why-does-the-crate...

tison··on Show HN: Cronexpr, a Rust library to parse and iter crontab expression
Could you elaborate a bit on the issue? I'm not sure you are commenting on cronexpr or other libraries.

In cronexpr, there is no requirement for a timestamp until you'd like to find the next scheduled time, and thus you need to provide a related point.

To decouple with certain datetime lib, I made a `MakeTimestamp` struct which provides multiple constructors. Later, I found it somehow like a function overload :D

tison··on Show HN: Cronexpr, a Rust library to parse and iter crontab expression
You're welcome! As described in the "Why do you create this crate?" section, I actually started this domain only a few weeks ago. So I experienced how those tribal rules can be confusing and hard to search over the Internet the understand their semantic.

And when I sorted out the parse structure [1] and finished the extensions [2][3], I believe I should write it down for others (and future me) :D

[1] https://github.com/tisonkun/cronexpr/pull/4

[2] https://github.com/tisonkun/cronexpr/pull/5

[3] https://github.com/tisonkun/cronexpr/pull/6

tison··on Show HN: Personal Blog based on Astro.build and styled by TailwindCSS
https://github.com/tisonkun/dacapo/pull/5

Added at /rss.xml, fully https://tisonkun.io/rss.xml.

Please check if it's desired and raise any issue if it can be improved further.

tison··on Show HN: Personal Blog based on Astro.build and styled by TailwindCSS
Thanks for your feedback!

I'd like to keep the page clean to focus on the content.

Yes, I noticed that the space is too tight on mobile. I may not be good at style, but I will try to improve it.

tison··on Show HN: Personal Blog based on Astro.build and styled by TailwindCSS
Yes. I opened an issue here https://github.com/tisonkun/dacapo/issues/3.

I found other templates implement this, so it should be easy to borrow from their code.

I have a large backlog of posts written in my native language that can be posted here. Stay tuned :D

tison··on The race to replace Redis
And here is an interesting conversation when Binbin came to the Kvrocks community: https://github.com/apache/kvrocks/pull/1581#issuecomment-163...

* Me: @enjoy-binbin Out of curiosity, do you have a fuzzer to test out Kvrocks? Your recent great fixes seem like a combo rather than random findings :D

* Binbin: They were actually random findings.I may be sensitive to this, doing code review and found them (also based on my familiarity with redis)

tison··on The race to replace Redis
For the first time, I know our (Apache Kvrocks, an alternative to Redis on Flash) committer Binbin Wang committed nearly 25% of the commits to the newer Redis version.

You can find his contributor for both at:

* https://github.com/apache/kvrocks/graphs/contributors

* https://github.com/redis/redis/graphs/contributors

Page 1 of 2Next →