HNHacker News
TopNewBestAskShowJobs

iamlintaoz

35 karma · joined September 19, 2024

submissionscomments
iamlintaoz··on We're no longer attracting top talent: the brain drain killing American science
This is not true at all. China’s education system is nationally standardized. Although economic development is uneven with far greater investment concentrated in major cities than inland regions, the structure of education itself is consistent across the country. Schools follow the same national curriculum and use the same core teaching materials.

Income disparities may have some impact on teacher quality, but the difference is often less significant than people assume. Broad access to education tends to matter more than whether a particular middle-school teacher is exceptional. In fact, students in some inland provinces frequently achieve very high scores on the national college entrance examination, driven in part by strong incentives to gain admission to top universities and pursue opportunities in more economically developed regions.

Among younger generations, illiteracy is virtually nonexistent. With nine years of compulsory education mandated nationwide, basic literacy rates are effectively at 100 percent.

iamlintaoz··on We're no longer attracting top talent: the brain drain killing American science
It’s true that learning Chinese as an adult—especially if you come from an English or other European language background—can be extremely challenging. I have several colleagues who have lived in Beijing for more than a decade, are married to Chinese spouses, and still can barely speak the language, it becomes even more challenging for reading.

This creates real difficulties in daily life. Today, almost all routine activities—online shopping, digital payments, banking, ride-hailing—are conducted through smartphone apps. If you can’t read Chinese, even basic tasks become complicated. In recent years, the number of foreigners living in China has declined compared to a decade ago. While political and economic factors clearly play a role, I suspect that the language barrier has also become a more significant obstacle.

Many Chinese people, especially younger generations, can speak some basic English, since it is a mandatory subject in school. As a result, interpersonal communication is usually manageable, and traveling in China is relatively easy. However, living there long-term is a very different experience from visiting as a tourist.

iamlintaoz··on EloqKV: Achieving Predictable P99.99 Latency on NVMe with Redis API
The disk storage EloqKV uses (EloqStore [1]) is optimized for batch updates because the upper Data Substrate layer manages buffering and the Write-Ahead Log (WAL), absorbing writes and guaranteeing durability. When durability is not required, the WAL can be optionally disabled.

[1] github.com/eloqdata/eloqstore

Disclaimer: I am the CEO of EloqData

iamlintaoz··on Scaling PostgreSQL to power 800M ChatGPT users
If you need so many tricks to support the infra, it will eventually come back to bite you. I am pretty sure that Google in year 2000 could have supported their workloads with existing technologies (Yahoo could, and it was a much larger company). But they did GFS and Bigtable, and the rest is history. Other companies struggled to catch up due to inferior infrastructure. A visionary company needs to be prepared and should not be hindered by infrastructure. Can you scale the single primary system another 10x or more? Because their CEO said that they will scale their revenue by that much within just a couple of years.
iamlintaoz··on Show HN: EloqDoc: MongoDB-compatible doc DB with object storage as first citizen
These are great questions. I appreciate you carefully reading through the documents. For the first question, we have detailed benchmark on EloqKV, with the same architecture (but with Redis API) in our blog, and we will soon publish more about the performance characteristics of EloqDoc. Overall, we achieve about the same performance as using local NVME SSD, even when we use S3 as the primary storage, and the performance often exceed the original database implementation (in the case of EloqDoc, original MongoDB).

As for the durability part, our key innovations is to split state into 3 parts: in memory, in WAL, and in data storage. We use a small EBS volume for WAL, and storage is in S3. So, durability is guaranteed by [Storage AND (WAL OR Mem)). Unless Storage (S3) fails, or Both WAL (i.e. EBS lost) AND Mem fail (i.e. node crash), persistence is guaranteed. You can see the explanation in [1]

[1] https://www.eloqdata.com/blog/2025/07/16/data-substrate-bene...

iamlintaoz··on Show HN: EloqDoc: MongoDB-compatible doc DB with object storage as first citizen
That's exactly the reason. S3 is better in almost all aspects compared with EBS, except the performance part, and I am glad that our Data Substrate technology solved this issue gracefully [1].

As for the compatibility, we are leveraging some of the code from 4.03 version (the last AGPL version), and we have a very good compatibility (we will show some results in later blog posts). As I mentioned in another reply post, the Mongo APIs are reasonably stable over the last few years, only seeing very minor changes. Most of the later versions improved upon performance and transaction supports, which we support natively with our underlying data substrate technologies. Still, if you have any specific API that you feel is needed, we'd be happy to implement and we welcome community contributions.

Multi-master/multi-writer means it is a fully distributed database. Of course you can run it in single node configurations and get all the single node benefits, but if deployed in a cluster, you do not need to worry about which node to write to, or how data are sharded. If you writes potentially can cause conflicts (i.e. write to the same data at the same time on different nodes), the concurrency-control will handle that for you. In fact, you will encounter the same issue even in a single node configuration, since a single node is still multi-threaded.

[1] https://www.eloqdata.com/blog/2025/07/16/data-substrate-bene...

iamlintaoz··on Show HN: EloqDoc: MongoDB-compatible doc DB with object storage as first citizen
For full transparency, I am the CEO of EloqData and I am happy to answer any questions you have. EloqDoc has been open-sourced for awhile, and we are glad to announce that the current code is stable for you to test on your workload.
iamlintaoz··on Show HN: EloqDoc: MongoDB-compatible doc DB with object storage as first citizen
Thank you so much for reporting this! My local testing on Chrome privacy mode was fine, but we will definitely prioritize checking and fixing the issue in Firefox's Strict Privacy Mode.
iamlintaoz··on Show HN: EloqDoc: MongoDB-compatible doc DB with object storage as first citizen
That is right. We leverage some of the AGPL MongoDB code for the parser, as indicated by the License. Our own code can be licensed differently, see a previous discussion on Hacker News [1]. The Mongo API are reasonably stable over the last few years, only seeing very minor changes. Most of the later versions improved upon performance and transactions, which we support natively with our underlying technologies. Still, if you have any specific API that you feel is needed, we'd be happy to implement and we welcome community contributions.

[1] https://news.ycombinator.com/item?id=44937978

iamlintaoz··on Show HN: EloqDoc: MongoDB-compatible doc DB with object storage as first citizen
That’s correct. FerretDB and DocumentDB both pursue PostgreSQL-based document models.

EloqDoc follows a fundamentally different architectural path. Instead of executing in Postgres, we leverage the existing MongoDB parser and executor to ensure maximum compatibility. Our main contribution is to replace the single-writer WiredTiger storage engine entirely with Data Substrate.

The Data Substrate is an abstract layer built specifically to handle distributed database fundamentals: scalable buffer pooling, concurrency control, durability, elasticity, and fault tolerance. You can read more: https://www.eloqdata.com/blog/2025/07/14/technology

This architectural choice is what enables EloqDoc to deliver features that Postgres-based solutions cannot easily match, including: full compatibility, native Multi-Writer capability, Object Storage First and extremely low-latency distributed transactions.

iamlintaoz··on Show HN: EloqDoc: MongoDB-compatible doc DB with object storage as first citizen
Yes, this is exactly the typical use case where EloqDoc shines.

Our auto-tiering feature allows you to manage a massive database (e.g., 5TB+) primarily on cheap, durable object storage (like S3), which is already built for cross-AZ replication and is up to 3x cheaper than traditional cloud block storage (EBS).

We use local NVMe SSDs (200K+ IOPS) purely as a high-performance cache layer to accelerate read access. Hot and cold data are automatically swapped across memory, SSD, and object storage tiers based on access frequency, ensuring high performance (up to 100x faster than archive solutions)

iamlintaoz··on EloqKV, a distributed database with Redis compatible API (GPLv2 and AGPLv3)
Thank you. Indeed we do already have several large multi-national companies using EloqKV in their production environments. Please contact us if you have any further questions. Moreover, we would be really interested to hear more details about your usage scenario. Metadata store for JuiceFS is a very interesting use case for us.
iamlintaoz··on EloqKV, a distributed database with Redis compatible API (GPLv2 and AGPLv3)
Great question. Currently, the database landscape is very fragmented. We are faced with a multitude of database choices (different ACID guarantees, data modalities, scalability, and so on). Data pipeline become very complicated, and we believe there must be a better solution. That’s why we developed a common architecture called DataSubstrate and built different APIs on top of it. EloqKV with a Redis API is just one of them; we also provide a MySQL-API RDBMS and a MongoDB-API JSON database (both open-sourced). Our goal is to create the next-generation database foundation to support the growing demand from new generation of applications. We believe future AI agent-driven applications will generate huge volumes of queries and data that will be difficult to handle with existing solutions.

EloqKV is only slightly slower than Dragonfly—about 10–20%—but for good reason. Dragonfly is a pure in-memory database with a highly optimized network layer and a very specialized design. EloqKV, on the other hand, is a full-featured database with all the checkboxes you can think of: fully consistent, durable, distributed transactions, fault-tolerant, tiered storage, and more. Despite this, we incur very little overhead compared with state-of-the-art, purpose-built databases when we have the same workload guarantee. Our thesis is that we may not need dedicated specialized solutions if we can achieve comparable performance (and cost) with general-purpose systems.

We will also add vector support very soon. Please stay tuned.

iamlintaoz··on EloqKV, a distributed database with Redis compatible API (GPLv2 and AGPLv3)
TiKV is a great project and we have a lot of respect for their work. EloqKV is based on a very different architecture [1], and we also have MySQL compatible [2] and MongoDB compatible [3] databases build on top of the same architecture. They all inherit the extreme performance, scalability, fault tolerance, and ACID properties due to the common underpinning.

[1] https://www.eloqdata.com/blog/2025/07/14/technology

[2] https://github.com/eloqdata/eloqsql

[3] https://github.com/eloqdata/eloqdoc

iamlintaoz··on EloqKV, a distributed database with Redis compatible API (GPLv2 and AGPLv3)
That’s right. You can simply choose GPL and ignore the AGPL part for EloqKV. The reason we use both is that, in other projects, we need to support both GPL and AGPL. EloqKV and all its dependencies are either developed by us or licensed under more permissive terms, so you can choose either license. However, EloqDoc is under AGPL (https://github.com/eloqdata/eloqdoc) and we cannot relicense it under GPL because it includes some AGPL-licensed code from MongoDB.
iamlintaoz··on EloqKV, a distributed database with Redis compatible API (GPLv2 and AGPLv3)
The reason is because we use a layer (our main innovation) called DataSubstrate to build modular databases that support distributed transactions [1]. Because we use some MariaDB code (GPLv2) and some MongoDB code (AGPLv3) in our other projects [2] [3], so we license our DataSubstrate code to be compatible with both, and therefore, we also license EloqKV under both licenses.

[1] https://www.eloqdata.com/blog/2025/07/14/technology

[2] https://github.com/eloqdata/eloqsql

[3] https://github.com/eloqdata/eloqdoc

iamlintaoz··on EloqKV, a distributed database with Redis compatible API (GPLv2 and AGPLv3)
Really appreciate it. I am the CEO of EloqData. We submitted the ShowHN about a year ago [1]. Since then we made a lot of progress, much of the work is based on the feedback from the great HN community, including:

1) Open Source (GPL and AGPL) (thanks PeterZaitsev).

2) Session based transaction in Redis API. (thanks fizx)

3) Better explanation of the architecture [2] (thanks apavlo).

4) Testing with Jepsen (internally, we will do it officially when we have the resource) (thanks jacobn, among others).

Again, thanks, really appreciate the community support. Please go to our website [3] or join our discord channel to provide more feedback.

[1] https://news.ycombinator.com/item?id=41590905

[2] https://www.eloqdata.com/blog/2025/07/14/technology

[3] https://eloqdata.com

iamlintaoz··on How to Migrate from OpenAI to Cerebrium for Cost-Predictable AI Inference
Why? Honestly, there are already tons of Model-as-a-Service (MaaS) platforms out there—big names like AWS Bedrock and Azure AI Foundry, plus a bunch of startups like Groq and fireflies.ai. I’m just not seeing what makes Cerebrium stand out from the crowd.
iamlintaoz··on Show HN: EloqKV – Scalable distributed ACID key-value database with Redis API
You're absolutely right—discussing the CAP theorem is essential for any distributed storage system. We are preparing a detailed blog post to cover the implications of CAP in our system. In brief, our system generally aligns with PC/EC in PACELC terms (https://en.wikipedia.org/wiki/PACELC_theorem), similar to other fully distributed databases like CockroachDB and TiDB. However, this configuration is flexible. For instance, in high-performance cache applications where durability is less critical, our system can be adjusted to favor PA/EL, akin to many NoSQL systems.

Regarding persistence, our Substrate component manages its own logging system while relies on external storage engines for checkpointing. This means that the system only depends on the performance and capacity of the external storage. Our database does not depend on the external storage engine for consistency. For example, DynamoDB offers virtually unlimited capacity and performance (in terms of QPS) in the cloud, and our Data Substrate is agnostic to whether the data is stored as a B-Tree or an LSM tree.

We are currently preparing the software for Jepsen testing prior to its General Availability (GA) release. More technical details will be shared as we move forward, so please stay tuned.

iamlintaoz··on Show HN: EloqKV – Scalable distributed ACID key-value database with Redis API
Thank you for the insightful reply. I agree, Redis's transaction API ("MULTI/EXEC/DISCARD/WATCH") is challenging compared to SQL's familiar "START/COMMIT/ROLLBACK," and it has key limitations:

1. No rollback on failure: If a command in EXEC fails, Redis can't roll back the transaction. EloqKV resolves this by starting real transactions on the TxServer node during EXEC. If any command fails, the changes are rolled back, maintaining atomicity. Additionally, EloqKV allows specifying isolation levels when using Redis transactions.

2. Lack of interactivity: This is where Lua shines. Users can embed business logic in Lua, functioning similarly to stored procedures in SQL.

3. No cross-shard transactions: Redis clusters can't redirect requests across shards, forcing clients to manage topology awareness. EloqKV addresses this with a fully distributed transactional design, eliminating these cross-shard complexities. For more on this, see our blog.

https://www.eloqdata.com/blog/2024/08/22/benchmark-cluster#w...

As for your point on Lua's risks in multi-tenant environments, I completely agree. Lua lacks robust ACLs or resource limitations, and managing SHA keys is cumbersome. We're considering enhancements to address these issues in the future. Finally, we are indeed working on a "real" transaction API for Redis. Like SQL, users will be able to begin transactions, read keys, apply transformations, and generate new keys, with the ability to commit or abort. EloqKV will also support configurable isolation levels and transaction protocols (OCC/Locking).

Do you have any additional API preferences beyond "START/COMMIT/ROLLBACK" for Redis transactions?

iamlintaoz··on Show HN: EloqKV – Scalable distributed ACID key-value database with Redis API
Thank you, Andy! This is Jeff, CEO and Chief Architect of EloqData. It's a great honor for us to have THE Andy Pavlo join the discussion on our first HN submission.

If I remember correctly, NuoDB uses a shared cache with a cache coherence protocol, whereas EloqKV uses a shared nothing (partitioned) cache. The former is a local read but needs to broadcast each write to all nodes. The latter has no broadcast for writes but may be a remote read. The tradeoff is evident and we are actively exploring opportunities to strike a balance, e.g., for frequently-read, rarely-write data items, use the shared cache mode.

We appreciate you pointing us to the CIDR paper. I had the pleasure of working with Phil for some time and fondly remember many discussions with Phil on various topics many years ago. To address your question, yes, we've been trying to solve the research challenges presented in the CIDR paper. The devil is in the details. We've developed numerous new algorithms and invested significant engineering effort into the design and implementation of our products. The benefits are as follows:

- Optimality: We believe we have an overall design that optimizes synchronous disk writes and network round-trips. For instance, when the design is reduced to a single node, its performance matches or exceeds that of single-node servers. As you might expect, a lot of innovation has gone into making distributed transactions as efficient as non-distributed ones, comparable to those in MySQL or PostgreSQL.

- Modularity: Our architecture allows us to easily replace the Parser/Compute layer and Storage/Persistence layer with the best existing solutions. This means we can create new databases by leveraging existing parsers and compute engines from current database implementations to achieve API-compatibility, as well as leveraging existing high-performance KV stores for the persistence layer. This allows us to avoid reinventing the wheels and to take advantage of decades of innovations in the database community.

- Scalability: The entire system operates without a single synchronization point—not even a global sequencer. We drew many inspirations from the Hekaton and your TicToc paper. All four types of resources (CPU, Memory, Storage, Logging) can be scaled independently, as we mentioned earlier. More importantly, they can scale dynamically to accommodate workload changes without service disruptions.

We look forward to sharing more technical details as we move out of stealth mode. I hope to continue this conversation with you in person in the near future.