HNHacker News
TopNewBestAskShowJobs

Sheldon_fun

153 karma · joined July 10, 2023

submissionscomments
Sheldon_fun··on Data Engineering 2025: Unified Batch‑Streaming, Iceberg Rise and Data Contracts
I’m curious: Have you started using data contracts in your pipelines? Is the unified batch/stream model worth the added complexity?
Sheldon_fun··on The Equality Delete Problem in Apache Iceberg
From my private conversations with several Iceberg PMC members, it’s clear that full equality delete support across major query engines will be slow — not due to lack of will, but due to complexity.
Sheldon_fun··on Why Postgres CDC to Iceberg isn't a solved problem: lessons from production
Although the last time I touched Debezium in 2020, it was too immature to adopt, I thought surely its problems had been solved by now. Apparently not. I really appreciate this in-depth list of real-world problems encountered by clients trying to pipe CDC-captured changes.
Sheldon_fun··on We Built an In-Memory-Class Architecture on Top of S3 – and Made It Work
No fluff, no hand-waving—just every mistake, lesson, and trick we learned along the way.
Sheldon_fun··on [dead]
Understanding Twitter's Data Infrastructure Challenges and Open-Source Solutions
Sheldon_fun··on The Rust Closure Pitfall That Silently Corrupted Our Metrics: A Debugging Saga
Key takeaways: Rust 2021's closure minimal capture breaks RAII patterns when only Copy-type fields are used in structs with Drop impl. Even with impl Drop, closures may capture Copy fields instead of the whole struct — a surprising edge case. Fix requires explicit ownership transfer via let stats = self.stats to override closure's partial capture.
Sheldon_fun··on We Ditched Etcd for PostgreSQL in Our Cloud-Native Streaming Database
etcd is primarily designed for bare-metal deployments, and its performance often suffered in cloud environments due to the relatively slower disk performance compared to on-premise setups.
Sheldon_fun··on Kafka at a Crossroads: Cloud Streaming, Batch Convergence and 10x Cost Cuts
Kafka has dominated data streaming for years, but cloud-native platforms (Snowflake, Redshift) now ingest data directly, batch-streaming convergence (Iceberg, lakehouses) is reshaping architectures, and cost-efficient alternatives (WarpStream, Redpanda) are cutting costs by 10x. This article explores whether Kafka can adapt—or if the streaming ecosystem is moving beyond it.
Sheldon_fun··on Streaming Databases and LLMs = Proactive AI Agents That React in Real Time
The core idea: an LLM subscribes to event-driven triggers defined in Streaming SQL (e.g., stock price surges, security alerts, IoT signals). When a trigger fires, the database pushes relevant context to the LLM, enabling instant decision-making without constant polling.
Sheldon_fun··on Time-Series vs. Streaming Databases: Key Differences and Use Cases
An in-depth comparison between time-series databases (TSDBs) and streaming databases, and when to use each.
Sheldon_fun··on Reflections from a Database PhD: Why Data Engineers Should Embrace AI
Founder of a data streaming startup (a PhD in database systems, ex-AWS Redshift/IBM researcher) explains why AI will fundamentally reshape data engineering within 24 months. Key insights: How text-to-SQL is becoming a commodity (with Snowflake hitting 90% accuracy) Why vector databases are a dead-end business model When AI replaces feature engineering (and what that means for Spark/RisingWave) The surprising way AI acts as "lossy compression" for storage Why database vendors must now answer: "Do we even need databases anymore?"
Sheldon_fun··on Roaring-rs: High-performance compressed bitmaps in Rust
roaring-rs is a Rust implementation of the Roaring bitmap data structure, originally introduced as a Java library for efficiently representing large sets of integers. This crate offers memory-efficient, high-performance compressed bitsets for Rust, and is compatible with the Roaring format used in other languages. Benchmarks and real-world datasets are included, and contributions are welcome!