HNHacker News
TopNewBestAskShowJobs

kwillets

655 karma · joined August 31, 2016

submissionscomments
kwillets··on A database kernel engineer's take on Apache Iceberg
I've been reading up on lakehouse stuff, and everything seems to agree with the idea that data warehouse functionality deteriorates as they try to accommodate raw data access.

The databricks people suggest that a better "cache format" is needed, but I don't see how that's different from ETLing the data into a regular warehouse.

kwillets··on The militarization of Silicon Valley
I don't know; my neighbors never talked about it.

But one guy came back from retirement after 9/11 and moved into a trailer in the Lockheed parking lot.

kwillets··on A database kernel engineer's take on Apache Iceberg
I see Iceberg as a nice complement to a real data warehouse, as a way to manage raw data files and ETL's that load them. Getting file updates as transactions a few times a day seems a lot cleaner.

But I constantly underestimate the tendency to take a minor capability and exaggerate it into a replacement for 50 years of heavily studied technology.

Even if people somehow get this to work like a real database, they will quickly become frustrated with the same features that the data lake avoids: access controls, locks, consistency, centralization, a schema, you name it.

kwillets··on Efficiently Generating a Number in a Range (2018)
Extended-width multiplication works, but the cost of extra random bits is often a lot higher than the range arithmetic.

Somewhere in my github there's an indefinite-width multiply that only adds bits while there's a risk of carry into the 1's digit; the check for that is quite cheap.

kwillets··on Ask HN: What are you working on? (July 2025)
Some folks I work with have a terabyte of short string records that they regularly scan with regexes to develop classification criteria; it's a prime application for a substring index that can accelerate their queries, but the scale is daunting.

I came up with a suffix-sorting index for this domain that's interestingly simple. Most algos for this use a generalized suffix tree that's built by concatenating all the strings into one giant string and feeding it into a conventional suffix sort, but that has some big constants on the indexing throughput, due I think to the overhead of handling one giant string instead of a bunch of small independent records.

In the latter case, by making the structure slightly simpler and search slightly harder, I can get indexing throughput in the GBps, at least for the sorting part.

The output of that in its simplest form is a 4n or 8n-sized set of int's, but it can be fed further into a compressed rank/select data structure for various space/indexing time/retrieval time tradeoffs, and I don't think those are slow (eg Roaring Bitmaps)

I'll post this on show HN if anybody's interested; I'm still writing up the details, as I've barely gotten the POC code working.

kwillets··on Mike Lynch estate and business partner owe HP Enterprise £700M, court rules
>The judge expressed his "sorrow at this devastating turn of events, and my sympathy and deepest condolences", adding that he "admired" Mr Lynch despite ruling against him.

The UK is looking increasingly like a third world country.

kwillets··on Analyzing database trends through 1.8M Hacker News headlines
Snowflake seems to have peaked; 2023 was hellish dealing with roomfuls of inexperienced devs and even architects convinced it was the fastest cheapest thing ever.
kwillets··on Why Koreans ask what year you were born
Kakao ended the English names a year or so ago due to inevitable confusion, but now people are supposed to add the "nim" (dear) suffix to each other's Korean names instead, which sounds creepy.

The funny part of the Kakao CEO asking to be called Brian is that there was a K-drama (Search WWW) with a fictional tech company, but they also made the CEO name Brian. I suspect if this idea had gone on every CEO would be called Brian.

kwillets··on Snowflake to buy Crunchy Data for $250M
Snowflake is becoming the Juicero of data.
kwillets··on AtomVM, the Erlang virtual machine for IoT devices
There's a why page in their docs, but basically it's cheaper than a multitasking OS for doing the same thing. The valueprop of Erlang is lightweight threads and message-passing.

My knowledge of microcontrollers is dated, but the frameworks for Arduino etc. seemed limited on their ability to do event or message-based programming; most example apps at least were a polling loop. The classic architecture of setting up interrupt/event handlers and going to sleep was not there.

kwillets··on Microsoft-backed UK tech unicorn Builder.ai collapses into insolvency
The UK has cemented its place in tech as the land without Sarbanes-Oxley.
kwillets··on A lost decade chasing distributed architectures for data analytics?
Stonebraker is one of the few whose criticism is listened to. He recently updated his "What goes around comes around" paper; it's worth a read:

https://db.cs.cmu.edu/papers/2024/whatgoesaround-sigmodrec20...

kwillets··on A lost decade chasing distributed architectures for data analytics?
Cloud and SaaS were good for a while because they took away the old sales-CTO pipeline that often saw a whole org suffering from one person's signature. But they also took away the benefits of a more formal evaluation process, and nowadays nobody knows how to do one.
kwillets··on Databricks acquires Neon
It's been a commodity for decades now. Metrics like price-performance have a long history, but the SnowBricks products fail at them quite dramatically. The difference is hard-sell vs. soft or no-sell.
kwillets··on Databricks in talks to acquire startup Neon for about $1B
It's only serverless in the way it commits transactions to cloud storage, making the server instance ephemeral; otherwise it has a server process with compute and in-memory buffer pool almost identical to pg, with the same overheads.
kwillets··on Bloom Filters
That's an interesting reference. I came up with what may be a simpler case of that -- a structure that estimates the last time a key was seen. It's an upper bound, as collisions make the time more recent, but never the other way. Instead of a semi-order it's a simple ordering, but it may be similar to the compact approximator.

Rate limiters apparently use count-min over a fixed time interval, which is bursty, but I took the same hash structure and came up with "timestamp-min" that allows even pacing (also without "fill up"). The one-sided error is also useful for checking if a cache entry is too old.

It doesn't fix everything (eg DDOS), but it can prevent any single client from over-requesting, or stealing another client's requests (collisions can be made as hard as necessary).

https://github.com/KWillets/RecencySketch

kwillets··on ClickHouse gets lazier and faster: Introducing lazy materialization
Nothing in C-store seems to have sunk in. In clickhouse's case I can forgive them since it was an open source, bootstrap type of project, and their cash infusion seems to be going into basic re-engineering, but in general slowly re-implementing Vertica seems like a flawed business model.
kwillets··on ClickHouse gets lazier and faster: Introducing lazy materialization
Late Materialization, 19 years later.

https://dspace.mit.edu/bitstream/handle/1721.1/34929/MIT-CSA...

kwillets··on Databases in 2024: A Year in Review
I don't know much about the financial side of the company, but it seems like a client-led effort by telco's etc. against the dreck that tech VC's keep pushing on them. That can't translate into decent salaries unfortunately.

Silicon Valley doesn't have a good record in the DB/DWH space; producing a fully-featured DBMS doesn't seem to fit the VC model.

kwillets··on Databases in 2024: A Year in Review
Zero is certainly on my scale. Bonus points if you build a server and keep it under your desk.
kwillets··on Databases in 2024: A Year in Review
I spent the past year puzzling over the DB market as well, but I don't feel like I'm much closer to understanding it.

It appears that a lot of attention is now directed at the folks doing 100 MB queries, and the high end has moved past everybody's radar. My idea of an exciting product is Ocient, who have skipped over Cloud and gone for hyperscale on-prem hardware. Yellowbrick is also a contender here.

I have a lot of experience with Vertica, and they seem to have gotten stuck in this niche as well, with sales tilted towards big accounts, but less traction in smaller shops, and a difficult road to get a SaaS or similar easy-start offering.

There's a crossover point where self-managed is cheaper than cloud, but nobody seems to have any idea where it is. Snowflake will gladly tell you that your sub-$1M Vertica cluster should be replaced by $10M of sluggish SaaS, and that you are saving money by doing so. These decisions seem more in the realm of psychology or political science.

DHH's cloud exit was a refreshing take on the expense issue, even if it wasn't strictly in the database space -- the cost per VCPU and so forth that he documented is a good start for estimating savings, and he debunked a lot of the "hidden costs" that cloud maximalists claim.

In the business/financial space the biggest news to me was the correction in Snowflake's stock price, which seemed to indicate that investors were finally noticing metrics like price-performance, but they added a little more AI and went back into irrationality.

I'm heavily in favor of DuckDB, Hudi, Iceberg, S3 tables, and the like. Mixing high-end and low-end tools seems like the best strategy (although settling on one high-end DWH has also worked IME), and the low end is getting better and cheaper, squeezing out the mid-range SaaS vendors.

In research I found Goetz Graefe's work in offset-value coding exciting -- he's wired it into query operators in a way that saves a lot of CPU on sorting and joins/aggregation. This is a technique that I've applied favorably in string sorting, and it was discovered in the DB community decades ago but largely forgotten. (This work precedes 2024, but I'm a slow study.)

kwillets··on How bloom filters made SQLite 10x faster
The animation, pseudocode, and join order discussion all imply that the cartesian product is being generated.
kwillets··on How bloom filters made SQLite 10x faster
The description of nested loop join is confusing; it's mainly a single pass through the outer table with one B-tree probe per row to each inner table.

The linked paper is clearer:

"However, the inner loops in the join are typically accelerated with existing primary key indexes or temporary indexes built on the fly."

"Note that SQLite probes the part table index for every tuple in the lineorder table."

The Bloom filter does not reduce the cardinality of the join, it simply replaces each B-tree probe and filter with a Bloom probe on a pre-filtered key set.

This technique is well-known; the paper cites several examples, and Bloom filter pushdown is common to many commercial systems.

kwillets··on Why we use our own hardware
SSD's are also a bit of an achilles heel for AWS -- they have their own Nitro firmware for wear levelling and key rotations, due to the hazards of multitenant. It's possible for one EC2 tenant to use up all the write cycles and then pass it to another, and encryption with key rotation is required to keep data from leaking across tenant changes. It's also slower.

We had one outage where key rotation had been enabled on reboot, so data partitions were lost after what should have been a routine crash. Overall, for data warehousing, our failure rate on on-prem (DC-hosted) hardware was lower IME.

kwillets··on Bit-twiddling optimizations in Zed's Rope
That's my guess as well.

Bitstring rank/select is a well-known problem, and the BMI and non-BMI (Hacker's Delight) versions are available as a reference.

kwillets··on Optimizers: The Low-Key MVP
I've never run across it, but I would describe the job of an analytics engineer as doing this over and over for analysts -- it's probably semantically clearer to do the join first and then aggregate, so I end up pushing it down for them.

Thanks for the reference; this question has been on my mind for some time.

kwillets··on Optimizers: The Low-Key MVP
To follow up, it's about 3x faster:

  SELECT 
      pickup.zone AS pickup_zone,
      dropoff.zone AS dropoff_zone,
      cnt AS num_trips
  FROM
      (select pickup_location_id, dropoff_location_id, count(*) as cnt from taxi_data_2019 group  by 1,2) data
  INNER JOIN
      (SELECT * FROM zone_lookups WHERE Borough = 'Manhattan') pickup
      ON pickup.LocationID = data.pickup_location_id
  INNER JOIN
      (SELECT * FROM zone_lookups WHERE Borough = 'Manhattan') dropoff
      ON dropoff.LocationID = data.dropoff_location_id
  
  ORDER BY num_trips desc
  LIMIT 5;
  ┌───────────────────────┬───────────────────────┬───────────┐
  │      pickup_zone      │     dropoff_zone      │ num_trips │
  │        varchar        │        varchar        │   int64   │
  ├───────────────────────┼───────────────────────┼───────────┤
  │ Upper East Side South │ Upper East Side North │    536621 │
  │ Upper East Side North │ Upper East Side South │    455954 │
  │ Upper East Side North │ Upper East Side North │    451805 │
  │ Upper East Side South │ Upper East Side South │    435054 │
  │ Upper West Side South │ Upper West Side North │    236737 │
  └───────────────────────┴───────────────────────┴───────────┘
  Run Time (s): real 0.304 user 1.791931 sys 0.132745
(unedited query is similar to theirs, about .9s)

And, doing the same to the optimized query drops it to about .265s.

kwillets··on Optimizers: The Low-Key MVP
One possible hand optimization is to push the aggregation below the joins, which makes the latter a few hundred rows.
kwillets··on How to Correctly Sum Up Numbers
It's a sad example of my C++ knowledge that I didn't know about the checked add intrinsics, but I do know that unsigned (x + y) < x compiles into checking the carry flag on most compilers.
kwillets··on ByteDance sacks intern for sabotaging AI project
Conspiracy theories are common in repressive regimes.
← PreviousPage 2 of 13Next →