HNHacker News
TopNewBestAskShowJobs

kwillets

655 karma · joined August 31, 2016

submissionscomments
kwillets··on Improving Parquet Dedupe on Hugging Face Hub
Strike the sharding idea as I misunderstood the CDC method.
kwillets··on Show HN: Vortex – a high-performance columnar file format
This paper compares the benefits of lightweight compression and other techniques:

https://blog.acolyer.org/2018/09/26/the-design-and-implement...

kwillets··on Show HN: Vortex – a high-performance columnar file format
Does this fragment columns into rowgroups like Parquet, or is it more of a pure columnstore? IME a data warehouse works much better if each column isn't split into thousands of fragments.
kwillets··on Improving Parquet Dedupe on Hugging Face Hub
One additional thought regarding query performance is that content-defined row groups allow localized joins and aggregations which are much faster than the globally-shuffled kind.

If the sharding key matches (or is a subset of) a join or group-by key, then identical values are local to a single shard, which can be processed independently.

This type of thing is typically done at large granularity (eg one shard per MPP compute node), but there are also benefits down to the core or thread level.

Another tip is that if no shard key is defined, hash the whole row as a default.

kwillets··on Improving Parquet Dedupe on Hugging Face Hub
One more: do you prefer the CDC technique over using the rowgroups as chunks (ie using knowledge of the file structure)? Is it worth it to build a parquet-specific diff?
kwillets··on Improving Parquet Dedupe on Hugging Face Hub
How does this compare to rsync/rdiff?
kwillets··on Improving Parquet Dedupe on Hugging Face Hub
I'm surprised that Parquet didn't maintain the Arrow practice of using mmap-able relative offsets for everything. Although these could be called relative to the beginning of the file.
kwillets··on Building data infrastructure that will last
OMG I'm that guy -- 3 straight Vertica roles with $B annual revenue.

I did learn a lot from watching MDS people try to beat it (in the end I'm also looking for what should come next), but mostly it was confirming the article and RDD. What they didn't know about data warehousing they also didn't know about performance or price-performance or selecting tools or managing projects or vendors, so costs exploded.

These folks were hilarious because they kept insisting that Vertica is not "modern", while it beat the pants off them with basic columnstore stuff.

kwillets··on Ask HN: Resources about math behind A/B testing
Microsoft has a seed finder specifically aimed at avoiding a priori bias in experiment groups, but IMO the main effect is pushing whales (which are possibly bots) into different groups until the bias evens out.

I find it hard to imagine obtaining much bias from a random hash seed in a large group of small-scale users, but I haven't looked at the problem closely.

kwillets··on I Wasn't the First Person to Find the NJ Coca-Cola Cocaine Factory (2016)
Once I had tried coca tea in Peru, it was clear what the taste of Coca-Cola comes from.

What became funny was the series of "tasters" who claimed they had figured out the formula with other ingredients like cinnamon, etc.

kwillets··on Let's consign CAP to the cabinet of curiosities
This seems like what I've noticed on MPP systems (a little before cloud): data replicas give a lot more availability than the number of partition events would suggest.

I likely need to read the paper linked, but it's common to have an MPP database lose a node but maintain data availability. CAP applies at various levels, but the notion of availability differs:

1. all nodes available 2. all data available

Redundancy can make #2 a lot more common than #1.

kwillets··on Is an All-in-One Database the Future?
We came close, as we had architects who knew little besides Snowflake.

While there are all kinds of issues it can't handle for CRUD, such as locking and small updates, it's also expensive to do small reads due to the reliance on auto-scaling instead of timesharing. If your app pageloads and issues dozens of queries to populate its widgets, SF will be both slow and expensive.

kwillets··on A reawakening of systems programming meetups
Papers We Love SF has been reconstituting lately; unfortunately I don't know if July has an event yet, but try https://www.meetup.com/papers-we-love-too/ .

We've been meeting online, and I'm interested in setting up physical venues if anyone has one.

The PWL Discord: https://discord.com/channels/1025104619975737446/10251046205...

kwillets··on A reawakening of systems programming meetups
SF Papers We Love might be interested in a physical space -- we've been meeting online, but IMO we would benefit by face-to-face.
kwillets··on Our great database migration
I'm guessing from spreadsheets riddled with per-cell SQL fetches.
kwillets··on Our great database migration
The latency before/after histograms unfortunately use different scales, but it appears that eg the under-200ms bucket is only a few percentage points smaller after the change, maybe 38 before and 33 after.

What I'm curious about is whether Neon can run pg locally on the app server. The company's SaaS model doesn't seem to support that, but it looks technically doable, particularly with a read-only workload.

kwillets··on Surprise, your data warehouse can RAG
Wouldn't a non-DWH SQL database work as well? Most RAG datasets seem like they would fit in a small DB.
kwillets··on Hacker confirms access through infostealer infection [withdrawn]
Ironically this model most resembles Teradata, which used to sell their own proprietary hardware/software combination at exorbitant rates.

Snowflake compute instance types cost about $.30-.40 an hour on EC2, so it's quite a markup.

As far as security I do believe they allow the customers to set their own storage keys, so there may be some isolation from a global breach.

kwillets··on Big data is dead (2023)
I've done it both ways. Look into Data Sketches also if you want to see applications.

The pros:

-- Samples are small and fast most of the time.

-- can be used opportunistically, eg in queries against the full dataset.

-- can run more complex queries that can't be pre-aggregated (but not always accurately).

The cons:

-- requires planning about what to sample and what types of queries you're answering. Sudden requirements changes are difficult.

-- data skew makes uniform sampling a bad choice.

-- requires ETL pipelines to do the sampling as new data comes in. That includes re-running large backfills if data or sampling changes.

-- requires explaining error to users

-- Data sketches can be particularly inflexible; they're usually good at one metric but can't adapt to new ones. Queries also have to be mapped into set operations.

These problems can be mitigated with proper management tools; I have built frameworks for this type of application before -- fixed dashboards with slow-changing requirements are relatively easy to handle.

kwillets··on Big data is dead (2023)
Back in the 80's and 90's NASA built a National Aerodynamic Simulator, which was a big Cray or similar that could crunch FEA simulations (probably a low-range graphics card nowadays). IIRC they found that the queue for that was as long or longer than it took to run jobs on cheaper hardware; MPP systems such as Beowulf grew out of those efforts.
kwillets··on Overflow in consistent hashing (2018)
I believe this problem diminishes as the key count goes up.
kwillets··on Launch HN: Baselit (YC W23) – Automatically Reduce Snowflake Costs
I was thinking about an AI to feed you the proper Snowflake sales pitch each time a query runs expensive or fails a benchmark. At my previous org it could replace several headcount.
kwillets··on Scaling to Count Billions
Yes, I almost added that -- there are several alternatives, but it takes some testing and tuning to be certain.
kwillets··on Scaling to Count Billions
I'm guessing the next stage will be to keep the raw data and aggregate dynamically. We went through a similar progression with ad audiences -- maintaining pre-aggregated data sketches and summaries was such a PITA that we likely saved money by moving to a larger raw-facts database.

> OLAP databases are not good at serving large volumes of requests with low latency within milliseconds.

Just use Vertica lol.

kwillets··on 1BRC merykitty's magic SWAR: 8 lines of code explained in 3k words
Multiplying digit bitfields by their respective powers of 10 and shift/adding via MUL is a (well?) known technique, see Lemire https://lemire.me/blog/2023/11/28/parsing-8-bit-integers-qui... .
kwillets··on SSDs have become fast, except in the cloud
The Nitro chipset claims 100 GB/s encryption, so that doesn't seem to be the reason.
kwillets··on SSDs have become fast, except in the cloud
AWS docs and blogs describe the Nitro SSD architecture, which is locally attached with custom firmware.

> The Nitro Cards are physically connected to the system main board and its processors via PCIe, but are otherwise logically isolated from the system main board that runs customer workloads.

https://docs.aws.amazon.com/whitepapers/latest/security-desi...

> In order to make the [SSD] devices last as long as possible, the firmware is responsible for a process known as wear leveling.... There’s some housekeeping (a form of garbage collection) involved in this process, and garden-variety SSDs can slow down (creating latency spikes) at unpredictable times when dealing with a barrage of writes. We also took advantage of our database expertise and built a very sophisticated, power-fail-safe journal-based database into the SSD firmware.

https://aws.amazon.com/blogs/aws/aws-nitro-ssd-high-performa...

This firmware layer seems like a good candidate for the slowdown.

kwillets··on SSDs have become fast, except in the cloud
EBS is also slower than local NVMe mounts on i3's.

Also, both features use Nitro SSD cards, according to AWS docs. The Nitro architecture is all locally attached -- instance storage to the instance, EBS to the EBS server.

kwillets··on Is the "modern data stack" still a useful idea?
You're on the right track with pricing. What I find ubiquitous about MDS people is a lack of old-school ideas like benchmarking and price-performance. Cloud salespeople talk endlessly about your company becoming data-driven until you bring up cost data.
kwillets··on Almost every infrastructure decision I endorse or regret
To make up for having a better schema in Terraform than in the database.
← PreviousPage 3 of 13Next →