HNHacker News
TopNewBestAskShowJobs

apavlo

696 karma · joined November 4, 2020

https://www.cs.cmu.edu/~pavlo/
submissionscomments
apavlo··on Databases in 2024: A Year in Review
> Wrt Andy, here are [1] somehow interesting views from (presumably) previous employees.

I am only seeing this now and I take the complaints about being "slightly racist and offensive" very seriously. I am checking with investors, former HR people, and co-founders. I was not made aware of any issues. If anything, I was overly cautious at the company.

I was openly transparent with our employees about every direction the company was pursuing up until the very end. The complaint that "He thinks he knows everything about business" makes me believe this person is just trolling because I was always the first to admit in meetings that I was not an expert in how to run a business. We had to fire people because of inappropriate behavior, but not because I had strong disagreements with how to run the company.

apavlo··on Databases in 2024: A Year in Review
> Wow, the reasons why Redis commands API suck in Andy's video (linked in the post) are the weakest ever.

In my example, the API on a key changes based on its value type. And the same collection can have different value types mixed together. You've recreated the worst parts of IBM IMS from the 1960s. However, the original version of IMS only changed the API when a collection's backing data structure changed. Redis can change it on every key!

We didn't get into the semantics of Redis' MULTI...EXEC, which the documentation mischaracterizes as "Transactions". I'm happy that at least you didn't use BEGIN...COMMIT.

apavlo··on Database Tools in 2024: A Year in Review
Instead of emailing me to ask if you can translate my end-of-year blog article to Chinese like you did last year, this year you decided to just copy my format? I will be sure to mention that when I release mine next week.
apavlo··on MVCC – the part of PostgreSQL we hate the most (2023)
If all table pages exist in memory and you are using cooperative GC, then O2N can be preferable. As workers scan version chains, they can clean up dead tuples without taking additional locks.

This is what Microsoft Hekaton does.

apavlo··on Show HN: EloqKV – Scalable distributed ACID key-value database with Redis API
> EloqKV brings significant innovations to database design

What is the novel part? I read your "Introduction to Data Substrate" blog article and the architecture you are describing sounds like NuoDB from the early 2010s. The only difference is that NuoDB scales out the in-memory cache by adding more of what they call "Transaction Engine" nodes whereas you are scaling up the "TxMap" node?

See also Viktor Leis' CIDR 2023 paper with the Great Phil Bernstein:

* https://www.cidrdb.org/cidr2023/papers/p50-ziegler.pdf

* https://youtu.be/tiMvcqIfWyA

apavlo··on Show HN: InstantDB – A Modern Firebase
This is from me. I didn't realize the connection to Lutris + Enhydra. It should be listed as a "Acquired Company" + "Abandoned Project". Wikipedia also says that it lasted until 2001. Usage is different from development/maintenance. I will update the entry for the old InstantDB and add an entry for this new InstantDB.

I think given that the original InstantDB died over two decades okay and is not widely known/remembered, reusing the name is fine.

apavlo··on Ask HN: BTree Alternative for Storing to Disk?
> Can you try to use a memory-mapped file

Definitely do *not* use MMAP for this.

apavlo··on Ash HN: What are some good resources on building a relational database?
You want our Intro DB Systems course not the Advanced one:

https://15445.courses.cs.cmu.edu

Lectures start next month. Or you can watch previous years. Learn to walk before you run.

apavlo··on CedarDB: German-Powered, PostgreSQL-Compatiable Freak of Nature Database System
CedarDB is the commercial version of the impressive Umbra database system from TU Munich.
apavlo··on An Empirical Evaluation of Columnar Storage Formats [pdf]
Lance v2 looks interesting. I like their meta-data + container story. Lacking SOTA encoding schemes though.

There is also Vortex (https://github.com/fulcrum-so/vortex). That has modern encoding schemes that we want to use.

BtrBlocks (https://github.com/maxi-k/btrblocks) from the Germans is another Parquet alternative.

Nimble (formerly Alpha) is a complicated story. We worked with the Velox team for over a year to open-source and extend it. But plans got stymied by legal. This was in collaboration with Meta + CWI + Nvidia + Voltron. We decided to go a separate path because Nimble code has no spec/docs. Too tightly coupled with Velox/Folly.

Given that, we are working on a new file format. We hope to share our ideas/code later this year.

apavlo··on An Empirical Evaluation of Columnar Storage Formats [pdf]
> Have a situation right now where I am bottlenecked by IO and not compute.

Can you describe your use-case? Are you reading from NVMe or S3?

apavlo··on Databases in 2023: A Year in Review
> it seemed like it might be genuine

It is genuine. Larry and I go way back.

Source: https://youtu.be/_UGUDVeSQSI?t=872

apavlo··on Building a faster hash table for high performance SQL joins
> just wanted to mention that I found bidirectional linear probing to outperform Robin Hood across the board in my Java integer set benchmarks

Research results from the last five years shows that Robin Hood hashing performs better than the other approaches under the right conditions. See this eval paper:

https://15721.courses.cs.cmu.edu/spring2023/papers/11-hashjo...

apavlo··on The subtleties of proper B+Tree implementation
> No, you don't.

This is correct. Most disk-oriented DBMSs use fixed-size pages. So you can't go larger than the node size (not entirely true if using auxiliary data structures to buffer changes like an inefficient -epsilon tree).

apavlo··on My favorite database shirts
No. I've been meaning to do an updated article but I've been unfortunately too busy. A bunch of companies send us shirts to give out to students:

* https://twitter.com/andy_pavlo/status/1659019035818729472

* https://twitter.com/andy_pavlo/status/1335045678876270592

* https://twitter.com/andy_pavlo/status/1125465168023048193

* https://twitter.com/andy_pavlo/status/996191088372322304

* https://twitter.com/andy_pavlo/status/862320227601850368

My highlights from the last 6-7 years:

* DuckDB (European embroidery!)

* Materialize (like Snowflake)

* Yellowbrick (wild designs)

* Timescale (old logo was popular)

-- Andy

apavlo··on ChDB: Embedded OLAP SQL Engine Powered by ClickHouse
Where can I find an SVG version of this logo? https://github.com/chdb-io/chdb/raw/main/docs/_static/snake-...
apavlo··on Show HN: SquareDB; a database in 200 lines of code
Do you have a logo?

200 lines is a bit misleading since RocksDB is doing all the heavy lifting.

apavlo··on Ask HN: Selling a database engine as a solo developer?
I am going to guess Helium?
apavlo··on Ask HN: Selling a database engine as a solo developer?
Why would any company take the risk on adopting a single-person passion project DBMS that is leftover from failed startup?

The best you can hope for is to make a logo, slap it on Github, and then I'll add it to my list:

https://dbdb.io/browse?tag=failed-company

apavlo··on Home – Database of Databases
I will fix. Thank you.
apavlo··on Home – Database of Databases
This is me. I only care about databases.

What was your system?

apavlo··on Garbage Collection for MVCC – MyRocks, InnoDB and Postgres
> I wonder why they chose such an unrepresentative dataset size, ~84MB.

Yo. This is me. The point of this paper was to evaluate the core MVCC algorithms under different contention scenarios. So you strip out all the internal features of the system that can influence performance that are unrelated to the experiments. This ensures that you can make it a true apples-to-apples comparison. And since everything is in memory, you don't need a large data set.

IIRC, we ran ran the same experiments with 100m tuples instead of 10m and it did not change the results.

apavlo··on Database Gyms [pdf]
> I had a feeling this was OtterTune's product through a reading of the abstract before I noticed the names.

No, this project is separate from OtterTune. I keep our CMU research strictly firewalled from OtterTune for legal reasons.

The Database Gym project rose up from the ashes of the NoisePage self-driving DBMS project. See my comment from last week about why it failed:

https://news.ycombinator.com/item?id=36355963

apavlo··on Building a new database management system in academia (2017)
See our 2014 paper on evaluating CC protocols on in-memory system with high contention / core counts:

https://www.vldb.org/pvldb/vol8/p209-yu.pdf

All the protocols regress to the same. This evaluation was only with stored procedures though. It would be worth doing a similar investigation with conversational DB protocols (e.g., JDBC, ODBC).

apavlo··on Building a new database management system in academia (2017)
LingoDB is an interesting system. Jana has done great work with it. I like projects that take unorthodox approaches to old problems.

The problem with (most) query optimizers is that they take a one shot approach at optimization. I think an optimizer should be built from the groundup to support adaptive query optimization. Something similar to Berkeley's Eddies project from 20 years ago.

apavlo··on Building a new database management system in academia (2017)
Yikes! Thanks for the heads up. Peter left Stanford so I guess they took over the domain name :-(
apavlo··on Building a new database management system in academia (2017)
Actually, it was a combination of three things:

1. OtterTune Start-up (https://ottertune.com)

2. Biological Daughter (https://twitter.com/andy_pavlo/status/1187841279260004355)

3. Pandemic

When the pandemic first started, I had a bunch of CMU students reach out to me saying that their summer internships were rescinded and that they were looking for a project to work on so that they wouldn't have a gap in their CV. I ended up taking on any student that could program C++ even if they hadn't taken my DB class before. It as my way of trying to help. But our research group grew to about 35 people. That was not sustainable and the code quality suffered greatly.

We ended killing the project and now all our self-driving work is done in the context of Postgres (https://db.cs.cmu.edu/papers/2023/p27-lim.pdf).

I also now realize that building the DBMS engine first then building the query optimizer second is the wrong order. Our future project is going to start with the optimizer first.

apavlo··on The Great Graph Debate: Revolutionary concept in databases or niche curiosity?
> Database engines are the last place you want to be playing in traffic if you are going to be on the hook for support.

This is a great line! I'm going to steal it if that's okay with you.

apavlo··on Ask HN: Papers to get up to speed on database research
Watch my course lectures from last semester:

https://15721.courses.cs.cmu.edu/spring2023/

apavlo··on Nobody cares about our concurrency control research [pdf]
I have no idea what it's supposed to say. I think I just put "Andy Pavlopolis" into Google Translate. It's meant to be a joke. Before I started at CMU, all of the database professors in the city of Pittsburgh (Carnegie Mellon + University of Pittsburgh) were Greek. So the joke is that the only reason that I got hired at CMU was because I pretended to be Greek with a fake Greek name.

It's in the video here:

https://www.youtube.com/watch?v=M2MEcvMHzkY&t=4507s

← PreviousPage 2 of 5Next →