HNHacker News
TopNewBestAskShowJobs

rwultsch

164 karma · joined August 19, 2015

submissionscomments
rwultsch··on PostgreSQL for Everything
"MySQL was also potentially faster as it did not implement all features of the SQL standard. "

This is a not great start. I assume it refers to MyISAM which has not been relevant for over a decade at this point. InnoDB made different design than PG decisions and was (and perhaps still is) faster at point lookups.

rwultsch··on Why Canada Should Join the EU
Does this not speaks to the rapid growth of Central Europe economies rather than a decline of Japan? I have read similar comparisons for both the U.K. and Germany vs Poland.

Also, a quick googling suggests both Poland and Japan both have fertility rates around 1.2. Working age of Poland is 65% of the population vs 60% for Japan. So a bit worse.

rwultsch··on Why Canada Should Join the EU
Oddly enough I am visiting Japan right now. There are lots of foreign workers in service jobs. They speak zone language and don’t jay walk.

Prices feel like a middle income country, but that is just the Yen sucking. Otherwise it feels very first world.

rwultsch··on Pattern of brain damage is pervasive in Navy SEALs who died by suicide
They are however incredibly loud. Being close to someone shooting 556 with a brake is a really annoying experience and I would not be shocked if there were follow on effects.
rwultsch··on First randomized trial of Ozempic for alcoholism shows big drops in drinking
I am an overweight person that drank more than I should. Ozempic helped me drop 15 lb and I don’t feel the draw of drink nearly as much as I had.

I was diagnosis ADHD as a kid. Since being on Ozempic I have noticed no difference in ability to hold attention.

rwultsch··on HBase Deprecation at Pinterest
“Introduced in 2013, HBase was Pinterest’s first NoSQL datastore.”

I don’t think this is correct. When I started in late 2013 Redis was being used as a persistent data store. And what pain it was. I convinced leadership in late 2014 this was a bad and they had me keep it alive until it was replaced by MySQL in mid 2015.

HBase was nothing but pain at Facebook where it was supposed to replace MySQL and then Pinterest where… I think there was hope it would replace MySQL. Once I automated MySQL at Pinterest I think it wasn’t so bad, particularly given the absurdly limited staff they gave the problem.

rwultsch··on Google to pause Gemini image generation of people after issues
I enjoy that “ma” has ambiguous meaning above. Does it mean mandarin question mark word or does possibly mean mother?
rwultsch··on How Pinterest scaled
A pin was a 1.2 KB json blob. There were other tables but pins was the big one. Why MySQL? It did not destroy data like the alternatives.

how storage became efficient https://medium.com/pinterest-engineering/evolving-mysql-comp... https://medium.com/pinterest-engineering/evolving-mysql-comp...

rwultsch··on How Pinterest scaled
Before I joined in late 2013, they had not known how to run schema change without downtime. Once we fixed the kernel the db’s were nearly completely untaxed in terms of performance. They did however need large instances due to disk usage.
rwultsch··on A fridge from 70 years ago has better features than the fridge I own now
I am getting a Sub-zero in a few weeks which is damn near the most expensive fridge. It does not have nice pull out shelves or the veggie compartments.

I expect the sealing and ethylene scrubbing will keep veggies fresh linger.

rwultsch··on Ask HN: How to find a small town to relocate for remote work?
We recently moved to Frederick, MD. Walking distance to the very lively downtown, inexpensive housing, close to an airport, yada, yada...
rwultsch··on Tell HN: I let my 6-year-old daughter design my website
OT: How is the tech scene in Taiwan?

My wife is from Taiwan. Every time we visit I am very sad to leave. We have thought about moving to Taipei.

rwultsch··on Junk – Mark Bittman’s History of Why We Eat Bad Food
And Taiwan.
rwultsch··on Why Uber Engineering Switched from Postgres to MySQL (2016)
FB had plenty of schema changes. I know, I wrote software to push them out. The important concepts for pushing schema at scale became part of Skeema.
rwultsch··on Bolsonaro fires health minister, calls to reopen economy
Taiwan.
rwultsch··on The sound of the Hagia Sophia, more than 500 years ago
Eh, I will kind of buy that. It was a church before it was a mosque and churches tend to have quite a bit more music than mosques.

The race war comment above seemed uncharitable.

rwultsch··on The sound of the Hagia Sophia, more than 500 years ago
The Hagia Sophia was the site of a significant massacre when it was forcibly converted from a church to a mosque.
rwultsch··on Command-line tools can be faster than a Hadoop cluster (2014)
I was on the DBA team there around that MAU. I had lots of company and support from SRO (aka jr DBA's) as well as other teams (provisioning, etc...).
rwultsch··on Facebook, Instagram go down around the world in an apparent outage
Short duration: network, bad software deploy Long duration: db. If you break data, it takes a while to unbreak.

Source: Me. My career has been spent managing db's for internet scale sites.

rwultsch··on Reasons to choose PostgreSQL 9.6
Postgres has had CTE for a while, but not that long. MySQL 8.0 (the next version) plans to them http://mysqlserverteam.com/mysql-8-0-labs-recursive-common-t...
rwultsch··on PostgreSQL 9.6 Released
You missed the point of my post. You are going to have one of the two issues, either looking through two index or indexes including the a large PK. At least with InnoDB you can make the choice. The strategy I suggested gets you the desired outcome of not including a large PK in all secondary indexes.
rwultsch··on PostgreSQL 9.6 Released
If you don't want a clustered index in InnoDB you can define the primary key as an auto incrementing uint.
rwultsch··on PostgreSQL 9.6 Released
The PG storage engine is not particularly awesome. It is basically COW (with exceptions) and compaction (called vacuum) has been quite painful for a long time. Every release it is supposedly fixed, but people keep complaining. This not to say PG sucks, their optimizer knows far more about their data than InnoDB and PG can perform far more types of execution plans.

We (Pinterest, I wrote most of the MySQL automation) make heavy use of MySQL replication which is vastly simpler to manage than PG. All queries still flow through SQL and unlike PG, we can force whatever execution plan we need. We do lots of PK lookups, and InnoDB is really good at that. In InnoDB all the data is stored in the PK while in PG it is just a pointer.

rwultsch··on Ingesting MySQL data at scale – Part 1
Hi cookiecaper, All our operational code (other than pinterest specific bits) is open sourced https://github.com/pinterest/mysql_utils

Both statement and row based replication can be very reliable in modern versions of MySQL. It is my experience the ways data is corrupted are: 1. read_only not being set on slaves, so random users can write to the slave. We set read_only on startup based on service discovery. 2. Bad automation for failovers. See https://github.com/pinterest/mysql_utils/blob/master/mysql_f... for how we do it. 3. Crashes without all the durability settings being on.

If you are having to run slave_skip_errors, you are doing it wrong. You should checkout out our automation for backups and restores. They can be found in mysql_restore.py and mysql_backup.py .

With regard to PG replication, I suggest you watch https://m.youtube.com/watch?v=bNeZYVIfskc&t=26m54s Uber had a 16 hours outage in large part caused by pg replication issues.

Vacuum issues are also no joke.

There are a host of other issues: MySQL can deal with large numbers of connections, PG needs middleware. MySQL is more efficient with "web" workloads where most queries only need to pull on row. etc...

-Rob

rwultsch··on Ingesting MySQL data at scale – Part 1
Hi, I was the first MySQL DBA hired Pinterest and before that I worked at Facebook and GoDaddy. At none of these places did we run active/actice. One of the first things I did at Pinterest was rip out the multi-master configure because it is dangerous.

Why MySQL? Really easy replication, no vacuum, and great point lookup performance.

rwultsch··on When should you store serialized objects in the database? (2010)
I was on the DBA team at FB and I spent the better part of a year working on the deployment system for online schema change. It was a pain. Other companies have done quite a bit of work on this as well (Shift from Square, etc...).

Later on I joined Pinterest as their first MySQL DBA. They had copied the sharding system from FB, but instead of having a bunch of columns, they just stored a JSON blob. This saved them from learning how to perform schema change until I joined the company. This is a pretty incredible feature.

We have a new feature under development (which will be open sourced as part of Percona MySQL) which will allow column level compression with an optional predefined dictionary. During testing, this resulted in a 30% additional reduction in spaced consumed versus InnoDB page compression AND doubles our peak QPS at lower latency. This would not work well with many individuals columns, but kicks ass for JSON blobs.

http://www.slideshare.net/denshikarasu/less-is-more-novel-ap... (slides 37, 40, 41, 42)

rwultsch··on Open-sourcing Pinterest MySQL management tools
It has not open sourced and I think (and am in no way authoritative) it is unlikely to be open sourced.
rwultsch··on Open-sourcing Pinterest MySQL management tools
(I am the author)

I wrote most of these tools last year and a lack of good example code definitely slowed me down. Even if we drop support for these tools, at least the code now is open sourced as a reference. This was one of the primary motivations for open sourcing this code.