HNHacker News
TopNewBestAskShowJobs

bdarnell

1,126 karma · joined August 9, 2010

Co-founder & CTO of Cockroach Labs
submissionscomments
bdarnell··on Clerihew
Google Reader's maintenance/downtime page was a Clerihew (preserved at readerisdead.com)

    Google Reader's
    built with electrons and leptons, meters and liters.
    We're off dealing with those particles
    so we can bring you your articles.
bdarnell··on Infectious Executable Stacks
This misfeature can show up in surprising places. Any use of assembly (`.s`) files in a Go program will by default add the executable-stack flag, as we discovered in CockroachDB: https://github.com/cockroachdb/cockroach/issues/37885
bdarnell··on Regex Crossword
I also like this larger hexagonal one: https://rampion.github.io/RegHex/
bdarnell··on Show HN: Bashfs – Run commands as your filesystem
It's mainly for fun, but to find cases where it might be useful, think about symlinks into this filesystem. `ln -s /bashfs/somescript.sh /etc/hosts` and then your hosts file can be dynamically generated without any changes to programs that read it. It can be a clever way to sneak dynamic behavior into contexts where there would normally just be a static file.
bdarnell··on Ask HN: How can we destroy AMP?
Don't destroy it, embrace it. AMP is the low-javascript web many of us pine for. If browsers implemented AMP natively (and it were properly standardized), it would be great, and it has tons of potential for things like RSS readers. Google's heavy-handedness here has a lot of problems, but if they've managed to get publishers on board with a move away from the JS quagmire we should be pushing in the same direction and not fighting against it.
bdarnell··on Relicensing CockroachDB
> if I'm using CockroachDB and don't want to manage it myself, my only option is Cockroach Labs.

Cockroach Labs is not your only option; you can also use other providers that have a license agreement with us. Our first partner in this area is ObjectRocket: https://www.objectrocket.com/blog/cockroachdb/introducing_co...

bdarnell··on Relicensing CockroachDB
I'd like to thank our friends at TimescaleDB for the idea to use schema control as the dividing line (they're doing the same thing in their license)
bdarnell··on Relicensing CockroachDB
(Cockroach Labs founder)

The details are in the "additional usage grant" clause: https://github.com/cockroachdb/cockroach/blob/8acfe8ffd0028c...

We decided to draw the line at whether the end user has direct control over table schemas. If users can specify the schema to be used, it's a database service and needs a license. If you're fitting everything into a generic schema (even if the user can specify things that look like new columns in the UI), it's an application and doesn't need a special license.

bdarnell··on CVE-2019-9193: Not a Security Vulnerability
I think this is an argument for not giving root/superuser all possible permissions by default. It's OK that granting the `pg_execute_server_program` permission gives access to this feature, but it should still be something you have to opt in to, instead of making database superuser equivalent to the host user that the database runs as.

For comparison, in CockroachDB (disclosure: I'm a co-founder of Cockroach Labs), we don't have any features that let you execute server programs, but we do have something analogous to `pg_{read,write}_server_files` via the BACKUP, RESTORE, and IMPORT commands. In order to use these commands with a target on the server's filesystem, though, the database `admin` role isn't enough. The server also needs to be started with the `--external-io-dir` flag (and file operations will be limited to that directory). This gives an extra layer of opt-in before filesystem operations are allowed.

bdarnell··on Emacs 26.2 Released
Here's the NEWS file for the curious: http://git.savannah.gnu.org/cgit/emacs.git/tree/etc/NEWS?h=e...
bdarnell··on Serializability vs. Strict Serializability: The Dirty Secret of Isolation Levels
> Note, I don't believe it's necessary for _all_ transactions to be strictly serializable, but I would make the argument that all read-write transactions should be in order to prevent the anomaly in the article. In other circumstances it's reasonable to make the tradeoff and accept potentially stale data.

How exactly does the "free checking" example work?

Is it two transactions, where each transaction is responsible for checking the total balance across all accounts? Or is it three transactions, where a third read-only transaction verifies the sum of all accounts?

If it's two transactions, I think the anomaly as described is prevented even by non-strict serializability, because the additional reads force the transactions to conflict. If it's three transactions, then what we're really concerned with is the read-only transaction. What does it mean to mix some strictly-serializable transactions and non-strictly-serializable ones?

bdarnell··on Ask HN: How should a programming language accommodate disabled programmers?
T.V. Raman (author of Emacspeak) has said that "S-Expressions are a major win" for visually-impaired users because they give you a good way to navigate the code structurally ("go to beginning of expression" is more accessible than "go up one line").

https://tvraman.github.io/emacspeak/manual/Emacspeak-And-Sof...

bdarnell··on Colorizing Stderr: racing pipes, and libc monkey-patching
This is significant for programs written in Go, which tends to make system calls directly instead of using the libc wrapper functions. LD_PRELOAD trickery often doesn't work with Go programs for this reason.
bdarnell··on Show HN: SQL Trainer – Learn SQL by doing live data exercises
In theory `information_schema` is supposed to be universal, so you could do `select table_name from information_schema.tables`. However, getting useful information out of `information_schema` tends to require fairly complex queries and still has quirks and differences between databases.
bdarnell··on Testing Memory Allocators: ptmalloc2 vs. tcmalloc vs. hoard vs. jemalloc
Performance isn't the only thing to consider. For example, tcmalloc and jemalloc both have good profiling/debugging tools, and these are the biggest reasons why I choose one of them for any large C/C++ project. I've also found that jemalloc is easier to integrate into complex build systems than tcmalloc, so jemalloc is my first choice in most cases.
bdarnell··on Python 1.0.0 is out (1994)
I learned Python 1.4 on DOS (with DJGPP). I think there were binary releases, but you couldn't use them if you needed any third-party C modules. There was no dynamic linking on DOS so you needed to statically link any C code you were using into python.exe. This was a huge pain but it also taught me a ton about both C and Python.
bdarnell··on GNU Readline for Better Productivity on the Command Line
I'm also a fan of emacs shell and tramp, and use functions like this to start shells on remote boxes without the extra dired/whatever step:

  (defun shell-foo()
    (interactive)
    (let ((default-directory "/ssh:bdarnell@foo:/home/bdarnell/"))
      (shell "*shell-foo*")))
bdarnell··on CockroachDB 1.1 Released: Production Made Easy
> When would you need to revoke an individual cert and why wouldn't that be better handled by just shutting down the VM or container instead?

You revoke a cert when it's somehow been compromised and something other than the VM/container that's supposed to has it gets a copy of it.

bdarnell··on CockroachDB 1.1 Released: Production Made Easy
What "standard auth options" do you have in mind for node-to-node auth? The ones that come to my mind are even more of a pain to set up than TLS certs.

It's true that setting up a secure cluster is kind of annoying right now. But the kubernetes templates (https://github.com/cockroachdb/cockroach/tree/master/cloud/k...) do support secure mode now, and the plan is to provide more like this so it's not something that everyone has to solve by hand.

If you know or can predict the addresses or hostnames you'll be using, then it's possible to generate one cert and reuse it for multiple nodes. This isn't ideal from a security perspective since you lose the ability to revoke individual certs, but since we don't (yet) support CRLs/OCSP it's not much of a loss.

Adding an option to skip hostname checks for node certs might make this less of a pain (it would then be trivial to share one cert for all nodes if that's what you want to do). We'll consider that and see if it compromises any important security properties.

bdarnell··on CockroachDB 1.1 Released: Production Made Easy
Some of those have been done, but not all, so it's still not possible to use CockroachDB with the Django ORM. The biggest blocker (IIRC) is that the default set of migrations that Django performs on a new database uses ALTER COLUMN TYPE.

Some of the features from those lists we have implemented include ALTER COLUMN SET DEFAULT, pg_table_is_visible(), UUID, extract(), and (some) schema changes in transactions. 1.2 will add (at least) INET types and sequences.

bdarnell··on Use pew, not virtualenvwrapper, for Python virtualenvs
The basic tools (pip and virtualenv) have been the same for years. Posts like this are about layers that people have built on top of the same underlying tools. There's a lot of them, but that's because different people have different preferences, not because those preferences are changing over time. There's no need to chase the latest trend here (if indeed there is a trend).
bdarnell··on Use pew, not virtualenvwrapper, for Python virtualenvs
Yes, exactly. bin/pip inside the virtualenv installs into that virtualenv.
bdarnell··on Use pew, not virtualenvwrapper, for Python virtualenvs
You can use virtualenv without `bin/activate`. Just refer to `bin/python` (and other scripts in `bin/`) explicitly. Don't bring in the magic until you need it. And when you need it, consider what kind of magic you need. If you're typing `bin/` too much, do you need something that adds `bin/` to your path, or do you need to write a script for this task that saves you from having to type both `bin/` and a bunch of other command-line flags?

(I do add my virtualenv's `bin/` directory to my path, but mainly so my editor can find the right tools, so I do it from elisp instead of bin/activate in a shell)

bdarnell··on The Limits of the CAP Theorem
Spanner offers "external consistency", which is similar to linearizability.

CockroachDB offers serializability (as the default isolation level). Most other SQL databases offer serializability as their highest isolation level, but default to something weaker.

Serializability and linearizability are not equivalent, although it's not always easy to devise a scenario in which the differences are apparent. The "comments" tests in Aphyr's Jepsen analysis of CockroachDB is one such scenario: http://jepsen.io/analyses/cockroachdb-beta-20160829#comments

bdarnell··on The Limits of the CAP Theorem
There's a lot of complexity here, but the short version is that nodes cannot quite come and go freely. A replica set keeps track of its members; a quorum of the existing members must vote to admit a new node, or to remove one.

Note that in both CockroachDB and Spanner a cluster contains many independent and overlapping replica sets. The data is broken down into "ranges" (to use the terminology of CockroachDB; Spanner calls them "spans"), each of which has its own replica set (typically containing 3 or 5 members).

bdarnell··on The Limits of the CAP Theorem
That's an important point. With mobile applications that support offline usage, you can no longer assume a single global source of truth, and the application as a whole is AP.

However, I'd argue that this tilts the balance even more in favor of a CP database on the backend. Even when the client application is not executing transactions on the database, consistency at the database level is what makes it possible to support secondary SQL indexes that work without surprises. An offline-capable mobile app buffers writes, moving the write to the server out of the critical path so server-side write-latency is not as visible to the user.

bdarnell··on The Limits of the CAP Theorem
Yes, that's correct. If you are partitioned away from the lease holder, then you can't read until either the partition heals, or the lease expires and a new lease is granted to a node that you can reach.

Writes require the lease too, so it is not possible for a quorum of nodes on one side of a partition to serve writes while a lease holder on the other side serves stale reads. This is a degenerate case of quorum leases (https://www.cs.cmu.edu/~dga/papers/leases-socc2014.pdf) for a single lease holder; in the future we're interested in supporting multiple lease holders to improve read latency (at the expense of write performance and availability).

bdarnell··on Be Careful with UUID or GUID as Primary Keys
One of the post's points is that UUIDs will scatter your writes across the database, and that for this reason you want a (more or less) sequential key as your primary key. This crucially depends on both your database technology and your query patterns.

In a single-node database or even a manually-sharded one, this post's advice is good (For Friendfeed, we used a variation of the "Integers Internal, UUIDs External" strategy on sharded mysql: https://backchannel.org/blog/friendfeed-schemaless-mysql).

But in a distributed database like CockroachDB (Disclosure: I'm the co-founder and CTO of Cockroach Labs) or Google Cloud Spanner, it's usually better to get the random scattering of a UUID primary key, because that spreads the workload across all the nodes in the cluster. Sometimes query patterns benefit enough from an ordered PK to overcome this advantage, but usually it's better to use randomly-distributed PKs by default.

For CockroachDB, my general recommendation for schema design would be to use UUIDs as the primary keys of tables that make up the top level of an interleaved table hierarchy, and SERIAL keys for tables that are interleaved into another. (Google's recommendations for Spanner are similar: https://cloud.google.com/spanner/docs/schema-design#choosing...)

bdarnell··on PyPy v5.8 released
In addition to being incompatible with (some) third-party libraries, pypy tends to use significantly more memory than cpython. It's also slower than cpython for scripts that don't run long enough to warm up the JIT, so you probably wouldn't want to use it by default. (Disclaimer: I'm basing this on experience with older versions of pypy and haven't verified it recently)
bdarnell··on The Gilectomy – How's It Going [video]
The big constraint (aside from backwards compatibility) is performance: Guido has indicated that he is unwilling to accept much (if any) slowdown of single-threaded code in order to remove the GIL. It's (relatively) easy to remove the GIL and replace it with a bunch of fine-grained locks (or atomic increments, etc), but doing so tends to slow things down. The challenge is in figuring out how to avoid synchronization overhead for common operations (mainly reference counts).

It's true that buffered refcounting probably means that `__del__` would no longer be called immediately as it is now, but I'm not sure if that's a requirement - pypy and jython don't do this either, and destructors are generally discouraged in favor of `with` blocks these days.

Page 1 of 5Next →