HNHacker News
TopNewBestAskShowJobs

bdarnell

1,126 karma · joined August 9, 2010

Co-founder & CTO of Cockroach Labs
submissionscomments
bdarnell··on Spanner: Becoming a SQL System [pdf]
(Cockroach Labs CTO here)

> > But you don't need an atomic clock to get Spanner's guarantees.

This comment continues "...but GPS clock-sources will do just fine for Spanner's 6ms quantum". Providing Spanner's guarantees with reasonable performance requires specialized hardware, but there are more options for that specialized hardware than just atomic clocks.

Note that Spanner itself uses both atomic and GPS time sources according to Google's publications; when we talk about "atomic clocks" we're usually talking about the entire category of specialized time-keeping hardware instead of distinguishing atomic clocks from GPS clocks.

> I always hear this from CockroachDB folks and fans, but no details. What are the downsides?

As we describe in our blog post (https://www.cockroachlabs.com/blog/living-without-atomic-clo...), CockroachDB on commodity hardware provides a slightly weaker consistency model than Spanner (serializable instead of linearizable), and latency is sometimes higher as we need to account for the larger clock offsets in certain situations.

If you do have a high-quality time source available, we have an experimental option to use a Spanner-like linearizable mode.

bdarnell··on CockroachDB 1.0
We have started parallelizing our tests with the new subtest feature: leaktest in the top-level test, t.Parallel in the subtests. This means we only check for leaks in between batches of parallel subtests. This works OK for us for now since our slowest "test" is really a huge data-driven test suite, and that's the only place we're currently parallelizing, although it would be better if we could parallelize more of our tests.
bdarnell··on CockroachDB 1.0
This bootstrapping problem is tricky. We publish kubernetes templates at https://github.com/cockroachdb/cockroach/tree/master/cloud/k... that contain our current best solution for the join/init problem.
bdarnell··on CockroachDB 1.0
I talk about this in the presentation I linked in another subthread (https://www.cockroachlabs.com/community/tech-talks/challenge...). The key to getting good performance out of any GC is to generate as little garbage as possible, and in our experience Go makes better use of stack allocation and value types keep many objects out of the garbage-collected heap. We've found that idiomatic go programs tend to produce less garbage than similar java programs, and in the presentation I discuss some tricks we use to get that even lower in critical paths. Admittedly, we're not JVM tuning wizards so maybe there's more that could have been done on the JVM side.
bdarnell··on CockroachDB 1.0
Thanks v3ss0n!
bdarnell··on CockroachDB 1.0
Currently, you need to run one node without --join for the initial bootstrapping (as soon as this bootstrapping is complete, you can and should restart it with --join to get everything into a homogenous configuration). I was hoping to make some changes here so you could start every node with --join from the beginning, but it was trickier than anticipated so it didn't make the cut for 1.0. Watch for improvements here in a future release.
bdarnell··on CockroachDB 1.0
(Cockroach Labs co-founder) I gave a talk on exactly this subject at the ACM Applicative conference last year: https://www.cockroachlabs.com/community/tech-talks/challenge...

Overall we've been happy with the choice. The GC is sometimes a performance issue, but it's manageable (and Go gives you better tools to limit the cost of GC than many other garbage-collected languages)

bdarnell··on The SQL layer in CockroachDB
Ah, right. We did turn on TCP keepalives so we probably don't need a protocol-level ping (although we often deploy behind haproxy, so tcp keepalives need to be configured at each hop)
bdarnell··on The SQL layer in CockroachDB
(Cockroach Labs CTO) Yes, the commit identifiers sound like something we could use, as long as they're not tied too closely to pg's WAL LSNs. For us, the token would be a 96-bit hybrid logical timestamp, and we'd want semantics roughly similar to HTTP cookies (the server sends it back with the response, and then that client will include it on any future requests. With our current implementation the client would only need to send the token on new connections, but including room for it on a per-request basis seems like a good idea).

Also on the subject of the protocol, we've run into a need for a (server-initiated) ping message (https://github.com/cockroachdb/cockroach/pull/10188).

bdarnell··on Spanner vs. Calvin: distributed consistency at scale
(disclosure: CockroachDB founder) The reason I've heard is that Spanner uses a separate mutation API instead of SQL DML because of a quirk of its transaction model. Writes within a transaction are not visible to subsequent reads within the same transaction (source: https://cloud.google.com/spanner/docs/transactions#rw_transa...). This is different from other SQL databases, so the use of a different API forces you to think about your read/write transactions in a Spanner-specific way.

(FWIW, CockroachDB does not have this limitation - transactions can read their own uncommitted writes, just like in other databases)

bdarnell··on VPNs are not the solution to a policy problem
I think "other risky servers" may refer to the lesser-known servers that streisand includes, like shadowsocks.
bdarnell··on VPNs are not the solution to a policy problem
I've used streisand on DO (while traveling in China) and it worked well. There's also a similar project called algo[1] which provides a single protocol with maximum security, in contrast to streisand's multi-protocol flexibility (and increased surface area).

https://github.com/trailofbits/algo

bdarnell··on In search of a simple consensus algorithm
The raft paper states that its use of an explicit leader election was chosen for understandability, not performance. Leaderless systems can potentially improve performance, especially when replicas are far apart (so you don't have the extra network hop to the leader before you can get started). The main drawback IMHO to ePaxos (as a hopefully representative leaderless algorithm) is not performance but its complexity, since it intrudes into the application to track dependencies and conflicts between proposed commands.

If I squint at the scheme proposed here, it looks like ePaxos taken to the extreme of a single key/value per paxos group, so conflict tracking becomes trivial. It has demonstrated good performance in the case of uncontended writes; I'd be curious to see how it behaves under contention, or when a downed node rejoins the cluster.

bdarnell··on CockroachDB beta-20161013
(Cockroach Labs CTO) CockroachDB provides two transaction isolation levels: SERIALIZABLE, which is our default and the highest of the four standard SQL isolation levels, and SNAPSHOT, which is a slightly weaker mode similar to (but not quite the same as) REPEATABLE READ. Unlike most databases which default to lower isolation modes like REPEATABLE READ or READ COMMITTED, we default to our highest setting because we don't think you should have to think about which anomalies are allowed by different modes, and we don't want to trade consistency for performance unless the application opts in.

All transaction interactions are localized to particular keys, so other transactions can proceed normally while there is contention in an unrelated part of the database.

(as for the acronym, we prefer CRDB instead of CDB)

bdarnell··on CockroachDB beta-20161013
(Cockroach Labs CTO) The major issue with clock offsets is stale reads due to the time-based read lease. Pretty much everything else is independent of the clocks. Stale reads are still enough to get you in trouble (some of the jepsen tests will show serializability violations if you disable the clock-offset suicide and mess with the clocks enough), but it's tricky to actually get bad data written into the database, as opposed to returning a stale-but-consistent view of the data.
bdarnell··on CockroachDB beta-20161013
(Cockroach Labs CTO here)

Jepsen is essentially a worst-case stress test for a consistent database: all the transactions conflict with each other. CockroachDB's optimistic concurrency control performs worse with this kind of workload than the pessimistic (lock-based) concurrency control seen in most non-distributed databases. But this high-contention scenario is not the norm for most databases. Most transactions don't conflict with each other and can proceed without waiting. Without contention, CockroachDB can serve thousands of reads or writes per second per node (and depending on access patterns, can scale linearly with the number of nodes).

And of course, we're continually working on performance (it's our main focus as we work towards 1.0), and things will get better on both the high- and low-contention scenarios. Just today we're about to land a major improvement for many high-contention tests, although it didn't speed up the jepsen tests as much as we had hoped.

bdarnell··on Introducing Cloud Spanner, a Global Database Service
(CockroachDB CTO here) We haven't implemented everything in the standard yet (Nor will we by 1.0 - there's a lot of stuff there!), but we are aiming to ultimately be compliant with the SQL standard. For example, when we introduced "time travel queries" (https://www.cockroachlabs.com/blog/time-travel-queries-selec...) we adopted the SQL-standard syntax "AS OF SYSTEM TIME" (as opposed to the non-standard out-of-band parameter used in Cloud Spanner)
bdarnell··on Introducing Cloud Spanner, a Global Database Service
(Cockroach Labs CTO here)

Google launching Spanner is generally a positive thing for our industry and our product. It's more proof that what we're aiming for is possible and that there's demand for it. We expect that in five years, all tech companies will be deploying technology like ours.

One of the big differences is that Spanner only uses SQL for read-only operations, with a custom API for writes. We use standard SQL for both reads and writes, which means we also work with major ORMs like GORM, SQLAlchemy, and Hibernate (docs should be live today or tomorrow). Spanner's custom write API will make it difficult to work with existing frameworks, or to convert an existing application to Spanner.

Cloud Spanner only works on Google Cloud and is a black-box managed service. CockroachDB is open source and can be run on-prem or in any cloud on commodity hardware. (We don't offer CockroachDB as a service yet, but may in the future)

At this point, both products are still in beta and are still missing features like back-up and restore (according to the Quizlet blog post). We plan to launch CockroachDB 1.0 with back-up / restore enabled.

* For anyone wanting to know more about how we make CockroachDB work without TrueTime, see our blog post: https://www.cockroachlabs.com/blog/living-without-atomic-clo...

bdarnell··on SQLITE: JSON1 Extension
The `sqllogictest` suite (https://www.sqlite.org/sqllogictest/doc/trunk/about.wiki) is a great resource that is not SQLite-specific. It's a huge collection of queries and expected responses (in an easily-parseable format) that mostly sticks to the SQL standard so it can be run against any database backend. We're using it in CockroachDB.
bdarnell··on How We’re Building a Business to Last
(Cockroach Labs CTO here) This could have been worded more clearly in the post. There will be two versions of backup functionality: a basic implementation for free (Apache license) and a faster distributed and incremental implementation as a paid feature (CCL). It's like the difference between mysqldump and an InnoDB-aware backup tool.
bdarnell··on How We’re Building a Business to Last
(Cockroach Labs CTO here) We don't really anticipate making a lot of changes like this, but yes, it's possible that as the product and market evolve we may change our minds and relicense some CCL features as APL. Of course, we wouldn't move in the other direction - once something has been released under the Apache license it will stay that way.
bdarnell··on Show HN: Use Ansible to Run a “friends and Family” OpenVPN Server on Digital Ocean
From another subthread, https://github.com/trailofbits/algo is an ipsec-only alternative to streisand that looks good (although it requires an app to be installed on android). They also get ipsec working on aws/gce, so apparently whatever obstacle streisand faces with this configuration is solvable.
bdarnell··on Show HN: Use Ansible to Run a “friends and Family” OpenVPN Server on Digital Ocean
I've been using streisand for a while (from China) and it's great; the main reason I can see that you might want to use the linked project instead is that it has a lot less surface area so it could be more secure (it's a lot easier to harden/audit openvpn alone than all the services streisand includes).

That said, if I was going to pick one protocol out of the selection offered by streisand, it would be l2tp/ipsec instead of openvpn (assuming you're hosting on digital ocean. The networks on AWS and GCE are more restrictive and you can't serve l2tp/ipsec from there). I found this to be easier to set up (the client-side software is already included in most operating systems) and to have the best performance.

bdarnell··on CockroachDB Stability Post-Mortem: From 1 Node to 100 Nodes
(Cockroach Labs co-founder here) The source files contain tab characters (gofmt enforces this for all go projects). We configure our editors to render tabs as two spaces. There's no deep reason for this; we've just gotten used to two-space indents from our time at Google and other projects that used this convention.

Since Go code is formatted with tabs, you can mostly get away with setting the indentation to whatever you want. The one practical problem with letting people choose their own values for the width of a tab is that it becomes tricky to enforce a uniform line length, so we've standardized on two-space indents (and 100-char line lengths) across the project.

bdarnell··on Goodbye Mac OS Forge, hello GitHub
The fact that there is a super-specialized tool to do this demonstrates that it's not a pointless feature. It's much nicer to be able to use macports to install all the versions of python that I need instead of using brew to install pyenv and then pyenv to install python (and then pip to install python packages...)
bdarnell··on Goodbye Mac OS Forge, hello GitHub
I still use macports, although it's mainly out of habit and I'll probably switch to brew on my next mac (I've also been trying out nix, and while I really like the idea, package availability and freshness hasn't been as good as macports or brew. Nix hasn't picked up go 1.7, for example).

My main concrete reason to prefer macports is that it has separate packages for every version of python so I can have them all installed at once for testing. I also like a lot of macports' technical decisions, like the fact that my regular user doesn't have write access to the installation directory. And macports has binary distributions now so the awful rebuild-the-world updates aren't as much of a problem anymore.

bdarnell··on Network protocols, sans I/O
Python has a standard interface just like Go's `io.{Read,Writ}er`: the `send` and `recv` methods as defined by `socket.socket`. The appeal of "sans I/O" protocol implementations is primarily for asynchronous I/O, which requires a completely different style of interface. The clever thing about Go is not the standardization of a reader/writer interface, but the fact that the goroutine I/O scheduler makes these synchronous interfaces nearly as efficient as asynchronous ones, so there is no need to introduce a separate family of asynchronous interfaces.
bdarnell··on Curl and http/2 on Mac
Or for macports: `port install curl +http2`
bdarnell··on Nigerians Dominate Scrabble Tournaments Using Five-Letter Word Strategy
Back in the day (2001) I worked on the PalmOS version of Scrabble. One thing that surprised me when I was tuning the AI was that it played better when all words of seven or more letters were removed from its vocabulary. Apparently searching through the longer words was a worse use of time than searching for other placements of shorter words (except on the highest difficulty, where it was given a larger time budget).

We ended up giving the lower difficulty levels a restricted dictionary anyway, but this was to affect players' perception of the difficulty rather than the actual difficulty (the AI on "beginner" mode shouldn't be playing a lot of words you've never heard of). We adjusted time budgets so that difficulty still ramped up as the dictionary expanded at higher levels.

bdarnell··on Import C++ files directly from Python
If you like cross-language import hacks like this then you might be interested in my codegenloader package, which lets you import .proto or .thrift files directly (and it's extensible for other code generators). Mine is a little less magical since it needs to be declared in the package where it will be used, instead of affecting all imports everywhere.

https://github.com/bdarnell/codegenloader

← PreviousPage 2 of 5Next →