HNHacker News
TopNewBestAskShowJobs

jdf

108 karma · joined November 13, 2010

submissionscomments
jdf··on Fast columnar JSON decoding with arrow-rs
It would be great if someone could implement the schema discovery algorithm from the DB research GOAT, Thomas Neumann, and add it to Apache Arrow: https://db.in.tum.de/~durner/papers/json-tiles-sigmod21.pdf
jdf··on Writing parsers like it is 2017
There's a Rust parser generator called LALRPOP that is apparently inspired by Menhir.

https://github.com/nikomatsakis/lalrpop http://smallcultfollowing.com/babysteps/blog/2016/03/02/nice...

I've never used Menhir so I can't compare how similar they are in practice, but I've enjoyed the times I played with LALRPOP much more than the many times I've battled various yacc derivatives.

jdf··on Google Cloud Platform is the first cloud provider to offer Intel Skylake
Not asking you to reveal any internal plans, but do you think there's any chance that some form of KNL could show up on GCP in the near future?

FWIW we're currently using GCP, generally love it, and I'm looking forward to trying out Skylake...

jdf··on How Does the SQL Server Query Optimizer Work?
There have been a lot of tuning systems that take a workload and try to optimize things like index structures. They are tools that get your database "tuned" rather than "self-tuning".

However, I think this (very recent) research paper is a much better peek at how self-tuning might work:

https://blog.acolyer.org/2017/01/17/self-driving-database-ma...

jdf··on Announcing Tokio 0.1
> The way in which we are safer is memory safety, nothing more.

I know you want to not overstate rust's claims given recent articles, but I think you're actually underselling a little here. For example, a rust `enum` make it much easier for the compiler to enforce code correctness. It's hard to go back to similar code in C or Go once you've gotten used to `match`.

jdf··on 2017 Rust Roadmap
Yes, this is a great point! I was a bit sloppy in my wording above.
jdf··on 2017 Rust Roadmap
> The only thing I can think of that Rust has that Go doesn't is, like, SIMD, and I'm sure your comment wasn't just referring to SIMD.

Just to clarify, are you referring to Rust using SIMD by way of LLVM, or by way of being able to use SIMD primitives / intrinsics directly in Rust code?

The former works much better than I had anticipated. I've been surprised by the extent that my iterator code ends up vectorized without me doing much work.

The latter does not give me warm Rust feelings today. There's a SIMD crate, but it doesn't look maintained and only works with the nightly compiler releases. I didn't think there was any stable way to do inline assembly, so I think linking C is my best bet here?

jdf··on 2017 Rust Roadmap
I forgot to add - I partially agree with you about the productivity of GC. For non-performance sensitive code, GC simplifies a lot. Otherwise, I find Go much nicer to reason about than Java since it has real arrays (value types!).

But I find it a mixed bag of whether the naive Go version of something that shares memory by GC is simpler than the naive Rust version of something that shares memory. Sometimes ref-counting (Rc<T> in Rust) is fine, although that's more expensive than GC. Sometimes Rust's ownership model nudges you to make the code much simpler and makes it clear that something only has a single writer. Sometimes you wish you were in C and just did it yourself...

jdf··on 2017 Rust Roadmap
I agree with some of your points, but think this is phrased a bit harsh. FWIW, I write both Go and Rust on a regular basis.

Here's where I would agree with you:

- Go makes it harder for someone to write overly abstract code (a common affliction!).

- Being able to occasionally do type assertions in an ergonomic way is surprisingly nice.

- I wish Rust had something in the stdlib like net/http.

- I like that go fmt is so unconfigurable and canonical.

Here's where I would disagree:

- I find ADTs (Rust's enums) super helpful for productivity.

- Removing nil pointer derefs is wonderful, particularly for refactoring.

- I spend too much time in Go rewriting bits of code that I would just use generics for in Rust or C++. Rust's iterators are wonderful and I end up using them over and over again.

- Maybe it's my C background, but I like being able to occasionally use macros. Even for tests it makes things much more readable.

- The borrow checker ends up moving many concurrency issues from runtime debugging to compile time debugging.

- I think cargo is more pleasant to use to manage code than using go + godeps/glide/etc.

jdf··on Unicorn: a simple and flexible abstraction of BigTable-like databases
Not to be a stickler about name collisions, but Facebook wrote a research paper about a graph database called Unicorn back in 2013:

https://people.csail.mit.edu/matei/courses/2015/6.S897/readi...

This appears to be unrelated, which is somewhat unfortunate.

jdf··on Turning the database inside-out with Apache Samza
MVCC doesn't necessarily mean no in-place updates, it just means that you can distinguish between multiple versions. For example, Oracle:

- keep most recent version of all keys in B-tree

- store updates in undo log ("rollback segments")

- queries for older versions dynamically undo recent changes

http://docs.oracle.com/cd/B19306_01/server.102/b14220/consis...

jdf··on After Docker: Unikernels and Immutable Infrastructure
If you are using a more minimal hypervisor (see my other comment on the parent), then there do seem to be some measurable gains. I've seen a few papers in this style:

https://www.usenix.org/conference/osdi14/technical-sessions/...

We describe the hardware and software changes needed to take advantage of this new abstraction, and we illustrate its power by showing improvements of 2-5 in latency and 9 in throughput for a popular persistent NoSQL store relative to a well-tuned Linux implementation.

That said, a simple application like memcached might be currently latency-bound by the kernel's network stack, but a more complex application that reads from disk (even SSD) won't be.

jdf··on After Docker: Unikernels and Immutable Infrastructure
If you're running a service, then you can use a much trimmer hypervisor, e.g.

https://github.com/siemens/jailhouse

Since the guest unikernel isn't a full kernel, the hypervisor interface is much more minimal, and the few host features it needs can be delegated to the CPU via VT-X (e.g. page table mapping).

At least, that's the dream. (I've never actually used Jailhouse or tried any of the research projects attempting this.)

jdf··on Pandas 0.15 has been released
Just FYI, when 'sudo pip install pandas' doesn't work (which it didn't for me recently), you'll get no love upstream:

https://github.com/pydata/pandas/issues/7517

As @mynegation notes, you can use Anaconda (or virtualenv).

jdf··on Ask HN: Do you still use an RSS reader?
I also use Bazqux, and endorse every point bonaldi made. It also works pretty smoothly on iOS with Feeddler.

Despite shifting a lot of article tracking to Twitter, I still find RSS to be a better way to track and consume long form content. Also, the signal to noise ratio of the average RSS feed is much better than that of the average Twitter feed for people whose article's I'd like to read.

jdf··on Revisiting 1M Writes per second
It seems like this info should be front in center in the test. 1M/s 100kb writes is much more impressive than 1M/s 16 byte writes.

That said, there's a previous benchmark linked to at the top of the post:

http://techblog.netflix.com/2011/11/benchmarking-cassandra-s...

The client is writing 10 columns per row key, row key randomly chosen from 27 million ids, each column has a key and 10 bytes of data. The total on disk size for each write including all overhead is about 400 bytes.

There are 3 replicas, so figure that in as well.

jdf··on Ur/web: pure functional, statically typed web programming
I've been using BazQux since the Google Reader close and have been pretty happy with it. I don't think I've noticed a single issue with BazQux so far.

The web interface is actually quite snappy, in contrast to a lot of the overly-designed alternatives I tried out. That said, I generally interface with FeeddlerPro on my phone/tablet rather than directly hitting bazqux.com.

jdf··on Symas Lightning Memory-Mapped Database (LMDB)
http://www.consul.io/intro/vs/serf.html
jdf··on Simple Binary Encoding
Good point, although it's not like the stock protobuf implementation (rather than protocol) is the best that can be had. For example, this implementation of protobuf deserialization is quite a bit faster:

https://github.com/haberman/upb

http://blog.reverberate.org/2011/04/upb-status-and-prelimina...

jdf··on The Writings of Leslie Lamport
Lamport's description of his failures in introducing the Paxos algorithm is a highly amusing story:

http://research.microsoft.com/en-us/um/people/lamport/pubs/p...

For those that don't know, Paxos is one of the most important algorithms in distributed systems, so it's amazing to see that it wasn't even published for 8 years due to the author's ... odd structuring of the problem.

jdf··on Dropbox Uploader: Bash script to upload, download, list, delete from Dropbox
Not if you're running a normal Linux distro on ARM - Dropbox doesn't provide a binary there. Which is a bummer since there are a number of nice, simple, and cheap machines you can buy nowadays that come with an ARM chip, e.g. a Raspberry Pi.
jdf··on Google Chromebook Under $300 Defies PC Market With Growth
I've always been a vim/screen/bash user rather than an IDE, so I can't speak to the experience of using Eclipse/IntelliJ on a Chromebook. The fact that I'm doing everything console based certainly lowers the resource needs since the chroot isn't running another windowing system.

I know the JVM isn't often associated with low memory applications, but it seems like it should be possible. As I mentioned in my other reply, the Chromebook has more resources than the average Android phone, so all those Java Android apps should run fine (albeit under Dalvik rather than Sun's JVM).

jdf··on Google Chromebook Under $300 Defies PC Market With Growth
For me, there are three potential types of applications:

a. Those that run fine on the Chromebook.

b. Those that would run fine on a larger laptop or desktop, but not on the Chromebook.

c. Those that need to be deployed to some sort of server.

I don't write anything that falls under (b). If a program is expecting to be run on 10 disks, or with 48gb of RAM, or across a cluster of 12 nodes, it's highly unlikely that my personal dev machine will work out, so (c) will be used. That's also how I'd want to test any production-level deployment.

The gap between (a) and (b) is actually quite small. The Samsung 3 Chromebook specs are a bit better than almost every smartphone, so pretty much any Android app could be placed under (a). Pretty much any unit test for a (c) app fits under (a).

Compiles would be faster with a beefier dev machine, but the fact that the Chromebook comes with an SSD already places it ahead of the default corporate machine (at least in my experience - perhaps employees get SSDs in their Dells nowadays). Certainly a MB Air has a faster CPU, but the difference isn't that dramatic for development purposes.

As one counterexample, I'll point out that my wife spends most of her work day in Photoshop and Illustrator. Even if those apps ran on Linux, my guess is that the resource needs would still make an Air or a MB Pro a much better choice.

jdf··on Google Chromebook Under $300 Defies PC Market With Growth
I feel like the Chromebook is actually a better experience than a laptop.

As a programmer, I honestly only use 2 windows: browser and shell. Running crouton to create chroot Ubuntu "images" I can have my normal full (text-based) dev environment, and alt-tab back and forth with the browser. Honestly it feels easier to use than Ubuntu - I don't use any of the builtin Google services (e.g. Drive), but they manage to stay out of the way.

It's pretty cheap, looks pretty reasonable, and comes in at a pretty low weight. The screen's not the greatest, but I'm not doing anything where that matters. So all in all a great dev machine.

Only caveat for me is that Dropbox doesn't have an installable app for ARM. I need to find something else that does a good job of seamlessly syncing my workspaces and NFS, rsync, etc don't fit as nicely as Dropbox.

jdf··on The most useful thing in bash
Alternatively, you can add ':p' to the third line. This will print out the command rather than executing it. Additionally, it's also added to your bash history, so you can add it by pressing the up arrow.

  $ chmod 755 foo
  $ cd !$:p
  cd foo
  $ cd foo
There are some similar "tricks" here:

https://news.ycombinator.com/item?id=5337558

jdf··on Google’s Dremel Makes Big Data Look Small
Dremel was Google's internal name. Their public API is called BigQuery:

https://developers.google.com/bigquery/

jdf··on Google’s Dremel Makes Big Data Look Small
Not sure why Cloudera is part of this article, seems like all the attention here should be on Google and the BigQuery team.

Here is an open source project similar to Dremel:

http://www.itworld.com/big-datahadoop/290026/new-apache-proj...

jdf··on Snowflake Server
There's a new project called Ansible that may be of interest to you:

http://ansible.github.com/

While that page has a long list of things they do, the important bits relevant to your comment are

1. a tighter focus on idempotence than Fabric 2. an easy-ish way to integrate package management so you could potentially use the same script to kick off either yum or apt depending on the box

jdf··on A standing desk for $22
I have the exact same setup at home, and am also a fan. It works well, is cheaper than a lot of other pre-built solutions (although obviously not as cheap as the original post), and seems to be pretty well made. I'm also much happier with the idea of a hand crank than an electric motor.

I also ended up getting an anti-fatigue mat as well. I thought that other standing desk folks were overrating the mats, but after the first few days standing at my desk (on a hardwood floor, no less) I saw the light.

jdf··on F1 - The Fault-Tolerant Distributed RDBMS Supporting Google's Ad Business
Clustrix relies on an Infiniband cluster interconnect, so the latency is several orders of magnitude smaller. This makes distributed transactions across shards much quicker at the cost of requiring more expensive hardware.

Cassandra's query language, CQL, is not really comparable since it only supports such a small subset of SQL. Also, Cassandra uses eventual consistency in place of doing distributed transactions.

Page 1 of 2Next →