HNHacker News
TopNewBestAskShowJobs

jaffee

263 karma · joined January 26, 2014

[ my public key: https://keybase.io/jaffee; my proof: https://keybase.io/jaffee/sigs/m6DntPZSNkB2cmrKeaITaqkrHBbXJpgYP3QX_fQWS6I ] Twitter: @mattjaffee
submissionscomments
jaffee··on Pilosa: An open source, distributed index
Good catch... that sounds pretty silly. It should probably read more like "converting relationships to be represented by single bits"

As a concrete example, we took the NYC taxi ride data set which is something like 300GB of CSV files and when it was indexed in Pilosa, the total size of all the bitmap files was closer to 40GB.

jaffee··on Pilosa: An open source, distributed index
Source code:

https://github.com/pilosa/pilosa

jaffee··on License Changes for Confluent Platform
You can rant about an "attack on freedom" all you want, but how do you propose that Confluent protect its business so that they can continue paying their engineers to keep working on tools for which they publish all the source code?

The alternatives seem to be 1. they keep all their stuff proprietary or 2. they leave it truly open and AWS takes the majority of their market and they slowly suffocate.

Aren't both of those strictly worse than the path they've taken?

jaffee··on Django Newbie Mistakes
I'm sure they could save a massive amount of money and improve their latencies if they did.
jaffee··on Django Newbie Mistakes
Ha, ok - Django is not difficult to maintain, Django is a very powerful, carefully built framework with tons of features that can get you up and running with a production ready website very quickly.

Python is difficult to maintain. Full stop. Why? Anything in Python can do anything to anything else - there are no boundaries which makes it basically impossible to reason about code at scale. You need very strictly enforced code quality standards across your entire codebase and libraries in order to be able to trust anything - otherwise you don't know if some method on a class has been swapped out right underneath your feet, or if some field access is actually calling a function which is accessing a database.

Go is super easy to reason about compared to Python - as long as no one is importing unsafe or reflect (easy to check), you have very solid guarantees as to what can happen at any point.

jaffee··on Django Newbie Mistakes
scalability and speed are two very different beasts. You can build scalable things in Go, or Python, or most any language.

At a certain scale though, the amount you'd save in infrastructure bills by using Go instead of Python is absolutely stupid. A 10x difference would not be surprising in the slightest.

jaffee··on Django Newbie Mistakes
I feel like I just got into a time machine - this is the top post on hacker news? People are learning Django in 2018?

I guess its fine if you have a small project and aren't going to see much traffic, but Python is so slow and difficult to maintian compared to (e.g.) Go. I say this as someone who worked in Python for a long time, and on several large Django projects.

jaffee··on Pain27 keyboard
If you like this, see http://www.40percent.club/

In particular: http://www.40percent.club/2016/11/gherkin.html

jaffee··on Notes on structured concurrency, or: Go statement considered harmful
I LOVE Go's concurrency model, but this article has won me over (pending some experimentation anyway).

If you just skimmed, this is actually worth a careful read. The parallels between "go" and "goto" are explained very clearly, and you get some awesome Dijkstra quotes to boot!

jaffee··on How should you build a high-performance column store for the 2020s?
Pilosa, https://github.com/pilosa/pilosa which is mentioned, is actually open source, and a relatively readable Go codebase if anyone is interested in what "an entire data engine on top of bit-vectors" looks like.
jaffee··on Pilosa Performance – Billion Taxi Ride Dataset
1. It is written in Go, and open source! https://github.com/pilosa/pilosa

2. It's going to depend on your use case, but generally you will be streaming writes into your current data store as well as Pilosa. Then read queries will be serviced by Pilosa, or by a combination of Pilosa and your system of record depending on the data you need.

We're more than happy to help explore how you might use Pilosa - get in touch with us on github, or through https://www.pilosa.com/about/#contact

jaffee··on Pilosa: open source, distributed bitmap index in Go
The use case for Pilosa is usually as an addition to something like Cassandra or HDFS where you have terabytes (or more) of data that you want to be able to query more flexibly - especially if there are a very large number of attributes that you want to filter and segment on (think tens of millions).

I think you're right though, that there are use cases which would benefit from a library exposing this functionality - you need to have quite a lot of data before compressed bitmaps representing the relationships in that data start overflowing memory on a single machine.

jaffee··on Pilosa: open source, distributed bitmap index in Go
I'll add to travisturner's response - Pilosa has a few bells and whistles beyond being a straight up bitmap index including:

- associating each bit with a timestamp (at various granularities) and queries over time ranges.

- adding arbitrary key/value metadata to each row or column

- automatic sorting/caching of bitmaps to support "TopN" queries

jaffee··on Pilosa: open source, distributed bitmap index in Go
User segmentation was actually the original use case that prompted Pilosa's development! Segmenting hundreds of millions of users with tens of millions of potential attributes.
jaffee··on Pilosa: open source, distributed bitmap index in Go
Good call - wikipedia has a nice overview: https://en.wikipedia.org/wiki/Bitmap_index

The key addition with Pilosa is that it's distributed and can scale horizontally :)

jaffee··on Go compiler: Initial support for concurrent backend compilation
On github, it'd be called a "pull request".
jaffee··on SpaceX Falcon rocket explodes on landing after delivering satellite to space
I want to know what exploded - I wouldn't have expected it to have very much leftover fuel at that point...
jaffee··on Ask HN: How to learn OpenCL
If you're a complete beginner in data parallel programming, and you're having trouble finding good intro material for OpenCL, it might almost be worthwhile to check out CUDA instead. In terms of the programming model, OpenCL and CUDA are identical - significant differences don't come about until you start optimizing for specific devices.

I learned CUDA first on my own, and then took an OpenCL class and found that the whole first section was completely redundant. There's also a pretty great wealth of CUDA material online and a few published books if that's your sort of thing.

← PreviousPage 2 of 2