As a concrete example, we took the NYC taxi ride data set which is something like 300GB of CSV files and when it was indexed in Pilosa, the total size of all the bitmap files was closer to 40GB.
263 karma · joined January 26, 2014
As a concrete example, we took the NYC taxi ride data set which is something like 300GB of CSV files and when it was indexed in Pilosa, the total size of all the bitmap files was closer to 40GB.
The alternatives seem to be 1. they keep all their stuff proprietary or 2. they leave it truly open and AWS takes the majority of their market and they slowly suffocate.
Aren't both of those strictly worse than the path they've taken?
Python is difficult to maintain. Full stop. Why? Anything in Python can do anything to anything else - there are no boundaries which makes it basically impossible to reason about code at scale. You need very strictly enforced code quality standards across your entire codebase and libraries in order to be able to trust anything - otherwise you don't know if some method on a class has been swapped out right underneath your feet, or if some field access is actually calling a function which is accessing a database.
Go is super easy to reason about compared to Python - as long as no one is importing unsafe or reflect (easy to check), you have very solid guarantees as to what can happen at any point.
At a certain scale though, the amount you'd save in infrastructure bills by using Go instead of Python is absolutely stupid. A 10x difference would not be surprising in the slightest.
I guess its fine if you have a small project and aren't going to see much traffic, but Python is so slow and difficult to maintian compared to (e.g.) Go. I say this as someone who worked in Python for a long time, and on several large Django projects.
In particular: http://www.40percent.club/2016/11/gherkin.html
If you just skimmed, this is actually worth a careful read. The parallels between "go" and "goto" are explained very clearly, and you get some awesome Dijkstra quotes to boot!
2. It's going to depend on your use case, but generally you will be streaming writes into your current data store as well as Pilosa. Then read queries will be serviced by Pilosa, or by a combination of Pilosa and your system of record depending on the data you need.
We're more than happy to help explore how you might use Pilosa - get in touch with us on github, or through https://www.pilosa.com/about/#contact
I think you're right though, that there are use cases which would benefit from a library exposing this functionality - you need to have quite a lot of data before compressed bitmaps representing the relationships in that data start overflowing memory on a single machine.
- associating each bit with a timestamp (at various granularities) and queries over time ranges.
- adding arbitrary key/value metadata to each row or column
- automatic sorting/caching of bitmaps to support "TopN" queries
The key addition with Pilosa is that it's distributed and can scale horizontally :)
I learned CUDA first on my own, and then took an OpenCL class and found that the whole first section was completely redundant. There's also a pretty great wealth of CUDA material online and a few published books if that's your sort of thing.