HNHacker News
TopNewBestAskShowJobs

usman-m

80 karma · joined May 12, 2015

[ my public key: https://keybase.io/usmanm; my proof: https://keybase.io/usmanm/sigs/otABJdz5-k0w9X8_CBU9PKNK0Foan2bxzcMulmDmyJU ]
submissionscomments
usman-m··on Kinase: A Framework for Building Web Scrapers in Chrome
Curious how you guys use contexts?
usman-m··on ContinuouSQL Triggers
Almost. Currently continuous views that contain an ORDER BY or LIMIT clause are not supported.
usman-m··on ContinuouSQL Triggers
They work better for non-sliding window queries. Triggers on sliding window queries are much more resource intensive, both for CPU and memory. Essentially, the trigger process has to keep track of tuples for each step in the window and combine them whenever any tuple is updated to get the new value.
usman-m··on Orchestra: An open-source system to orchestrate teams of experts and machines
Great stuff! For now it seems like this leaves out the job of finding crowd workers. Could it be possible to do something like designing microtasks in MTurk and "stitching" the workflow together using Orchestra?
usman-m··on Making Postgres Bloom
Our native implementations of all probabilistic data structures use MurmurHash3, so this isn't a problem. The dumbloom implementation is in no way a good Bloom filter, as the name suggests :)
usman-m··on Making Postgres Bloom
It probably uses a HyperLogLog--the 2% error rate kind of gives it away. Bloom filters approximate set membership queries, HyperLogLogs approximate set cardinality queries. COUNT DISTINCT is a set cardinality query.

We actually support a HyperLogLog backed COUNT DISTINCT aggregate too: http://docs.pipelinedb.com/aggregates.html#general-aggregate...

usman-m··on Making Postgres Bloom
Oops, it was meant to be:

  SELECT <user_id> FROM (SELECT DISTINCT user_id FROM user_actions);
You're absolutely right that both those queries will give the same result. I guess I was trying to motivate the basic problem of finding whether some user exists in a set of users, and `SELECT DISTINCT` is the SQL way of representing a set.

Fixed the post, thanks!

usman-m··on Making Postgres Bloom
ahachete, I'm not sure if I totally understand your question.

Continuous views are consumers for streams. You can think of them as high throughput real-time materialized views. The source of data for the stream can be practically anything. Logical decoding on the other hand is a producer of streaming data--it's basically a human readable replication log. So you could potentially stream the logically decoded log into PipelineDB and build some continuous views in front of it.

usman-m··on Making Postgres Bloom
Mostly that. I've also been thinking about how we could incorporate some machine learning algorithms, like online perceptrons.
usman-m··on To Be Continuous
Not right now, but it's definitely a feature we're thinking about.

Awesome--let us know what you think about it!

usman-m··on To Be Continuous
It's not so much as how different it is from an architectural standpoint as it is about the sheer magnitude of such a feature. All open-source communities have processes which help maintain high quality, but also add a bit of red tape. I don't think we'd be able to operate at the pace we want if we were pushing every change upstream.

However, we love Postgres and plan on actively merging upstream releases!