- "On Writing Well" by Zinsser
7,398 karma · joined March 16, 2011
- "On Writing Well" by Zinsser
FWIW, Keebo (https://keebo.ai/) tries to solve this problem & reduce your Snowflake bill by using Data Learning techniques. It can be configured to return exact results or approximate results.
We're looking for software engineers on the database development team which is responsible for development of the Citus extension and related tools, and providing custom solutions for customers. Programming is done primarily in C, but without all the usual messiness thanks to Postgres' elegant internal APIs. We also build tools to help customers in whichever language is appropriate. If you're interested in working on distributed SQL, high availability, distributed transactions, seamless scale out, and other parts of a distributed database, and you're excited about working in a distributed organisations with engineers from companies like Amazon/Heroku/Google/Uber then Citus might be the place for you.
To see our other positions visit: https://www.citusdata.com/jobs/
Apply by sending your resume to hadi@citusdata.com or imagine@citusdata.com.
If my data were at order of 10GBs, I would choose Cloud SQL, RDS, etc. At order of 100GBs, I would try both Cloud SQL, RDS, etc. and Citus, etc. to see which one fits my usecase. At order of terabytes, I would choose Citus or some other distributed database.
(I'm a former Citus employee and Current Googler in a non-Cloud SQL team)
1. I think from the 'storage' point of view, a column-store will use much less space than a collection of key/value PgSQL tables:
A) Each PgSQL table comes with several system columns, so each key/value pair will have the large overhead of (key-size + system column size), which will most likely be larger than the actual size of data. This will defeat one of the goals of columnar stores which is to load less data in-memory (to achieve less disk I/O, assuming more frequent columns will get eventually cached), unless you have so many columns and use only very few of them.
B) PostgreSQL doesn't compress each column. Column-stores usually use some kind of light-weight compression to use less disk. When I worked on cstore_fdw, this was one of the desirable features that attracted users.
2. I haven't done PostgreSQL benchmarking in last ~10 months, but unless selectivity of your query is high, fastest of joins won't be even close to a sequential table scan. If the selectivity is low and your query only needs to scan small sub-set of rows, then joins+indexes should be faster, but then again you have overhead of indexes.
One method I found useful in PostgreSQL is to create an index of columns of very frequent queries, and then tune the system to use index-only-scan for these queries.
[1] https://www.edx.org/course/how-win-coding-competitions-secre...
result = init
for value in values:
result = func(result, value)
return result
So, "foldl add_endpoints [] bs" will translate to: result = []
for (x1, h, x2) in bs:
result = add_endpoints(result, (x1, h, x2))
return result
If you expand the 3rd line, you get: result = []
for (x1, h, x2) in bs:
result = result ++ [(x1, height x1), (x2, height x2)]
return result
where "++" is the list concatenation, and "height x" is a function which finds the skyline height by finding the tallest building with "x1 <= x && x < x2".I think the 2nd solution should be easy to understand if you understand the 1st solution. Probably the only new thing is that I've sorted the buildings by height before passing them to "foldl".
In the 3rd solution, the following lines:
skyline bs = (skyline (take n bs), 0) `merge` (skyline (drop n bs), 0)
where n = (length bs) `div` 2
mean: first_half = first_n_elements(bs, length(bs) / 2)
second_half = remove_first_n_elements(bs, length(bs) / 2)
result = merge ((skyline(first_half), 0), (skyline(second_half), 0))
You may wonder what are those 0s? In the merge function I need to keep track of current height of each half, and the initial height of left and right skylines are 0.Then, in merge function:
merge ([], _) (ys, _) = ys
merge (xs, _) ([], _) = xs
means return the other list if any of the lists become empty. The underscores mean a variable whose value is not important for us. We don't care about the current height values here, so I've put _'s instead of real names.Then the other cases:
merge ((x, xh):xs, xh_p) ((y, yh):ys, yh_p)
| x > y = merge ((y, yh):ys, yh_p) ((x, xh):xs, xh_p)
| x == y = (x, max xh yh) : merge (xs, xh) (ys, yh)
| max xh_p yh_p /= max xh yh_p = (x, max xh yh_p) : merge (xs, xh) ((y, yh):ys, yh_p)
| otherwise = merge (xs, xh) ((y, yh):ys, yh_p)
First case, "x > y" simply swaps the two args. This ensures that in the following cases we have x <= y.Second case is probably easy to understand.
In third case, we know that x < y. Just before reaching x, skyline has height "max xh_p yh_p". When we reach x, height of skyline changes to "max xh yh_p". If these values are not equal, we have a height change. So we a construct a new list with head "(x, new height)" and the result of merging the rest of skylines.
If the height doesn't change, we just ignore the change at x and continue with the rest of skylines.
(One of them started practicing 10 years ago [1]. The two other are active at topcoder since 2010 [2] and 2011 [3]).
[1] http://en.wikipedia.org/wiki/Gennady_Korotkevich
[2] http://community.topcoder.com/tc?module=MemberProfile&cr=229...
[3] http://community.topcoder.com/tc?module=MemberProfile&cr=230...
https://us.pycon.org/2014/events/letslearnpython/
http://therealkatie.net/blog/2015/feb/17/young-coders-why-tw...
http://pycon.blogspot.com.tr/2013/03/how-kids-stole-show-you...
* A Demonstration of Agda (https://www.youtube.com/watch?v=8WFMK0hv8bE)
* Agda Tutorial (http://people.inf.elte.hu/divip/AgdaTutorial/Index.html)
I stopped participating in programming competitions when I was about 24. But now I am 27 and I have started doing them again. I am not as strong as I was (or even close to it), but I am aiming to become strong again. One of my goals is to qualify to the GCJ finals next year (this year I was in Round 3, but couldn't make it to finals).
Joy of solving problems is the greatest joy I've ever experienced.