"Here from HackerNews? This was originally posted several months ago. Check back in two weeks for an updated benchmark including newer versions of Hive, Impala, and Shark."
28 karma · joined December 8, 2012
"Here from HackerNews? This was originally posted several months ago. Check back in two weeks for an updated benchmark including newer versions of Hive, Impala, and Shark."
Very useful if you are into distributed systems and would like to see a practical implementation with clear examples.
Infact, one of the solutions proposed for ngRepeat issue (with lot of caveats) has 300+ stars no github
https://github.com/Pasvaz/bindonce
Including me, there are lot of folks who would find your solution very useful. You are addressing a fundamental O(n) scaling problem in AngularJS.
https://chrome.google.com/webstore/detail/hackernew/lgoghlnd...
How about asking BigCos their top pain points once in a while? That would be a 1200x treasure for current/future entrepreneurs.
Does this imply that mac forwarding tables(on switch) and arp cache(on hosts) need to have entries only for their immediate neighbours ?
Curious to know how much modified the host network stack is. Also, how do you provision a new server with the right IP ? Is this mechanism in L2 ?
>There are many other reasons that layer2 wasn't a good choice for us, and that layer3 makes a lot of sense. I'd be happy to discuss more of these as well.
I am sure people would find that very useful. Thanks for the excellent writeup !
On a serious note, take a look at cumulus networks[1] who claim to solve the hw/sw disaggregation problem. They have a linux OS distro which can run on h/w of multiple vendors ( not the popular ones like cisco/jnpr/hp/brcd etc.. since those are closed platforms).
Talking about any disadvantages, I can say that you need someone of Cheriton's ability to really guide you during the initial build-up of the project. His ideas of having the compiler do a lot of checking to avoid bugs in the code later are really cool. But the learning curve is pretty steep. Newer paradigms (since the course was originally designed) like STL, BOOST, and tips in Effective C++/STL etc can replace some of the concepts he espouses. But he is pretty clear that some of them are really inferior. It might be true for some cases, but I have usually used them in production just fine..
http://www.rocksclusters.org/rocks-doc/papers/two-pager/pape...
Rocks is used at production cluster installations with thousands of servers.
http://www.rocksclusters.org/rocks-register/index.php?sortby...
for 10M loop count, method 1 -> 1.599 s, method 6 -> 1.91 s
for 30M loop count,method 1 -> 4.967 s, method 6 -> 5.871 s
Summary: The KISS s1 += s2 always wins
Great to hear that this is built on the awesome boto library. Will serve as an useful reference for boto developers.
BTW, the person asking the last couple of questions is Ed Bugnion, one of the co-founders of VMWare. He is a faculty now at EPFL.
http://www.reuters.com/article/2013/09/03/us-syria-crisis-us...
Think we agree on case ii). On case iii) I still think that having 1 local and N-1 remote might be useful when we want to prioritize a particular writer over others. Borrowing the sales example from commenter jchrisa, consider a new user sales-head (local rule). He syncs the data uploaded by his sales folks (remote rule) and then goes offline. When he is done editing and comes online, he wants to make sure his delta takes precedence irrespective of any previous changes by sales folks during his being offline. Since he has local rule, his update will just win. Further, he does not want sales guys who were offline and come online after him to overwrite his last update immediately. I am assuming that the sales folks with "remote" rule will see the data with newer server version and accept it.
Let's say I have 10 devices d1,d2....d10 making updates to "a" on the server and went offline. a==20 and last update was by d5 before everyone went offline.
When the devices come back up, the fate of "a" depends on the rulesets. Following are 3 possible high-level combinations.
i) All devices have "remote" rule. On reconnection, everyone rollback "a" to 20. They are essentially back to the time before going offline. Even the device which did the last update(d5) before going offline is rolled back too, which seems bit odd. Still simple to reason with..
ii) All devices have "local" rule. On reconnection, the last device to reconnect updates "a". It is then broadcasted to all other devices. Note that it is not the last device to update "a". Rather it is the last to reconnect (Now, even if all of them reconnect at same time, depending on the queueing at server, the one at the tail wins). Not really simple..
iii) Mix of "remote" and "local" Let's say d1 had "local" rule and all others had "remote". On reconnection, d1's "a" will be propagated to everyone. This is irrespective of the order of reconnection (I am assuming that between reconnections "a" is not modified). This is pretty simple and perfectly predictable. Now, if we have more than one "local", we start getting non-deterministic, and at the extreme move to case ii)
Startups are a different story as time to market is crucial and you have almost no slack.
^^^^^^^^
From my experience in bay area & india, I never saw shortage of resumes for open positions. Large enough even if you assume people applying to multiple companies. The shortage though was for talented folks meeting the expectations of mgr/co-workers.That's wildly optimistic. The reality is unemployment rate for indian engineering graduates now ranges from an optimistic 50% upto 80%..
[1] http://www.livemint.com/Industry/HCWB4sLvFBxfIFyNBYtqOP/Degr...
[2] http://articles.economictimes.indiatimes.com/2013-06-18/news...
Let me take a concrete example closely mirroring my experience. There are 4 fixed performance buckets. Top 5%, Next 20%, Next 65%, Bottom 10%. They get hikes of 20%, 8%, 4%, 0% respectively. Again, these are fixed numbers. Let us assume there are 4 people A,B,C,D and out of a hypothetical score of 100, score 95, 90, 87, 85 respectively based on various parameters. You would assume that since D differs in ability with A by 10%, he would get 90% of A's hike. But sorry, due to the stack implementation he gets 0%, while A gets 20% ! Let's say if the scores of A, B, C, D were instead 100, 50, 25, 5, the hikes would have make much more sense.
Summary: Discrete curve of benefits works well only when it closely matches the curve of people productivity. This is rare. So it just ends up being unfair and creates an unhealthy rat race.
For large web programs ( html/js etc), once you use Sublime, no going back to VI. The plugin ecosystem for js fwks are awesome.
Second that. I have had videos which couldn't be played on VLC (codecs or missing frames) play nicely with mplayer. Easily the best versatile player, especially if you are conversant with the command line options ( or take a quick look at man page).