592 karma · joined July 28, 2011
Anyone can post as me, but you have to know the name of that guy that I lost an argument to that one time. I just ended up rambling some nonsense about the afterlife at the end.
Do play by the rules, though.
It doesn't seem uncommon for very young (25+) top engineers in the Bay Area to be making anywhere from $175k/year -- 250k/year in base salary + bonus + equity. Over 30 years, that's a more or less guaranteed $5m -- 7.5m before taxes from a single income. (Ignoring things like promotions and older engineers possibly needing to retrain.)
If you're not an early employee at the next Google, or a co-founder of a to-be-talent-acquired company, is it likely at all that you will beat those numbers at a string of startups? And if you do beat those numbers at a startup, is it because you're basically doing the same work you would be doing at Google anyway, likely ads or infrastructure?
I guess it seems like "make a business version of X" is always sensible, so maybe this is like the business version of http://about.me/ ?
Personally, the resume process whenever I have applied anywhere has been: (1) make a beautiful resume in LaTeX, (2) export beautiful LaTeX resume to PDF, (3) get asked for an MS Word Doc or plain text so they can put it in their recruitment system.
Does passively posting a resume work for people? Does it lead to actual good recruiter interactions, rather than noise? Or is this for freelancers?
However, for what it's worth, the two best resources I've seen for systems design are:
Hints for Computer System Design: http://cseweb.ucsd.edu/classes/wi08/cse221/papers/lampson83....
How to Design a Good API: http://www.youtube.com/watch?v=aAb7hSCtvGw
What else should be in this bibliography?
Silicon Valley seems like a good multiplier on execution, as well as a place where it is easier (and more expected) to increase the original product vision. VCs insist on the team trying to take over a $100B+ market, many top infrastructure engineers, mobile developers, UX people, and so on won't stay if they don't think the equity will be worth anything, and there is a huge amount of highly mobile web talent. More talent breeds more ideas, expanding the concept and improving the execution.
Could Facebook have scaled their engineering team to handle 2x-5x the number of users every year (with maybe 1m users on day one of the company), with a huge portion of those users viewing the website every day, without being in a place with (a) tons of web talent and (b) a culture where leaving an established job to work at a web startup was reasonable?
In terms of upper level management, half of Facebook's C-level executives appear to be ex-Google. Could Facebook have gotten that level of operational and advertising expertise in Boston?
To answer your question directly, I suspect there are huge recruiting, operational, and business development challenges to running a large scale consumer web startup outside of Silicon Valley. That said, it is hard to know how that would impact any particular company, or any particular space (e.g., location-based services).
(Groupon obviously had to scale sales, but that's not what I mean here.)
Here's tumblr.com vs blogspot.com on Google Trends:
http://trends.google.com/websites?q=tumblr.com%2Cblogspot.co...
Here's foursquare.com vs facebook.com on Google Trends:
http://trends.google.com/websites?q=facebook.com%2Cfoursquar...
Here's etsy.com vs ebay.com (from another comment):
http://trends.google.com/websites?q=ebay.com%2Cetsy.com&...
Here's the other four you list vs facebook.com:
http://trends.google.com/websites?q=facebook.com%2Cgroupon.c...
I can name almost no consumer web startups from the past decade that weren't built in Silicon Valley. Stack Overflow? MySpace? Some fashion startups in NYC? Anything else?
For most of its early life (and perhaps still), Facebook had less than one engineer per million users. Is that kind of talent available outside of the Valley? How many people that have worked at the scale of Google/Facebook/Yahoo!/Twitter (with all that it entails in terms of coding standards, infrastructure, and so on) exist outside of the Valley? Maybe the Seattle area has enough ex-Amazon/ex-Microsoft people to launch a big consumer web startup, but even there, a substantial part of Bing seems to be based in Silicon Valley.
I guess what I'm wondering is: is OSX/iOS versus Linux/Android the operating system flame war of our time? In this thread, I see almost entirely ad hominem attacks on Stallman. (Not unlike a Daring Fireball thread in reverse.) Is he Glenn Beck or is he Rush Limbaugh? Is he just slightly impractical or is he extremely impractical?
What I don't see is any substantial discussion of the OP. At the moment, there are 2/72 comments that mention the iOS App Store at all, which is a major point that Stallman seems to be making here.
More generally, is no one worried about what Stallman is worried about? Jobs' explicit stated goal was to destroy the only Linux mobile platform (through patents), and his policies on the App Store and elsewhere were explicitly hostile to Free Software. Had Jobs been successful and perhaps lived another ten years, we might be living in a world with no Free Software on our personal devices at all. Apple laptops with software installed from the Mac App Store to develop iOS apps for the iOS App Store (with both App Stores of course rejecting GPL code).
Given that almost all startups rely heavily on Free Software, from emacs to gcc to Xen to Ruby to every poorly constructed library on Github, that would be tremendously negative for the startup ecosystem, no?
For example, one very common performance optimization in web applications is batching writes. Why not send the very compact results of a few A/B tests to the database/store once every ten requests, rather than once every request?
However, GAE doesn't really provide a way to do this. You could just store the results in local memory and then persist to the datastore every 10 requests, but there is no guarantee your instance won't get shut down. You could also store to memcache, and do what the author suggests, persisting from memcache to the datastore periodically. But my understanding is that there aren't any guarantees as to what the memcache size or revocation policy is, so it's not really possible to guarantee that all data gets persisted. (Losing random data intended to be persisted may be fine for A/B testing, but probably not for many other things.)
Likewise, it would be nice to delay datastore puts as the author does, or put them in a task queue, but both of those seem to take as long as doing a datastore put in the first place.
Is this analysis wrong? Does Google's memcache implementation provide some guarantees that mean that your data won't get removed as a result of your or someone else's actions without you knowing? Is there some reasonable pattern for storing persistent data on GAE but not having to force the user to wait?
Having read Matt's account of his actions, I agree with your reading. It feels like he came to believe academia was bullshit, but didn't want to hurt his close friends in academia. (Which is fine! You are using a throwaway account, no?)
Over time, I came to view CS academia as bullshit as well, but I am starting to reverse this view. Academia is inefficient, anachronistic, political, bureaucratic, and doesn't do a very good job measuring its own output. However, that seems to apply to countless other institutions, especially institutions with billions of dollars.
Is it possible that academia is a flawed system that sometimes has good results (e.g., new knowledge, startups, quality engineers) that would not have happened otherwise? Or have new inventions like incubators or online education arrived that make it easier to achieve those good results without the old, flawed system?
In other words, it would be interesting to see the differences between Akzidenz-Grotesk, Arial, Helvetica, Helvetica Neue, Roboto, Univers, and Vera. Something like:
http://en.wikipedia.org/wiki/File:Helvarial.svg
My guess is that other than typographers, people mostly cannot tell any of the Helvetica-ish fonts apart, and there are much bigger differences between weights and such.
That said, what really came out of left field was the OP saying that he and his friends are more and more using Go.
Is Go adoption happening? I'd love to have a systems programming language that isn't C or C++, but I'd always assumed that Go was dead-on-arrival specifically because it wasn't C or C++.
1. Ruby/Rails developers use Ruby/Rails even though it is slow because storage access is slower (and thus the bottleneck).
2. Storage access is about to get much faster as people switch from spinning disks to SSDs.
3. Therefore, people will start using faster performing languages because app server language will become the bottleneck.
However, there are two questions I have about this argument:
1. How much faster are SSDs in common database scenarios? Presumably, they are much faster for point queries? But is it 2x, 10x, or 100x? How much faster are they per dollar (e.g., how do SSDs compare to tons of spinning disks in RAID)? Will they impact common caching scenarios (e.g., memcache)?
2. Are app servers ever the bottleneck? Modern web development seems designed for stateless, horizontal scaling ("scale out") at the app server layer. Further, this stateless, horizontal scale out requires no effort for even the smallest shops, through things like Heroku, App Engine, EC2, and similar services. Will it ever make sense to trade off expensive programmer time (by coding in a lower level language) for less app servers at all but the most extremely popular websites? Is there some compelling reason the app server layer should not be stateless?
But it seems a little fishy that the OP got a -50 penalty from Google due to SEO, and now the OP is trying to get a blog post featuring one anecdotal data point about Google's results being unfair to the top of HN, no? I mean, the appearances seem a bit weird.
I mean, I'm totally inclined to believe that Google is somehow promoting G+, but this does not strike me as convincing evidence. Neither does "I've seen lots of other examples!" Is there some credible information to be had, here?
Frankly, the output for this query suggests to me that Google isn't doing a very good job on this query in the first place. The photos of Guy aren't until the second page of results, and for some reason AllTop (possibly for SERP diversity reasons) is above many of his social media profiles.
Lastly, the G+ link is the seventh link. Only 50% of searchers even look at what the seventh link is! Is the claim that they are just trying to juice the results subtly?
The project page is very unclear, so I ended up on Wikipedia instead. Wikipedia actually has a leaked memo that seems to do a really good job describing the purpose of the language.
Specifically, the language is meant for three environments: server-side, compiled to JavaScript client-side, and fast native client-side once there is browser support. (The main goal is better performance on the client-side, which is deemed to be very difficult with JavaScript.)
However, the language looks so Java-like, one wonders why they didn't just use Java and extend GWT with a native Java client in Chrome. Did it just not make sense to bet the farm on Java when Oracle controls it?
Also, what does the "structured" in "structured web programming" mean?
What tools do you end up using? Some sort of custom apt repository plus some shell scripts? What does it look like?
I ended up literally printing out the wiki. (And the wiki seemed to be in a state of pretty extreme flux and/or disagreement with what blog posts suggested was best practice.)
The business model of the companies promoting Puppet and Chef seems to be to charge for support and/or hosted services. Which is fine. But is it leading to abysmal documentation?
Have you tried it?
I understand that people overwhelmingly tend towards the extremes of rating scales. But who cares? There's so much noise anyway.
Here's what I want: (1) include (all) films and television series and (2) include top 1000 lists for as many niche categories as possible (foreign films about relationships, teen comedy, and so on).
I don't even care if it is personalized, because the personalization is always horrible. (I've tried Netflix, Jinni and others and they're all terrible.) Just having the average for the movie is actually amazingly accurate on many measures (within 10--20%).
Is there something like this? The closest I've gotten so far is finding a good IMDB user list or two.
The analogy between music and tech suggests that the two hits-based industries selling digital goods are not so different after all. But why would our morality in one case be so different from the other case?
If anything, this allegory seems to suggest to me that the recording industry has committed one of the biggest public relations bungles of our lifetimes. People who wouldn't think to pirate a mobile app would pirate music without a second thought thanks to widespread lawsuits, pay-for-play, price fixing, and of course, ease.
Given that, NoSQL adoption often appears driven by success stories. But Cassandra seems to have the opposite: numerous failure stories. Facebook, Digg, Reddit and a number of others have all tried Cassandra in production, and have either had serious complaints or moved off to either SQL or other solutions like HBase.
Of course, these failure stories are anecdotes, and numerous unrelated factors (like bad interactions between Cassandra and Amazon's EC2) could be at fault. But I'm not sure it matters.
Has anyone on HN had a really good experience with Cassandra? (This may be the wrong thread to ask for obvious reasons.)
However, I still don't quite get it. What is Riak? It seems to be some sort of Dynamo implementation (like the ironically ill fated Cassandra), but apparently it has a workflow engine? What do people use it for? What is it best at?
Right now we're using PostgreSQL, Redis, and S3. PostgreSQL gives us ACID, Redis gives us fast in-memory access, and S3 gives us an infinite KV store. Is there some reason to use Riak? Would Riak just replace S3?
The strategy as I understand it is:
1. Convince a high profile conference in ${FIELD} that reproducibility is important.
2. Create a special group within that community to test submitted code to see if it matches the results presented in papers.
3. Give a special carrot to authors (a special mention in the program, a piece of text in their paper) who meet the expectations of this group.
4. Hopefully, eventually readers come to see papers with the markers indicating reproducibility as the only legitimate ones, and writers are then required to make the significant time commitment (and take the significant risks) of releasing their code.
As it happens, (1), (2), and (3) have happened in a few systems communities. For example, SIGMOD has (more or less) the same setup as you describe.
However, I have deep doubts about whether (4) will ever happen. The three issues are:
1. The group doing the evaluation of the code for the conference has a boring, unappreciated job. They are also reading terrible, likely buggy code. A natural outcome is that the evaluation group will make bold claims about how all of the code they evaluated had significant issues potentially impacting research results, making everyone who submitted look bad, and leading to disincentives for future submitters. In fact, the evaluation group may even write papers about how bad specific code they reviewed was. I believe this has happened in other communities.
2. I briefly alluded to this in my original post, but many actors have extremely good reasons (at least on their face) for not releasing their code and/or data. This is why I mentioned how researchers embedded at companies modifying large proprietary code bases are extremely unlikely to ever be part of this evaluation regime. (And, no one wants to kick such researchers out of the academic community.)
3. In order for a stigma to be attached to non-reproducibility according to the conference, there has to be a strong correlation between the highest quality work and reproducibility. However, it is likely that much of the highest quality work will not be reproducible, either because it comes out of (or in conjunction with) corporate research labs, or because it uses some very difficult to get proprietary data. Likewise, the most easily reproducible results may be the least significant.
Do you think that these issues are solvable in the long term?
We have discussed this topic on HN a number of times, for example:
http://news.ycombinator.com/item?id=2735537 http://news.ycombinator.com/item?id=2006749
Many of the comments in those threads do a better job summing up than I ever could. However, briefly, literally all of the incentives are aligned against publishing code and data.
If a writer's code is wrong, they are embarrassed (and there is no culture of being embarrassed by not publishing code).
If a writer publishes their code and it is actually good, someone else can scoop their follow-on results.
If a writer does not publish their code, and it is actually any good, they can potentially commercialize it thanks to the Bayh-Dole Act.
If a writer publishes their code and people intend to use it, the writer needs to clean it up, check it for correctness, and handle support requests. These activities are probably more time consuming than writing the code in the first place.
If the writer publishes their code, and other people in the writer's field do not, the writer is usually at a disadvantage. Others will appear to have more publications, the basic currency of academia. (Many people have great reasons for not publishing their code or data, especially researchers embedded at large companies making changes to large proprietary systems.)
So overall, yes, it would be great if CS paper writers gave out their code. What they are doing is not reproducible science in the philosophy of science sense.
But what is Jacques (or anyone else) doing to fix this system of incentives, and what could anyone do?
That said, is there something that would be lost if everyone just switched from MySQL to PostgreSQL tomorrow? What benefits does MySQL have over PostgreSQL these days?
From the discussion here, it sounds like:
1. futuremint was developing a bunch of Rails applications in Rails 2 + Ruby 1.8.x.
2. futuremint got bored of coding in Rails, was tired of Ruby being slow, and yearned for an IDE.
3. futuremint checked out Smalltalk/Seaside and Common Lisp. He liked it for a while (for reasons that seem unclear). There doesn't seem to be much discussion of continuations (which seem most salient to me as an outsider, but who knows!). Refactoring in Smalltalk is easier (which isn't surprising given that it's a really simple language).
4. futuremint returns to Rails 3, Ruby 1.9. Things are faster, he uses an IDE, and he is content with his ability to meta-program even if the syntax is ugly. futuremint is tired of lots of boilerplate and conventions because Smalltalk is so small, and he's tired of not having a community like Ruby/Rails in either Smalltalk/Seaside or Common Lisp.
Is Twitter Bootstrap what everyone should be using to make their MVP when they don't have a designer? Or is it sufficiently non-cross-platform compatible and specific to Twitter's requirements (no jQuery/jQuery UI?) that one should just use another admin theme from Theme Forest?
(Also, Blekko somehow supports grep on its dataset, but doesn't support it without voting? Is this a social mechanism, or does their infrastructure actually work in a way that this is an expensive operation?!)