532 karma · joined September 23, 2010
Currently: Movin' money at Square
Previously: Backend developer at Airbnb, mobile developer at Khan Academy, lead developer on Burner, SWE at Google
Shameless plug: If you're looking to benchmark or load test CouchDB a bit, I wrote one at https://github.com/mgp/iron-cushion. Hopefully someone out there will use this to decide if CouchDB's performance meets their needs, because migrating away from any database is painful...
Searching for "about me" (with quotes) returns the "Distributed Version Control is here to stay, baby" article at http://www.joelonsoftware.com/items/2010/03/17.html first, and http://www.joelonsoftware.com/AboutMe.html second. While it's pretty stupid that the "About Joel Spolsky" text is in a <div> tag versus a <hX> tag, maybe you could weight the title more, because likely people will be using Swiftype to search curated corporate sites that are typically less spammy than the general Internet.
Anyway, great job so far, and I'm very excited to see this and where this could lead!
I also find it amusing that in Oracle's slides, "Java On Our Computers" is demonstrated with the Java update screen. Sigh.
Meh.
For the lazy like me: http://itunes.apple.com/us/app/overnight-buses-magazine/id49...
http://www.dataprotocols.org/en/latest/couchdb_replication.h... https://github.com/couchbaselabs/TouchDB-iOS/wiki/Replicatio...
From the end of http://wiki.apache.org/couchdb/Replication
It sounds like they designed it to stay within the rate limits, but had a bug, so it didn't.
"We’re a bit surprised that the key was disabled almost immediately after we reached the limit."
Is it crazy to think that Flickr abides by the rate limiting they advertise? Isn't that just like "truth in advertising"? The onus is on you to stay within that rate.
"We thought about creating a new api key but didn’t know if that would be flagged as abuse."
Is there really any part of your gut that says no? Do you think Flickr would implement a rate limiting policy if they were okay with people creating an unlimited number of keys so that it serves no point?
Sorry to be so negative, but... Is this really news?
MKPointAnnotation *point = [[MKPointAnnotation alloc] init];
point.coordinate = CLLocationCoordinate2DMake(35.01234, -115.56789);
point.title = @"The title!";
point.subtitle = @"The subtitle!";
[point release];I had one referral decline an offer because he wouldn't know what project he would be working on until he was hired, and therefore, didn't know if it would be more enjoyable than his current job. I also knew PhDs who were hired and expressed disappointment that the project they were working on didn't make use of the specialized material they studied towards their degree. And I wouldn't think that the majority of new engineers assigned to ads projects expressed advertisement as a preference with their recruiters.
But most new hires, well, get over it.
Yes, there are some cases where people are recruited with specific expertise to fulfill specific roles, e.g. John Barton of Firebug or Sebastian Thrun for the self-driving car. But this is the exception and not the rule.
I'm trying to bootstrap a startup now, and if it fails to gain traction then justin.tv would be one of the first places I'd interview. (Hey, I like watching competitive TF2, and they provide a hell of a platform for casts, can I say.) Now I know this blog post was sort of written for the purposes of recruitment, but it's sort of making me think twice about whether I'd want to interview there. Bleh.
It was pretty obvious at class and obvious hours he was crazy smart -- but I had no way of knowing he was Fields Medal smart. But unlike other crazy smart professors I've had, he's a very gifted teacher as well. I mostly learned by reading the books and considered the lectures as an ancillary learning aid, but his lectures were very illuminating. I'm glad he now has a blog to teach a wider audience at http://terrytao.wordpress.com/, but unfortunately most of it is beyond what I can understand. If you're a math die-hard, be sure to read that.
Every year or so Stonebreaker feels compelled to come down from his ivory tower to troll industry. Remember a few years ago when he co-authored the paper "MapReduce: A major step backwards" (http://databasecolumn.vertica.com/database-innovation/mapred...)? Nevermind that Google probably processes petabytes of data using MapReduce every day that ends up getting served on practically every Google property. How can he argue with results? Can't we just ignore him?
I wrote up a Java implementation in my free time, and it worked well and was quite fast. I pinged the authors of the paper, asking if I could open-source it, because I came across a web site with their names on it that seemed like they were going to commercialize the idea. They replied that they had abandoned the idea of online codes because, even though their approach was faster than all other coding schemes, it still violated patents held by Luby and his company Digital Fountain, which has since been acquired by Qualcomm. That was a bummer, because I thought the online codes paper was very elegant, and the implementation was quite simple.
1. See item 3.
2. VP changes rarely happen compared to a PM change. And a PM who comes may change some of the goals, but the focus area of your product doesn't change, e.g. your optimization tool for display advertising will continue to be an optimization tool for display advertising. It would be hard to find your project suddenly disagreeing with some "existing strategy."
3. New products that launch should have a UX designer contributing or at least advising. There should be mocks for the frontend engineers to follow, and the backend engineers are just told to make it work.
4. Yeah, there's a checklist, and it is long. SREs want to minimize the number of pages there, but pages happen because something's going bad, and they want you to have proper monitoring in place to address the pages quickly. The launch checklist also focuses a lot on addressing potential security vulnerabilities. Google is tough on security.
5. Google's all about the software compensating for commodity hardware. I don't think too many teams have specialized hardware. You can ask for what resources you need, but would have to fight for what form you get them in.
7. Can't speak for this. Why reorganize marketing?
8. While I agree that Google likes to fail fast, in my opinion some Google products probably didn't get the love or time they needed to take off. I wasn't really privy to how those decisions are made.
- The ability to assign types as ClassName<Protocol>, instead of the protocol specifying the type completely. Neat, I guess...
- Because it's message passing, you get duck typing for free, albeit with compiler warnings (you send a message, and wouldn't you know it, the selector exists!). Not that I recommend this approach over protocols.
- The awesomeness of @selector, especially if you're coming from Java and used to writing anonymous classes everywhere simply for the sake of making functions first-class entities.
- The lack of public/protected/private visibility modifiers and how variables are suddenly part of the public API when you define them in properties. Say you have a variable you want to keep private, but you need to frequently update its value. You go ahead and define @property (nonatomic, retain) for it so within the class so that you can simply do self.varName = newVarValue to update it. But now all your clients/outside clients will see; you need to write the release-and-retain setter yourself in the .m file to prevent this.
- The minimalism of Xcode. Hey, it's faster than Eclipse, thank $DEITY for that. But when I want to alphabetize the filenames in the left sidebar, I have to go to Edit -> Sort -> By Name for _every_ file group?
The kicker with Git is that all these branch operations take on the order of milliseconds because Git stores the entire project history -- all past revisions of files, etc -- locally. Sure, you pay in a bit of disk space, but consequently almost all common operations can be done without hitting the network and are crazy fast. Even when you execute a commit, you're not committing to a server, but to your working copy of the repository. Later, you can push your changes somewhere else and merge accordingly.
If you really have a tin cans and string physical layer, Google's larger IW could be more disruptive to other connections on the link: A vanilla TCP connection would ramp up cwnd from 1 segment until congestion is observed (in the slow-start phase) and then grow cwnd conservatively (after ssthresh is first set). If congestion would be observed at a cwnd value less than 10 segments, then starting with IW at 10 segments could be very disruptive to others sharing the connection.
Mind you, this argument feels very... academic. As you pointed out, Google's connection would converge toward fairness anyway. (Unless so little data is actually transmitted, then the connection probably isn't open long enough for that to happen.) And most shared links don't saturate at 12 segments. I'd guess that a high-capacity link would only be at risk if it has a lot of connections (so that every connection has a cwnd not much greater than 10) and there are always many (albeit, short-lived) connections to Google always being created (which could appear as fewer long-lived connections with a constant cwnd of 10, more than would be fair).
Net neutrality is not the only way to privilege your flows on the Internet: There's nothing to stop me from writing a crude application-layer protocol atop UDP that implements reliability but not congestion control. (You maybe had to implement something like this for your networking class in school; otherwise you could start with a protocol like DCCP.) If I were to use that to send data as fast as I could to some remote computer, I could be sending more data than the smallest-capacity link could handle. Other TCP/IP connections sharing that link would detect data loss and thus reduce the amount of data they put in transit, but my protocol wouldn't have to. I can monopolize that link.
So assuming your physical layer is tin cans and string, what he's arguing is that if you have a link with a capacity of 12 segments, then data from Google will use 10 of them and a client will never expand its outstanding data beyond 2 segments. If both used vanilla TCP/IP, they should share the link evenly.
Of course, speed is a critical factor for Google. Android by default uses TCP Westwood+.
It's been six years since I tinkered with TCP/IP and really focused on networking, so someone please correct me if I'm wrong >_<
You can treat it as an implementation detail, but it's a good idea to let the client know that the underlying key-value collection is actually sorted by key. This allows the client to iterate over the key-value pairs in sorted order without dumping the keys to an array and then sorting it, retrieving the k smallest keys through forward iteration, retrieving the k largest keys through reverse iteration, etc. I think Java solves this nicely by creating the SortedMap subinterface of Map; if you have a SortedMap, you're assured these properties hold. (The typical implementation of SortedMap is a balanced tree, while the typical implementation of Map is a hash table.)