HNHacker News
TopNewBestAskShowJobs

shadowmatter

532 karma · joined September 23, 2010

https://github.com/mgp https://twitter.com/omgitsmgp http://omgitsmgp.com/

Currently: Movin' money at Square

Previously: Backend developer at Airbnb, mobile developer at Khan Academy, lead developer on Burner, SWE at Google

submissionscomments
shadowmatter··on SpaceX Falcon 9/Dragon Launch Webcast (starts at 12:00am PDT)
As someone who bowed out of the SpaceX interview process because he couldn't see himself coding in C++ anymore, and now focuses on web and mobile instead... I wish I shared your confidence right now. That launch gave me goosebumps that no app ever has >.<
shadowmatter··on Goodbye, CouchDB
I wish this article had some hard numbers for availability, performance, and the size of their data as opposed to hand-waving.

Shameless plug: If you're looking to benchmark or load test CouchDB a bit, I wrote one at https://github.com/mgp/iron-cushion. Hopefully someone out there will use this to decide if CouchDB's performance meets their needs, because migrating away from any database is painful...

shadowmatter··on Swiftype (YC W12) Builds Site Search That Doesn’t Suck
I love the idea, and I think there is a huge, ignored opportunity here. But relevance is what makes or breaks search engines, and dabbling with a search engine I made for Joel on Software (like in the example video), I think you should work on tuning your scoring function.

Searching for "about me" (with quotes) returns the "Distributed Version Control is here to stay, baby" article at http://www.joelonsoftware.com/items/2010/03/17.html first, and http://www.joelonsoftware.com/AboutMe.html second. While it's pretty stupid that the "About Joel Spolsky" text is in a <div> tag versus a <hX> tag, maybe you could weight the title more, because likely people will be using Swiftype to search curated corporate sites that are typically less spammy than the general Internet.

Anyway, great job so far, and I'm very excited to see this and where this could lead!

shadowmatter··on Rob Pike - Unix trivia
Until Mark Pincus wins a Turing Award for Farmville, it's probably not fair to compare the two.
shadowmatter··on Google's opening slides in the trial v. Oracle
On slide 52, which is supposed to demonstrate that the Android and Java source code implementations differ, it's clear that the Android definition did not come from Android: The anotherString parameter is not referenced, and there are no offsets as parameters.

I also find it amusing that in Oracle's slides, "Java On Our Computers" is demonstrated with the Java update screen. Sigh.

shadowmatter··on Sergey Brin: Facebook and Apple a threat to Internet freedom
Kinda sick of hearing this. I mean, you've been able to download an archive of everything you've posted to Facebook for awhile now. The Facebook APIs are way more built-out than the Google+ APIs. And here are like a billion other web sites out there that only make information available after logging in and so also qualify as walled gardens, but apparently aren't worthy of criticism.

Meh.

shadowmatter··on Launch HN - Overnight Buses Travel Magazine for the iPad
Congrats on the first launch. Looks very professional!

For the lazy like me: http://itunes.apple.com/us/app/overnight-buses-magazine/id49...

shadowmatter··on Android: Faster emulator with better hardware support
I was testing this last night. It reduced the time for the emulator to cold boot from 120-130 seconds down to 45-55 seconds. Not as snappy as the iPhone simulator, but over cutting the time by over half is really substantial. Props to them.
shadowmatter··on CouchDB ships 1.2.0, Windows packages, new website
I found the following two documents on the replication protocol:

http://www.dataprotocols.org/en/latest/couchdb_replication.h... https://github.com/couchbaselabs/TouchDB-iOS/wiki/Replicatio...

From the end of http://wiki.apache.org/couchdb/Replication

shadowmatter··on Snapjoy’s Flickraft Promised To Rescue Flickr Photos — Until It Was Blocked
"We built the system to stay within the limits of 3600 calls per hour, however it seems that a surge of imports pushed it over the threshold before we could throttle it back."

It sounds like they designed it to stay within the rate limits, but had a bug, so it didn't.

"We’re a bit surprised that the key was disabled almost immediately after we reached the limit."

Is it crazy to think that Flickr abides by the rate limiting they advertise? Isn't that just like "truth in advertising"? The onus is on you to stay within that rate.

"We thought about creating a new api key but didn’t know if that would be flagged as abuse."

Is there really any part of your gut that says no? Do you think Flickr would implement a rate limiting policy if they were okay with people creating an unlimited number of keys so that it serves no point?

Sorry to be so negative, but... Is this really news?

shadowmatter··on IOS Boilerplate: A base template for iOS apps
Looks good. Just FYI: There's already a MapKit class called MKPointAnnotation that provides a simple implementation of MKAnnotation, so your Place class is redundant:

    MKPointAnnotation *point = [[MKPointAnnotation alloc] init];
    point.coordinate = CLLocationCoordinate2DMake(35.01234, -115.56789);
    point.title = @"The title!";
    point.subtitle = @"The subtitle!";
    [point release];
shadowmatter··on I am nothing
Good post. Reminds me of that quote by Oscar Wilde: "Be yourself; everyone else is already taken."
shadowmatter··on Why HN Got Slow
Mainstream media'd. Meaning someone from the mainstream media posted a link to that post, and the surge in traffic made everything slow.
shadowmatter··on LevelDB: A Fast Persistent Key-Value Store
Jeff and Sanjay wrote the original protocol buffer implementation. The project was taken over by Kenton Varda, who rewrote the C++ and Java parts; this is what was open sourced. See http://temporal.fateofio.org/files/resume
shadowmatter··on A note to Google recruiters (and on Google hiring practices)
Like you, I disagree with some of the points in her blog post, but I can't agree with your statement "As for what team you'll be working on, you as a candidate have a lot of power in that regard."

I had one referral decline an offer because he wouldn't know what project he would be working on until he was hired, and therefore, didn't know if it would be more enjoyable than his current job. I also knew PhDs who were hired and expressed disappointment that the project they were working on didn't make use of the specialized material they studied towards their degree. And I wouldn't think that the majority of new engineers assigned to ads projects expressed advertisement as a preference with their recruiters.

But most new hires, well, get over it.

Yes, there are some cases where people are recruited with specific expertise to fulfill specific roles, e.g. John Barton of Firebug or Sebastian Thrun for the self-driving car. But this is the exception and not the rule.

shadowmatter··on I Broke Justin.tv
That's what caught my eye too. I like how it's soon followed by "Pushing code fast and often does have a cost, but the benefit in productivity is well worth it." That's certainly true if it's mostly _tested_ code, but there is nothing productive about spending a day trying to find and fix a bug like this in production. Been there, done that.

I'm trying to bootstrap a startup now, and if it fails to gain traction then justin.tv would be one of the first places I'd interview. (Hey, I like watching competitive TF2, and they provide a hell of a platform for casts, can I say.) Now I know this blog post was sort of written for the purposes of recruitment, but it's sort of making me think twice about whether I'd want to interview there. Bleh.

shadowmatter··on Terence Tao's General Exam
I did a math minor at UCLA, and he taught the upper-division linear algebra class I took. He didn't like the book, so he decided to write his own lecture notes, which formed a book unto themselves. They're still available online at http://www.math.ucla.edu/~tao/resource/general/115a.3.02f/. If you ever have the itch to learn linear algebra, read them, they're quite excellent.

It was pretty obvious at class and obvious hours he was crazy smart -- but I had no way of knowing he was Fields Medal smart. But unlike other crazy smart professors I've had, he's a very gifted teacher as well. I mostly learned by reading the books and considered the lectures as an ancillary learning aid, but his lectures were very illuminating. I'm glad he now has a blog to teach a wider audience at http://terrytao.wordpress.com/, but unfortunately most of it is beyond what I can understand. If you're a math die-hard, be sure to read that.

shadowmatter··on Stonebraker trapped in Stonebraker 'fate worse than death'
His "me" page explains that he's a database engineer at Facebook, and before that worked on performance-oriented systems and software at Wikipedia. It's part of his job to be aware of the drop in flash costs and employ it where he can while balancing performance with operating costs and expenses. No one is arguing that flash memory isn't as permanent as disks, but it's certainly not as cheap, especially on the scale of data that Facebook has.

Every year or so Stonebreaker feels compelled to come down from his ivory tower to troll industry. Remember a few years ago when he co-authored the paper "MapReduce: A major step backwards" (http://databasecolumn.vertica.com/database-innovation/mapred...)? Nevermind that Google probably processes petabytes of data using MapReduce every day that ends up getting served on practically every Google property. How can he argue with results? Can't we just ignore him?

shadowmatter··on Michael Abrash on Quake: "Finish the product and you’ll be a hero."
Note that fountain codes are very, very patent-protected by Luby et al. Five years ago I came across literature on rateless codes, which are digital fountain codes that allow practically infinite encoding, or a practically infinite stream of data from which the original could be reconstructed. One really promising type of rateless code was online codes, which required O(1) time to generate a single block and O(n) time to decode a message of length n, which was much better than the LT rateless codes that came before it. (The paper can be found at http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.12....)

I wrote up a Java implementation in my free time, and it worked well and was quite fast. I pinged the authors of the paper, asking if I could open-source it, because I came across a web site with their names on it that seemed like they were going to commercialize the idea. They replied that they had abandoned the idea of online codes because, even though their approach was faster than all other coding schemes, it still violated patents held by Luby and his company Digital Fountain, which has since been acquired by Qualcomm. That was a bummer, because I thought the online codes paper was very elegant, and the implementation was quite simple.

shadowmatter··on Birth and Death of Microsoft Bing
I'm an ex-Googler, and I disagree. Before I left the GWS team (http://www.theregister.co.uk/2010/01/29/google_web_server/) became a model for how frequent, stable releases should be done within the company. That's all I'll say.
shadowmatter··on Mistakes Google made in scaling its organization
As an ex-Googler, I disagree with this second-hand account. Yeah, political infighting sometimes happens and things get shook up, but this person paints the picture of no project ever launching. Addressing some points:

1. See item 3.

2. VP changes rarely happen compared to a PM change. And a PM who comes may change some of the goals, but the focus area of your product doesn't change, e.g. your optimization tool for display advertising will continue to be an optimization tool for display advertising. It would be hard to find your project suddenly disagreeing with some "existing strategy."

3. New products that launch should have a UX designer contributing or at least advising. There should be mocks for the frontend engineers to follow, and the backend engineers are just told to make it work.

4. Yeah, there's a checklist, and it is long. SREs want to minimize the number of pages there, but pages happen because something's going bad, and they want you to have proper monitoring in place to address the pages quickly. The launch checklist also focuses a lot on addressing potential security vulnerabilities. Google is tough on security.

5. Google's all about the software compensating for commodity hardware. I don't think too many teams have specialized hardware. You can ask for what resources you need, but would have to fight for what form you get them in.

7. Can't speak for this. Why reorganize marketing?

8. While I agree that Google likes to fail fast, in my opinion some Google products probably didn't get the love or time they needed to take off. I wasn't really privy to how those decisions are made.

shadowmatter··on First Glance at Objective-C and the iPhone SDK (by a Java and C Developer)
I started about a month ago... As someone with a background in Java/C++/Python, here are some more that have caught my attention:

- The ability to assign types as ClassName<Protocol>, instead of the protocol specifying the type completely. Neat, I guess...

- Because it's message passing, you get duck typing for free, albeit with compiler warnings (you send a message, and wouldn't you know it, the selector exists!). Not that I recommend this approach over protocols.

- The awesomeness of @selector, especially if you're coming from Java and used to writing anonymous classes everywhere simply for the sake of making functions first-class entities.

- The lack of public/protected/private visibility modifiers and how variables are suddenly part of the public API when you define them in properties. Say you have a variable you want to keep private, but you need to frequently update its value. You go ahead and define @property (nonatomic, retain) for it so within the class so that you can simply do self.varName = newVarValue to update it. But now all your clients/outside clients will see; you need to write the release-and-retain setter yourself in the .m file to prevent this.

- The minimalism of Xcode. Hey, it's faster than Eclipse, thank $DEITY for that. But when I want to alphabetize the filenames in the left sidebar, I have to go to Edit -> Sort -> By Name for _every_ file group?

shadowmatter··on Git Immersion
What he really means is that branches are now a part of your daily workflow. You can sync up to head (or wherever you pulled your working copy of the repository from), create a branch, and start working on some feature on that branch, mucking up a hundred files along the way. Now say some some urgent issue comes up that you need to fix immediately -- you can easily go back to the point before you branched, and create a new branch from there on which you can fix the issue. Meanwhile, your previous branch still exists, so when the urgent issue is fixed you can resume where you are working on. Another great thing about branches is say you're working on some feature in a branch and you come to a point where you can implement something multiple ways, but don't know which one will turn out elegant. You can create a new branch from that point in your current branch, and if it doesn't pan out, revert it to where you were.

The kicker with Git is that all these branch operations take on the order of milliseconds because Git stores the entire project history -- all past revisions of files, etc -- locally. Sure, you pay in a bit of disk space, but consequently almost all common operations can be done without hitting the network and are crazy fast. Even when you execute a commit, you're not committing to a server, but to your working copy of the repository. Later, you can push your changes somewhere else and merge accordingly.

shadowmatter··on Dating Denial of Service attack
This is similar to the Sybil attack in peer-to-peer networking. From http://en.wikipedia.org/wiki/Sybil_attack: "A Sybil attack is one in which an attacker subverts the reputation system of a peer-to-peer network by creating a large number of pseudonymous entities, using them to gain a disproportionately large influence."
shadowmatter··on Google and Microsoft Cheat on Slow-Start. Should You?
I thought that the minimum value cwnd could assume is IW, but looking at RFC 2581 that isn't true: "Upon a timeout cwnd MUST be set to no more than the loss window, LW, which equals 1 full-sized segment (regardless of the value of IW)." They even explicitly call out what I erroneously believed, so you are right -- I apologize!

If you really have a tin cans and string physical layer, Google's larger IW could be more disruptive to other connections on the link: A vanilla TCP connection would ramp up cwnd from 1 segment until congestion is observed (in the slow-start phase) and then grow cwnd conservatively (after ssthresh is first set). If congestion would be observed at a cwnd value less than 10 segments, then starting with IW at 10 segments could be very disruptive to others sharing the connection.

Mind you, this argument feels very... academic. As you pointed out, Google's connection would converge toward fairness anyway. (Unless so little data is actually transmitted, then the connection probably isn't open long enough for that to happen.) And most shared links don't saturate at 12 segments. I'd guess that a high-capacity link would only be at risk if it has a lot of connections (so that every connection has a cwnd not much greater than 10) and there are always many (albeit, short-lived) connections to Google always being created (which could appear as fewer long-lived connections with a constant cwnd of 10, more than would be fair).

shadowmatter··on Google and Microsoft Cheat on Slow-Start. Should You?
"Benefiting _their_ users," yes. But they're not benefiting whoever else is sharing the smallest-capacity link -- what he's arguing is that Google is crowding them out.

Net neutrality is not the only way to privilege your flows on the Internet: There's nothing to stop me from writing a crude application-layer protocol atop UDP that implements reliability but not congestion control. (You maybe had to implement something like this for your networking class in school; otherwise you could start with a protocol like DCCP.) If I were to use that to send data as fast as I could to some remote computer, I could be sending more data than the smallest-capacity link could handle. Other TCP/IP connections sharing that link would detect data loss and thus reduce the amount of data they put in transit, but my protocol wouldn't have to. I can monopolize that link.

So assuming your physical layer is tin cans and string, what he's arguing is that if you have a link with a capacity of 12 segments, then data from Google will use 10 of them and a client will never expand its outstanding data beyond 2 segments. If both used vanilla TCP/IP, they should share the link evenly.

Of course, speed is a critical factor for Google. Android by default uses TCP Westwood+.

It's been six years since I tinkered with TCP/IP and really focused on networking, so someone please correct me if I'm wrong >_<

shadowmatter··on A zombie keyboard, an app-store rejection, a call from Steve Jobs
Agreed. When you make an exception for developers to rely on a private method, you "promote" an implementation detail to its public interface. Now you can never remove the method, change its signature, or change its behavior. Nevermind that when the bug in the public method is fixed, you now have two methods to accomplish the same task. You've introduced cruft, or legacy code you need to maintain. Ick.
shadowmatter··on Rob Pike: Notes on Programming in C
"Balanced binary tree versus hash table is an implementation choice for the associative array abstract data type."

You can treat it as an implementation detail, but it's a good idea to let the client know that the underlying key-value collection is actually sorted by key. This allows the client to iterate over the key-value pairs in sorted order without dumping the keys to an array and then sorting it, retrieving the k smallest keys through forward iteration, retrieving the k largest keys through reverse iteration, etc. I think Java solves this nicely by creating the SortedMap subinterface of Map; if you have a SortedMap, you're assured these properties hold. (The typical implementation of SortedMap is a balanced tree, while the typical implementation of Map is a hash table.)

shadowmatter··on Google Engineering Management Mistakes
You can give a peer bonus of $200 to any individual with their manager's approval (which is very easy to get). They encourage giving peer bonuses for when someone goes the extra mile on something, but very few engineers initiate giving a peer bonus. After tax, it's about $100, hence the "I can poop $100" bit.
shadowmatter··on Does anyone have a great idea for what to name a new datastore?
How about Newd? For "New Datastore," of course. And all your clients can say "I'm going newd" when they adopt it.
← PreviousPage 2 of 3Next →