HNHacker News
TopNewBestAskShowJobs

indygreg2

862 karma · joined May 4, 2011

submissionscomments
indygreg2··on Notes from Facebook's Developer Infrastructure at Scale F8 Talk
It is much easier to bisect linear history than history with merges.
indygreg2··on Notes from Facebook's Developer Infrastructure at Scale F8 Talk
Except it isn't internal only. Many of Facebook's tools are open sourced. And not in the "throw it over the wall" sense. Their open sourced projects tend to gain traction and get a significant amount of community contributors. By turning a number of their core tools into successful open source projects, they are leveraging the community effect to lesson the ongoing maintenance burden for these projects.

Yes, they are paying a high initial cost to develop these offerings. But since they tend to produce high quality products that attract nearly-free-to-Facebook labor via open source, the investment tends to pays off in the long run. It's a savvy business move. And one that can arguably only be pulled off by a talented engineer organization.

indygreg2··on Notes from Facebook's Developer Infrastructure at Scale F8 Talk
The site is hosted via GitHub pages. Maybe it's not available to you due to the lingering GitHub DoS. If you care to read HTML and can manage to get through to GitHub: https://github.com/indygreg/indygreg.github.com/blob/master/...
indygreg2··on Bazel – Correct, reproducible, fast builds for everyone
I invented the Python bits of the Firefox build system (moz.build files). I learned after I implemented them that Google's internal approach with Blaze was very similar. It felt reassuring that I independently reinvented a similar solution :)

There are a handful of Blaze derivatives built by Xooglers. Pants and Buck come to mind. They also share the trait of using sandboxed Python to define a build configuration. I'll take it over make syntax any day!

indygreg2··on Git client vulnerability announced
More back story: https://twitter.com/indygreg/status/545701974671233024
indygreg2··on Spidermonkey has passed V8 on Octane performance
If you sign in to Chrome, your bookmarks and full browsing history are uploaded to Google's servers. Only your passwords are encrypted locally before being sent to Google.

https://support.google.com/chrome/answer/1181035 contains info on how the encryption works.

chrome://terms/ links to https://www.google.com/intl/en/chrome/browser/privacy/ which links to http://www.google.com/policies/privacy/ to define "how we use information we collect." From that page: "We use the information we collect from all of our services to provide, maintain, protect and improve them, to develop new ones, and to protect Google and our users. We also use this information to offer you tailored content – like giving you more relevant search results and ads."

Your full browsing history is a treasure trove of information useful for making Google's core services (search and ads) more effective. They would be stupid not to use it to improve the quality of their services. I challenge your assertion that Chrome is an altruistic endeavor.

indygreg2··on Ask HN: How do you use Docker in production?
I blogged about this the other day and would love to hear about your experience!

http://gregoryszorc.com/blog/2014/10/13/deterministic-and-mi...

indygreg2··on Deterministic, bit-identical and/or verifiable Linux builds
The path towards deterministic builds is definitely not clear. As many in this thread have pointed out, it's a difficult technical problem. The difficulties are multiplied by a project at Firefox's scale.

Further complicating matters is our platform breakdown. The majority of Firefox users are on Windows. Deterministic builds on Windows are very painful. And that's before you figure PGO into the mix. Tor works around this by compiling Firefox with an open source toolchain and doesn't use PGO. But that's a non-starter for us because choosing an open source toolchain over Microsoft's would result in performance degradations for our users. Believe me, if we could ship a Windows and Mac Firefox built with 100% open source to no detriment to our users, we would. There's work to get Firefox building with Clang on Windows (but only for doing ASAN and static analysis, not for shipping to users). That gets us one step closer.

All that being said, there has been exploratory talk lately of serving segments of our user base with specialized Firefox builds. e.g. a build with developer tools front and center that caters to the web development community. If that ever happens, I imagine a deterministically-built Firefox with things like Tor built in could be on the table. The way you can make that happen is to direct noise directly at the Mozilla community. Send a well-crafted email to firefox-dev (https://mail.mozilla.org/listinfo/firefox-dev) explaining your position. Anticipate that people will likely reply by asking you to prioritize this against existing goals, such as shipping 64-bit Firefox on Windows and shipping multi-process Firefox. We don't have nearly unlimited resources like some of the other browser vendors, so we can't just do everything. Again, I implore people to directly contribute to Mozilla any way they can. https://www.mozilla.org/contribute/

indygreg2··on Deterministic, bit-identical and/or verifiable Linux builds
I filed the linked bug and am the technical owner of Firefox's build system.

There were efforts made and discussions outside of the linked bug. To say "nothing" was done is just not true.

It would be more accurate to say that we just can't justify working on this right now because the timing isn't right and it's high cost for perceived low reward. The time of everyone involved to implement this would be better spent on improvements that benefit the general Firefox population. Some of those improvements include overhauling Firefox's build automation to better support things like building with Docker. That lays the groundwork for (easier) deterministic builds in the future. Even then, I'm not sure if this will happen. Brendan's post called on the larger community to make requests of Mozilla. That front has been surprisingly quiet. If you really want this, I would suggest making noise on the mozilla.org domain. Even better, contribute some patches, like the Tor Project has done: I will happily review them! #build on irc.mozilla.org.

indygreg2··on Facebook's git repo is 54GB
I think you are missing the point. Versioning and package management problems can largely go away when your entire code base is derived from a single repo. After all, library versioning and packaging are indirections to better solve common deployment and distribution requirements. These problems don't have to exist when you control all the endpoints. If you could build and distribute a 1 GB self-contained, statically linked binary, library versioning and packages become largely irrelevant.
indygreg2··on Facebook's git repo is 54GB
SOA isn't a magic bullet.

What if multiple services are utilizing a shared library? For each service to be independent in the way I think you are advocating for, you would need multiple copies of that shared library (either via separate copies in separate repos or a shared copy via something like subrepos).

Multiple copies leads to copies getting out of sync. You (likely) lose the ability to perform a single atomic commit. Furthermore, you've increased the barrier to change (and to move fast) by introducing uncertainty. Are Service X and Service Y using the latest/greatest version of the library? Why did my change to this library break Service Z? Oh, it's because Service Z lags 3 versions behind on this library and can't talk with my new version.

Unified repositories help eliminate the sync problem and make a whole class of problems that are detrimental to productivity and moving fast go away.

Facebook isn't alone in making this decision. I believe Google maintains a large Perforce repository for the same reasons.

indygreg2··on Facebook's git repo is 54GB
Having all code in a single repository increases developer productivity by lowering the barrier to change. You can make a single atomic commit in one repository as opposed to N commits in M repositories. This is much, much easier than dealing with subrepos, repo sync, etc.

Unified repos scales well up to a certain point before troubles arise. e.g. fully distributed VCS starts to break down when you have hundreds of MB and people with slow internet connections. Large projects like the Linux kernel and Firefox are beyond this point. You also have implementation details such as Git's repacks and garbage collection that introduce performance issues. Facebook is a magnitude past where troubles begin. The fact they control the workstations and can throw fast disks, CPU, memory, and 1 gbps+ links at the problem has bought them time.

Facebook made the determination that preserving a unified repository (and thus preserving developer productivity) was more important than dealing with the limitation of existing tools. So, they set out to improve one VCS system: Mercurial (https://code.facebook.com/posts/218678814984400/scaling-merc...). They are effectively leveraging the extensibility of Mercurial to turn it from a fully distributed VCS to one that supports shallow clones (remotefilelog extension) and can leverage filesystem watching primitives to make I/O operations fast (hgwatchman) and more. Unlike compiled tools (like Git), Facebook doesn't have to wait for upstream to accept possibly-controversial and difficult-to-land enhancements or maintain a forked Git distribution. They can write Mercurial extensions and monkeypatch the core of Mercurial (written in Python) to prove out ideas and they can upstream patches and extensions to benefit everybody. Mercurial is happily accepting their patches and every Mercurial user is better off because of Facebook.

Furthermore, Mercurial's extensibility makes it a perfect complement to a tailored and well-oiled development workflow. You can write Mercurial extensions that provide deep integration with existing tools and systems. See http://gregoryszorc.com/blog/2013/11/08/using-mercurial-to-q.... There are many compelling reasons why you would want to choose Mercurial over other solutions. Those reasons are even more compelling in corporate environments (such as Facebook) where the network effect of Git + GitHub (IMO the foremost reason to use Git) doesn't significantly factor into your decision.

indygreg2··on Facebook's git repo is 54GB
They aim for a completely linear history. They may even have a policy of not allowing merge commits. It is described in various places on the internet. I like https://secure.phabricator.com/book/phabflavor/article/recom... because it and its sister articles on code review and revision control are terrific reads.
indygreg2··on Prefer mercurial to git
I can relate to these comments because when I was a Git user forced to use Mercurial for Firefox development, I initially thought much of the same. I have since come around [1].

Mercurial has come a long way in the last few years. While I used to see repo corruption semi-frequently, I have not seen it once in the last year or so. This can be attributed to bug fixes and less reliance on mq. mq is a giant hack on top of Mercurial's storage model and there were many corner cases in older Mercurials where mq could lead to repo corruption. I use the experimental evolve extension now, but I can't yet recommend that to the masses because it's very rough around the edges. Hopefully in the next 6-12 months.

I strongly disagree with the statement that Mercurial is less flexible than Git. I find Mercurial to be more flexible. If you don't take my word for it, ask Facebook [2]: "Our engineers were comfortable with Git and we preferred to stay with a familiar tool, so we took a long, hard look at improving it to work at scale. After much deliberation, we concluded that Git's internals would be difficult to work with for an ambitious scaling project." The article goes on to describe some key areas where Mercurial is more flexible.

Some things possible in Mercurial that aren't with Git:

* Extending the wire protocol. Git's wire protocol is the exchange of objects (key-value pairs) and refs to said objects. Mercurial's is command-based and you can have client and server talk their own commands.

* Revision sets [3]. Extremely useful feature. See [4] for how I've used this as Mozilla.

* Phases. Mercurial knows when you are changing a published and should-be-immutable changeset/commit and by default prevents you from footgunning yourself.

* Changeset evolution. The mindset about "pushing rebases is evil" has its roots almost completely in limitations of tools. Changeset evolution removes that limitation.

As I wrote at [1], I believe the future of Mercurial is bright. Don't discount Mercurial because of past experiences with ancient versions or because you assume the Git way is the only and right way.

[1] http://gregoryszorc.com/blog/2013/05/12/thoughts-on-mercuria... [2] https://code.facebook.com/posts/218678814984400/scaling-merc... [3] http://www.selenic.com/hg/help/revsets [4] http://gregoryszorc.com/blog/2013/11/08/using-mercurial-to-q...

indygreg2··on Amazon’s Silk Browser tracks users' visited URLs
Chrome's sync (i.e. sign in to Chrome) uploads the cleartext of your full browsing history to Google by default: http://gregoryszorc.com/blog/2012/04/08/comparing-the-securi...

From a privacy perspective Chrome and Silk sound the same.

Now, maybe Amazon is more actively using that data and doesn't provide means to change privacy settings (Chrome allows you to locally encrypt). I dunno.

indygreg2··on Storing Passwords Securely
Assuming all the steps in the article are followed, yes, you are correct.

I still think any article talking about verifying credentials is obligated to mention that string comparison could be an attack vector.

Like I said, it plants a seed. And, I've seen way too many naive implementations where it is needed (like simple token-based auth systems) to know that this seed needs to be spread a lot more.

indygreg2··on Storing Passwords Securely
If you are going to go through all the effort to do it properly, you might as well use a proper comparison function. If nothing else, it reinforces the knowledge that string comparisons can be part of security (which goes overlooked by many).
indygreg2··on Google removed 1m DMCA links last month, 540,000 at the behest of Microsoft
The fact Microsoft is the #1 requester isn't a surprise to me. The 2nd most popular search engine probably automatically identifies "bad" content then siphons off this list to Google. If Google had more paid products (that weren't ad driven), I'd expect Google to be the largest referrer of Bing DMCA requests.
indygreg2··on Experimental Page Layout Inspired By Flipboard
Works great for me in Firefox Nightly. Couldn't notice any difference from Chrome, despite the message saying it works best in WebKit.
indygreg2··on Linux 3.3 released
If they keep releasing new versions so rapidly, I may have to switch to Chrome^H^H^H^H^HWindows.
indygreg2··on Chrome Overtakes Firefox Globally for First Time
Firefox's Sync synchronizes more frequently now than it did a few releases ago. If it doesn't synchronize fast enough for you, power users can always open about:config and fuddle with the services.sync preferences. services.sync.syncInterval is the one controlling the default interval (in milliseconds).

While I'm writing this, I should also point out that browser sync is a great example of how Mozilla and Google take a different approach to solving the same problem. Firefox's sync encrypts all data locally using a cryptographically secure randomly-generated key then uploads it to Mozilla's servers. Chrome's sync, by contrast, only encrypts passwords locally by default, leaving bits like your browsing history unencrypted on Google's servers. Chrome does have an option to encrypt everything, but you have to enable it in the preferences. (Firefox has no option to disable client-side encryption.) Even when you enable client-side encryption in Chrome, your data is encrypted with your Google password. This is less secure than Mozilla's approach because 1) your password likely isn't sufficiently complex or random 2) Google sees your password periodically (e.g. when you log in to Google services), meaning they possess the key to unlock your data. With Firefox Sync, Mozilla never sees your private key, so there is no way for them to see your data. Ever.

Google's business model means they have an inherent interest in your synced/private data. Mozilla has no such interest in it. Therefore, Mozilla locks the door and throws away the key.

indygreg2··on ØMQ: Mission Accomplished
Through versions 2.x, there were assert()s peppered throughout the code that could be triggered if malformed data was sent over the wire, making it unsuitable for internet exposure. There was also behavior in durable sockets where messages delivered to an identity that had disappeared would accumulate in memory, effectively manifesting as a memory leak.

3.x removes the asserts so you can't segfault remotely. 3.x also transparently drops messages on the floor destined for unroutable identities. This silent dropping sounds annoying, but you can work around it via a well-designed application protocol. Remember, 0MQ is effectively a transport layer, not a full-blown messaging system.

indygreg2··on Design mistakes in mixed C/C++ and Lua projects
There are some follow-up posts in the thread:

http://thread.gmane.org/gmane.comp.lang.lua.general/84469/fo...

To summarize, use the FFI library (http://luajit.org/ext_ffi.html) which enables LuaJIT to directly access C types and functions, bypassing the Lua C API. Or, write more of your code in pure Lua, reducing usage of the Lua C API bridge.

indygreg2··on Git 1.7.7 changes affecting the everyday developer
Blogofile works for me. http://blogofile.com/
indygreg2··on One Million Concurrent TCP connections
Technical details would be interesting. Until then, here's Urban Airship's post from last year on 500k connections on Linux:

http://urbanairship.com/blog/2010/09/29/linux-kernel-tuning-...

indygreg2··on Facebook privacy, Chrome extension for 2-clicks "like" button
For Firefox users, may I recommend ShareMeNot (http://sharemenot.cs.washington.edu/). It also nukes the tracking behavior of Twitter, LinkedIn, StumbleUpon, Digg, and Google +1.

Or, if you are an Adblock Plus user and want to kill Facebook on non-Facebook sites, add the following rule:

  ||facebook.*$domain=~facebook.com|~127.0.0.1
indygreg2··on Sprite3D.js: a javascript library for 3D positionning
It is worth noting that support for CSS 3D Transforms was committed to Firefox this week. See https://bugzilla.mozilla.org/show_bug.cgi?id=505115 for all the gory details.
indygreg2··on Remote Timing Attacks Are Practical
This was published in 2003: http://portal.acm.org/citation.cfm?id=1251354

That doesn't change the importance of securing against timing side-channel attacks, however.

indygreg2··on Dropbox Lied to Users about Data Security, Complaint to FTC Alleges
An identical file encrypted with key X will be a different blob from one encrypted by key Y. About the best you can do is slice the file into N byte blocks and dedupe on that (this is what some modern filesystems like ZFS do).
indygreg2··on "The 411 Parable": Goog411 vs Bing411 - Make sure you are playing the same game.
If I were the product manager of a search engine and our chief competitor had an offering (411) that we didn't (at least branded in our name) and the cost to deploy said offering was cheap (because we already had one very similar due to a recent acquisition), wouldn't it be worth the relatively trivial investment (from a large company's perspective) to deploy said offering? Even if it were for reasons of "me too," so what? I believe most brand managers would reach the same conclusion, especially if the costs were low.

Furthermore, does the track record not show that Microsoft re-brands (assimilates) its acquisitions eventually? How do you (or the author of the linked post) know that the deployment of a Microsoft-branded 411 service was not part of this process?

Finally, everyone in the speech recognition industry knows that it requires a large sample set (utterances) to refine a speech recognition engine. If you are building one from scratch (Google), you will do anything and everything in your power to collect utterances. To think that Microsoft and everyone else in the industry did not recognize what Google was doing with GOOG-411 is preposterous.

← PreviousPage 2 of 3Next →