Meta developer tools: Working at scale
engineering.fb.com
engineering.fb.com
Things like Eden were a great idea, but tools had all sorts of issues they gloss over (a virtual file system can be really slow if you have a ton of small files), dev environments would randomly fail a lot, really the only tool they had that nobody disliked was their log-searching thing (can't remember what it's called) but it was still lightyears behind something likes splunk.
It was my conclusion that "Wow you have 15 people working on a dev-tool compared to a public company with 100 building the industry-standard version over 10 years with actual product managers and UI experts, no wonder ours looks like crap... wouldn't it be cheaper just to take .01% of your salary and buy standard dev tools"
I guess "Clunky" is the word I'm looking for. "Blow it away and make a new one" was a phrase that happened with some regularity for dev-envs, repo-checkouts, etc. And iirc restarting your dev box took like >30min.
---
Side note - the other strangest thing was some of these tools people agreed were terrible (restart takes 30min). So you'd expect thousands of engineers to be swarming any system with any UI bug, edge-case, or whatever. But it just didn't work that way.
I also felt like a lot of the developer tools looked nicer than their google counterparts at the surface but had major reliability problems under the hood where you would need to do a lot of turn it off and on again style operations.
It’s possible to run a really disciplined Slack where updates go in a ready-only channel, and people broadcast their work appropriately - but it’s really hard. Most people default to non-discoverable DMs and then stream-of-thought in public channels, which are both bad for discovery in their own way.
Posts end up being a bit more considered than a slack message, but still lightweight, and more discoverable than docs.
Shame Google nixed their “Google plus for workplace”, I would love to see this workflow used more widely.
Mentioning splunk, I'm assuming you are using more paid industry tools. I suspect that opens things up a bit more, but realize that most of the industry doesn't use those due to prohibitive pricing. And I don't even know what the appropriate paid tool for building would be.
We seem remarkably primed to just hate whatever tooling we have at our disposal in ways that baffles me.
Regardless, many companies underspend by an order of magnitude on dev tools. If a tool makes you 1% more efficient, it's worth them putting 1% of your annual salary into it (multiply this by engineers at the company who need the tool). Soon you'll see that a tool costing hundreds of thousands to license a year is often a steal.
So you may say "oh splunk is overkill and unfair expectation" but in their CI/CD system, if your job failed (frequent occurrence) and you wanted to search for something in there you'd manually DOWNLOAD a 300 meg file to your machine and grep it for errors.
Basically a fang engineer's time costs $3 a minute minimum, so if a query could save 1 minute then the break-even cost is paying up to $3 to run that query.
More amusing is when you see someone build a giant Elastic Search pipeline to a Kibana dashboard so that they can get what would have been table stakes metrics if they had used the normal service templates. Without a ridiculously large bill.
Or folks that think they can warp around the high cardinality traps of making a metric out of everything. Assuming if you can make a system that works for the testing environment, of course it will work when you open the floodgates.
Seen people argue that sampling shouldn't be used, "because you could miss data?" Reservoir sampling is a thing, for very real reasons.
At any rate, I don't mean to just yell about splunk, from all I've heard it is nice. I am annoyed that folks seem to ignore the OLAP and OLTP divide, such that they think your metrics system should somehow be optimized for both. At the same time.
Pulling back to builds, though. What, exactly, is the gold standard in industry? I have yet to see it. Python builds, in particular, strike me as not good.
Hear me, nameless internet stranger, that you might learn from my mistakes: This is not true. Efficiency is only worth money if it increases profits - concretely this means your efficiency gain must result in the following:
1. Delivering features faster
2. Delivering features with meaningfully higher quality
3. Delivering the same features with lower headcount costs
(2) is hard to measure, so you can generally only sell on (1) and (3). During the 80s the business world learned the hard way you can throw away a lot of money on useless efficiency. If that interests you I recommend "The Goal" by Goldratt for a fictionalized account of those learnings.
* There's also a subtler point that corporate finances may mean that even if efficiency is perfectly captured, a 1% efficiency increase may only be worth 0.1% of your salary.
I’ve been through a few savings exercises where everything is gonna go open source dashboards etc. Once the guy who understands how it all works disappears, it migrates back to splunk.
One security incident response will pay for the entire splunk infrastructure.
If you can, that seriously awesome. The point is, many can’t.
Another thing is that sometimes it's truly impossible to use one-size fits all for some tools if your company scale is large enough. I've heard from friends at Amazon that the dev experience and tools are totally different in the hardware space than it is for those working in the retail or AWS orgs.
That's why big orgs are now opensourcing their tools. But it's not to clean up all "too internal" stuff from a tool.
Some of this is just having worked in the industry for a while. Guess what, 15 years ago "Put all your code in a web service running on AWS" was not nearly as slam-dunk a proposition as it is now; and tools optimized for on-prem were faster and more reliable.
The other thing is that you don't always have to reinvent things to the point that Meta does. You can just wrap around an Open Source project, or even contribute to it. This is what all of my previous jobs did -- make custom tooling from industry-standard OSS building blocks; and contribute back when it made sense for the larger community.
I actually think Meta was actually trying to do this with Mercurial a while ago [1], you probably have a much better idea than us as to why that didn't work out.
----------------------------------------
[1] https://engineering.fb.com/2014/01/07/core-data/scaling-merc...
It did work out.
One guess that I can make from Sapling docs is that perhaps FB/Meta needed an ability to pull in Git repos where needed; and I can see Mercurial devs not being super enthusiastic about that being a first-class workflow in Mercurial itself.
The needs of a large internal code base are often quite different than what open source projects need.
Anther issue is that tools like Git and Mercurial need to be compatible with positively ancient repositories and can't really break backwards compatibility whereas within Meta, it is easier to move faster and deal with breaking format changes since developers's checked out repos on their laptops can be silently upgraded since the machines are well managed with Chef and can be upgraded behind the scenes. In the outside world, you can't assume people will be running the latest version and can't upgrade them.
Finally, Google is also moving away from Mercurial. Perhaps the JJ developers want to respond why (and check out Jujitsu -- it's a novel and very interesting system y'all should check out).
Sure. I presented about that at Git Merge 2022. https://github.com/martinvonz/jj#disclaimer has links to the slides and the recording from there. I'll summarize the problems we have with Mercurial here:
1. Performance/scalability. That's partly because Python is slow. Both the Mercurial project and Meta have rewritten many parts of it in C and Rust as a result. Maybe more importantly, there are many assumptions in Mercurial's design that don't scale well. We have extensions for downloading only a slice of the repo. We slice it in both file space and in version space. However, it can get very expensive to change afterwards. For example, checking out an old revision that they user hasn't previously downloaded is very slow (it requires rewriting all local revisions after that point).
2. Consistency. Mercurial was designed for local file systems, so when we store repos in our distributed file system, we run into write races that can corrupt repos.
3. Integrations. We integrate with Mercurial by running the `hg` binary and parsing the output. That's unnecessarily complicated and slow.
We also see several opportunities by switching to jj (in addition to hopefully fixing the problems above):
1. Simpler workflows. Things like: working-copy commit (no "dirty working copy" errors, for example), undo, first-class conflicts (no interrupted rebases, for example). See the GitHub project for details.
2. Cloud-based repos. The repos will be stored in a database instead of being stored in files on top of a distributed file system. That makes them much easier for our server to work with, and it opens up for many kinds of integrations that were not feasible before.
3. Simpler architecture. We designed jj from the beginning to be easy to integrate with our internal systems, so there should be much fewer workarounds.
4. Simpler code base. You can typically add a command without worrying about concurrent commands, a dirty working copy, or conflicts. An example I like to mention is how I spent about two weeks trying to implement a command for amending into an ancestor commit in Mercurial. Then I implemented a more powerful version of that (can move changes from any commit to any other commit) in an hour in jj.
For many, the world has moved on a lot since, and those tools feel obsolete but so embedded in practices that it’s unthinkable to apply better approaches. One example of that is tracking data sources: legacy is slow, dysfunctional, and even fairly straightforward questions time out because you have to load so much meta-data. That’s because no one really uses it: most of that information is carried by the Data engineers who have built the system, and they are over-whelmed with questions that could be answer with a good system, but derive power over analysts. They benefit from it through comments they copy in their review, so fixing it isn’t anyone’s priority.
Others have evolved because the internal demand pushed things forward. PMs want Deltoid to work, so that system has moved to be state-of-the-art or interestingly unique in many ways: scaling, not implementing MCC, integration with the metric definitions.
> most of that information is carried by the Data engineers who have built the system, and they are over-whelmed with questions that could be answer with a good system, but derive power over analysts.
Is that the kind of question you're talking about here?
As someone who vastly prefers Slack over all its competitors, I'm very interested in hearing how Meta's compares. What makes it better than Slack for you?
However the combination of the "newsgroup" and chat makes it kinda better than slack on it's own.
Would I use workplace at a new startup? probably not. I would use slack though.
I’ve seen sub 500 employee companies and Slack kicks ass, but it seems once it goes over 1k it’s like glhf managing channels and bots.
Butterfly bot integration to workplace was really something.
Workplace def overkill sub ~ 600 users though.
On the other hand, building all these connections between the disparate apps is basically my job, so it’s got its positives.
Once your company reaches a sufficient size (maybe 300+ employees?), Slack channels just become so clunky and tedious for trying to keep up with projects and discussions. Using Slack keyword notifications helps cut through the noise, but I really think applying Facebook's ranking/Groups expertise to work-related content has been the best solution I've seen so far. You're no longer inundated with unread channels that may or may not have messages relevant to you, Workplace just surfaces the relevant content (and obviously there's chat/mentions for more pressing issues).
Early career this is pretty huge for growth. I was there until 2019 though, so not sure how it is now.
My point was that the cohesion of all these internal tools makes information discovery frictionless. cross-pollination of functional space in a business is like butter because everything is built by Meta and behind their intern tool. The other companies I've worked at just don't have this. They lack the Eng capacity to build it, unfortunately.
I personally loved things breaking at times. The monorepo and all the tools were so open that enabled me to follow along and try my own fixes in some cases (sometimes being the one to fix it!).
At that point in my career, that kind of exposure was like a rocket ship for personal growth. Others shared similar sentiments.
This is a curious thing to see.
As a mere user, and speaking purely about FB not any other Meta IP…
Other than unbreaking things that break when the OS and browsers change under you, has Facebook shipped even one thing in the last five years?
Don't get me wrong, I know it takes a lot of effort to stand still — forgetting the Red Queen effect was Musk's obvious mistake with Twitter even before stuff broke — but Facebook seems completely unchanged.
But I guess that implies you didn't mean the website itself when you wrote "FB" in the comment I replied to? :)
(One of the most well known consumer brands shipping tooling rather than stuff the consumers use directly is, ah, apt for a company called "meta").
They'd be just as fast with Slack or whatever, for the most part.
EDIT: as someone posted below:
> The open-ness of code, visibility, diffs. It was perfection. Something broke in my env suddenly? OH, I just checked recently pushed diffs that affect my realm. Hey there it is, security pushed something weird. I'll just revert the part that affects me and tag them. No meeting, maybe a SEV for visibility and review, maybe not. Easy peasy.
This is the kind of empowerment developers are given and expected to handle (both the explicitly ability to revert previously committed code and the implicit responsibility that the teams code that was overridden must deal with it once they're notified rather than push back)
To your point, other experiences have shown me the ego that rears it’s ugly head when trying to move that way in an env that didn’t have that culture. At FB there was a lot of candor, but in other environments I feel like I’m going to hurt someone’s feelings in code review, or even just watching what I say in slack messages. People take work too seriously in some companies. It’s so rewarding to have fun with it.
They used to say “Nothing at FB is someone else’s problem”
I hope they still do.
Meta cut through that by focusing on measurable improvements.
The downside is it can create silos, a negatively competitive environment. IE, teams not sharing resources or credit etc because they are only looking out for themselves at the end of the day, because any measurement you don't capture is one you can't claim as yours.
I also argue it can breed short therm thinking. Meta even had special teams from what I understand that were "exempt" from the typical metrics driven review cycles because they'd create the wrong incentives.
I think the general ethos can be really powerful though, but I'd peg it to collaboration and value driven measurements (and value is loose here, I'm not strictly thinking monetary) rather than strict "user based metrics".
These days at different companies it’s all planning and hashing out all those details before implementing, and then guess what? Surprises happen anyway. And in the end we become master planners and never actually build anything. Those blueprints sure do look great though.
Abstraction layers tend to be slow.
What they should focus on is plumbing. Take what exists and connect it in a smart way.
Take React: the idea of applying functional programming and spitting out HTML was a great idea. HTML existed before react. Then they also implemented a virtual-DOM, which was the unnecessary part.
Or react-native: rendering UIs in a functional way is great, but stuffing JavaScript into everything is not necessary.
But nowadays there are plenty of solutions that do not use virtual DOM (or even use a virtual DOM solution that is just plain faster) but still allow you to use declarative rendering, basically the best of both worlds. They are just not as popular or for some people not as interesting to them as React is. Or they are content with Reacts performance already.
Phabricator was pretty sweet... and the internal version of mercurial was a dream!
The sapling workflow is better than git. I was skeptical when I joined, but between no branches and excellent stack and merge support, it's just a better, more intuitive workflow.
Eden is fast. Crazy fast, in fact, if you look at the size of the monorepo. The comment about small files is odd to me. Source code are small files. The entire monorepo is just small files.
Now if we are going to complain about something, it would be the custom Android environment we've got. Android Studio integration is actually clunky, bordering on non-existent. Compared to writing vanilla Android code, which I have done for years before joining Meta, Meta folks wrapped almost every API, so all your pre-existing Android knowledge is useless. I've never been so unproductive when writing Android code.
For example for apps you may like to deal with abstractions that can express navigation and pages. They are the same across web, iOS, Android. Unlike Activity, Fragment, View (is that still a thing, I stopped doing android 10 years ago lol) and some slightly different set of abstrations over iOS.
When I was at Meta, if you could measure it, you could ship it and mark that as a win in your performance review. So a lot of projects got shipped off of A/B tested metric wins.
The entire system replicated the real system (not user data but the software) and you could spin up test addresses and accounts and all sorts trivially. Beautiful!
Why would someone need a more than one copy of a dev/feature branch?
It's slowly changing but I wouldn't consider Amazon's tooling to be industry standard by any definition.
I'm not sure, maybe?
When I say industry-standard tools I mean "Best you can purchase." So I mean a UI as good as github, a platform as good as AWS (meta's doesn't hold a candle to AWS), full-text log searching as good as splunk, chat as fast/searchable as slack (workplace doesn't hold a candle), video chat as fast/clear/clean as Zoom (workplace doesn't hold a candle)
Some problems are harder at scale (feature toggles interactions just get harder with size), but some of them are just a mess of their own making (restarting taking 30m, butterfly rules having >20m+ delays, eden not being optimized to work with buck, which needs to read tons of files out of tons of directories)
I agree the observability tools I interfaced within Meta was subpar a few years ago, but I think you're being ungenerous to take that and extend it to the coding tooling specifically.
https://engineering.fb.com/2020/12/10/developer-tools/probab...
> Ultimately, our goal with PFS is not to assert that any test is 100 percent reliable, because that’s not realistic. Our goal is to simply assert that a test is sufficiently reliable and provide a scale to illustrate which tests are less reliable than they should be.
??
The ticketing system had a rich tagging system (which GitHub has). The notifications system was completely unified and extremely useful - Workplace activity, relevant ticket updates, PRs to review, reviews, etc - all relevant and timely. From what I've seen, GitHub tends towards just being a useless uncoordinated firehose.
Finally code reviews. Having a code review for every commit is a bit to get used to but their tooling makes it super easy to post a stack (i.e. PR). Each commit in the PR can be independently reviewed & merged. This means you can land improvements progressively as code reviews complete (i.e. short cleanup commits go in faster while larger commits are reviewed longer) This also encourages people to size the commits appropriately & to be standalone.
Which one... there are like 3 (or more)? The ones I worked with were the single biggest cause of SEVs. I wrote some tooling around trying to make it better, but it honestly needed to be completely replaced.
>Workplace is a fantastic replacement of Wiki that surfaces relevant & interesting content to you.
Yea? I found it horrible to have to read through post-after-post-after-post to try and keep up-to-date with what's going on. Not my way of ingesting information.
>Their chat tools are way better than Google Chat / Microsoft Teams (not sure how it stacks up against Slack since I'm not a huge user of them).
100% disagree... telling my dog to bark at the neighbor because I want something from them is better than Workplace chat. Constantly switching between internal beta and production to try and get the features I wanted, with something stable. Go ahead and send me a link to that message in a chat... oh yea, you can't :-|
You’re measuring the wrong thing. The user actions are slow and take several seconds to complete usually. No one gives a shit how big the repo is. They care how long their operation takes.
I understand the tradeoff. That doesn’t mean it’s “fast”. It’s slow but the other alternatives are slower.
Though we arguably rely on a lot of Apache products, whatever we use that's internal only is great to work with. Daiq** recently started supporting notebooks, which has been a game changer for us as well as for the teams we work with. Phabricator is great as well, and makes shipping stuff super easy. Only Ben**, the internal notebook solution, I find meh. Especially compared to Google Colab. But the rest has simply been a joy to work with.
For those interested, a former DE made this nice repo that maps internal tools against "real world" products: https://github.com/thijsessens/xmeta2external
However, all the other tooling around code were just superior in every way. I miss Phabricator, and mercurial and the way it seamlessly integrates into a team. I also miss all of the command line helpers that let me manage all of it.
It used to be the case that legitimate teams could form around and focus just on building internal tools for other teams if it improved engineering velocity. Not sure if that's still the case in this new era of layoff-happy Meta, but it was definitely true when I was there.
s/Meta's tooling/Meta/g
s/inferior/very inferior/g
As a backend engineer, the dev experience was just incomparably better at FB. Some things I most most are probably phabricator / stacked diffs workflows, buck dep management with everything in a monorepo, deep integration across all the the tooling etc.
Curious if anyone has used Sapling after lots of time using git. Is it the future?
I used to work at FB, and while sapling is quite nice to use in practice, without the internal version of Phabricator to do code review (and, in all likelihood, mononoke), I don't think I'd pick it up again.
I've noticed that the system seems to repeatedly throttle me where it never did before. Logging in often results in being given an error until I try it enough times.
Sometimes attempting to pull recent changes from a repo will fail, and tell me that I pulled the recent changes even though I am stuck on an old commit. I tried several things to fix this and the only thing that seems to work is doing a hard reset on the latest commit id. I never noticed github doing this before.
I agree with you on the token system. It was a pain to figure out.
I've actually run into issues with throttling and stale fetches before, but it was due to a company proxy intercepting all requests. So effectively our entire org was hitting GitHub through a small number of proxy gateway IPs. Happened with AWS APIs too -- major throttling issues with terraform that seemed to be triggered based on the shared IPs.
Might be a longshot, but I'd check to make sure you aren't going through some proxy or VPN because I don't think your experience is typical.
Meta's Sapling (1) is definitely one of them. But there is also `jj` (2) and `git-branchless` (3). These tools target a smaller set of workflow where there is 1 main branch inside a big repo and everything else are short-lived branch/topic that could be treated as ephemeral stack of patches, constantly being uproot / rebase on top of the main branch to derive final result.
If that's the workflow you use daily, then you should give these tools a try.
(1): https://github.com/facebook/sapling (2): https://github.com/martinvonz/jj/ (3): https://github.com/arxanas/git-branchless
The UX just has lots and lots of polish in small ways and it has a nice amount of good features. It has fewer verbs than Git, but it's still a bit different from Mercurial. Having a built in 'undo' command that basically always works is nice.
I have replaced Git with Sapling (and a similar-but-not-the-same system, Jujutsu) in most of my own personal workflows at this time. If that's a good enough endorsement for you, then I suggest trying it out. You might be surprised.
Now, hg is amazing and when I leave I will be extremely sad to go back to git. The UX is well thought out, with commands mapping to operations (how do you undo a commit? hg uncommit vs git playing around with the reflog).
Amending deep commits is pretty good -- it's still tricky and absorb works on some pretty limited heuristics. But the merge tooling is pretty good around moving around in a stack, and the general UX over interacting with the stack is way better: hg histedit edits history, while hg rebase moves the stack around.
> While it is used in production within Meta, it currently does not build in an
> open source context and is not yet supported for external usage.
May be this is a feature not a bug? They want to filter out some Resume-Driven developers.
A bit cynical, but maybe not so unpopular opinion.
Meta as business is all about locking users to the ecosystem.
I know I’ll be downvoted for taking an ethical rather than financial or technological position, but ethics matter. Especially in technology and finances.
Better analogy augmenting yours: just cars. Accidents between cars or cars and pedestrians kill lots of people every year.
This works with every tech. Phones have made it easier to coordinate crimes. Printing press and even televisions made mass propaganda possible. There are always tradeoffs.
Facebook has done a lot of harm to the world, but it also has done plenty of good.
Aside from that, your analogy is incredibly weird. Do people drive drunk to... work? Who does that?
SWEs need to learn new tools all the time. This shouldn't be a concern. Besides, it's not like "the real world" only use a single tool. Each company has their own tweaks, even those using open source stuff.
* How do I setup a new project? * How do I build code? * How do I pull internal dependencies? * How do I publish artifacts?
Every company makes a horrible copy of these systems with off-the-shelf tools. At Amazon there were certainly flaws in the internal tools, but overall it was much better than the "real world".
The one thing that is super exciting to me is the stacked pull request support. Using Github for this kind of workflow is enormously painful. Conversations constantly get outdated, and its nearly impossible to track whether comments have been addressed.
I know they're working on an improved UX/experience there, but it seems like it'll be a good long while, especially for enterprise server customers.
git falls over and dies really, really badly when the repo gets stupidly large.
There's an article alluding to the discussion here: https://engineering.fb.com/2014/01/07/core-data/scaling-merc..., but I can't find the original thread on the git mailing list.
[0] https://web.archive.org/web/20210119051414/http://git.661346...
Stockholm syndrome. I've used Git for 15 years, early GitHub user, etc. Yes, you can solve many of these things, but until recently even things like "I am changing patch 2 in a series of 5 and need to rebase the following 3" were ridiculously painful. This is a common workflow many people like (including the Linux kernel devs) and Git was bad at it.
Git submodules. I'm not even going to go into this, they're so bad. That's a problem I wish Git had never "Solved" to spare us the burden.
There are tons of minor nits in Git all over the place. "Solving" something is completely different from actually having something that can be easily used for your team. There's no amount of contributing Facebook could have done to fix Git, because they'd be turning Git into something else that it fundamentally is not. And it doesn't matter if you have a trillion dollars, it's often not practical to just overhaul someone else's whole project when these goals don't align.
Is an intellectually dishonest fantasy invented for the sole purpose of using it to discredit and distract from criticism of the actions of the inventor of the phrase, so should never be ascribed as the source of a position you want to argue against unless your intent is to signal that your own position lacks a reasonable argument and you are just choosing to character-assassinate the opposition to cover for that.
I am not a fan of FB, but they tried - you can find them on the git mailing list where they got told they are doing it wrong for things like "scaling" or "productivity". Which is always ironic since basically nobody in open source generates or uses any real data about productivity, it's all just gut feelings about users.
I used it briefly but couldn't get most people to adopt it widely enough.
So, no. They wont support Phorge. They really never supported open-source Phabricator after Evan left and made it his own with Phacility.
Talking with a Meta person, it sounds like Phab really needs Mercurial to work well, at least for Stacked Diffs because you need to be able to identify commits independent of their location in history to properly maintain the Stacked Diff associations.
I'm a bit worried about the dominance of VS Code. I can't stand the editor and it's popularity just grows and grows.
- refactoring (extracting functions and variables, renaming across entire project etc.)
- context aware selection (expand selection from cursor logically)
- stack trace / error parsing
- comparing anything (diff selection with clipboard for example)
- DB schema support even in SQL formatted as strings (e.g. when using psycopg) directly from DB by just connecting to a DB
- find in files (I've never seen a VS code user have an easy time finding stuff)
Etc.
In our case, up until Dec 2022, parts our pre-acquisition monorepo were still separate, but gradually components and workflows, such as tests and reviews, were moved into the main repo. From my awareness, we didn't even start merging into the main monorepo until more than a year after we were acquired, though of course there were exploratory efforts before.
In general, I'm pro-monorepo, it makes sense to be able to update multiple interconnected components in lockstep. For the startup, we were still in research mode, so it was less urgent to spend eng/sci/TPM to incorporate with anything on the FB monorepo side...until it was.
(I am the author)
- Where are you going?
- I'm going on vacation.
- Have you finished your project?
- Not yet. Just submitted the diff.
[^_^]
It's funny that the Sapling command is "sl", that's going to conflict with installations of "stream locomotive".
The tl;dr is that Google's developer tooling and ifnrastructure is superior in almost every way. Examples:
- When I started at FB we used Nuclide, an internal fork of the Atom editor. While I was there it was replaced by VS Code. It's better but honestly they should've built their tooling off of Jetbrains products. Jetbrains make IDEs. VS Code is a text editor like vim or emacs. There's a massive difference;
- Buck should've been killed and replaced by Bazel. I can't speak to Buck2 but this seems like a pointless investment;
- Thrift should be killed and replaced with gRPC/Protobuf. Same deal;
- FB's code search is just grep. It's literally called BigGrep. Grep can get you pretty far but it's just not the same as something with semantic understanding. Google has codesearch, which does understand code, and it's miles ahead. This has all sorts of weird side effects too, like Hack code at FB can't use namespaces or type aliasing because then grep wouldn't be able to find it. When there were name conflicts you'd sometimes be forced to rename something to get something to compile;
- Tupperware (FB's container system) is a pale shadow of Borg;
- Pushing www code at FB is a very good experience overall. You commit something and it'll get pushed to production possibly within an hour or two or, at busier times, it might take to the next day. This requires no release process or manual build. It's basically automatic; Google's build and release process tends to be way more onerous;
- The big achillees heel in FB's www code is that it is one giant binary. There's no dependency declaration at all. This means there's an automatic system to detect if your change affects other things and that process often fails. This leads to trunk getting broken. A lot.
- Because of the above problem there is a system to determine what tests to run for a given commit. This is partially about what the affected components are but also longer-running tests aren't run-on-commit and often those tests would've found the problem. There is no way to say "if this file is modified, run this test". That's a huge problem;
- FB has a consistent system for running experiements and having features behind flags (ie gatekeeper). This wasn't the case when I was at Google. It may well have changed;
- Creating a UI for an internal tool or a new page is incredibly easy at FB. There are standard components with the correct styling for everything. If you want to write an internal tool, you can start at 9am and have it in production by noon if it's not terribly complicated;
- The build system for C++ at FB is, well, trash. For Buck (and Bazel), the build system creates a DAG of the build artifacts to decide what to build. FB C++ might take 2 minutes just to load the DAG before it builds anything. This is essentially instant at Google because a lot of infrastructure has been built to solve this problem. This is a combination of SrcFS and ObjFS. Incremental builds at FB to run tests doesn't really work as a workflow;
- All non-www builds at FB are local builds. Nothing at Google (on Google3 at least) is built locally, including mobile apps. This is way faster because of build artifact cachcing and you have beefier build machines.
- There tends to be less choices as to what to use for FB code (eg storage systems). I consider this largely a good thing. You will typically find 5 different way of doing anything at Google and then need to consider why. You will often find different teams solving the same problem in slightly different ways or even the exact same way.
- There are people at FB who work on system-wide refactors (eg Web security, storage). These people can often only commit their diffs that might touch thousnads of files on weekends.
- A lot of generated code is committed at FB that isn't at Google. This exacerbates the previous problem. FB has a ton of partially and completely generated files that mean a change to the generating code has a massive effect. At Google, for example, the protobuf generated code is genearted at build time and isn't in the repo.
There's probably more but that's what comes to mind.
Pedantically, codesearch is also grep. But codesearch calls out to Kythe (nee Grok) which has a semantic graph of the code.
But.... internally everyone thinks the Kythe team is just codesearch anyways, so that's about right lol.
Kythe is partially open source, but critically a lot of the postprocessing to get it to work at massive monorepo scale is not, so FB would have a bunch of work to do to replicating it.
Also we don't have a PHP indexer, because nobody's written one.
What's nice about this scheme at Google is all manner of generated code is indexed, so you can navigate up and down the caller-callee graph between artisanal code and generated code. Really works well.
Even then, it's not like rewriting a billion lines of BUCK files to use Bazel was even practical. Realistically any solution had to have a direct migration path from Buck1 without rewriting everything. I don't even work there, this is just pretty obvious from the design constraints and talking with the dev team, though. Frankly, I'm pretty impressed they were both able to meet the goals they had (migration from Buck1, better performance, more extensible), while still jam-packing the thing full of good design decisions and features like they have. It's good work.
(I am a pretty happy, non-Facebook user of Buck2 already, FWIW.)
My issue is more organizational. Meta, as a company, in my experience, does not put in sufficient investment to build a mature, robust open source project with few exceptions. Even things like the www test infra, which are core to the company when I was there had ~1 FTE SWE. That's not open source but you get my point. Buck (and Thrift) seem to have been woefully underinvested for years. Thrift was originally an intern project at a time when Stubby (Google's protobuf-based RPC) was not open source.
Buck2 is much faster than Buck and pretty great for the set of problems it's trying to solve. (I'm skeptical that the corner that Meta has painted itself into is good, but assuming that that can't change, buck2 is great.)
Remote builds are common now, from what I've heard.
Sapling is the name for Meta's fork of Mercurial. Piper is not a reimplementation of Perforce, by the way. The closest equivalent that Meta has to Piper is their source control server Mononoke.
Yes, it is. You are probably confusing Piper, the internal name for the VCS, with Google Piper [1], which is a completely separate and unrelated project. That project seems a closer match to Monomake.
This was a common problem with Google, actually. There would be an internal name but once released it would take on a different name. Sometimes that name would conflict with a different internal name, which then made internal searches impossible. I forget the specifics, but Buzz became an external product but was an internal product name for something different.
Dart (the language) was internally known as Dash prior to release. IIRC Google Dash was an external thing for advertisers or something like that.
So when I say Piper I really do mean the internal rewrite of Perforce.
> FB's code search is just grep. It's literally called BigGrep.
This hasn't been true for a long time, codesearch at FB is more complex than just grep and has some semantic understanding. Here's some discussion here when some of this infrastructure was open sourced: https://news.ycombinator.com/item?id=28365880
> FB C++ might take 2 minutes just to load the DAG before it builds anything.
Yeah, this is pretty terrible, but is mostly a description that doesn't apply to buck2.
> This is essentially instant at Google because a lot of infrastructure has been built to solve this problem. This is a combination of SrcFS and ObjFS.
SrcFS and ObjFS aren't what solve this problem at Google. And to the extent that they do, FB's sapling and buck integration do the same.
> All non-www builds at FB are local builds. Nothing at Google (on Google3 at least) is built locally, including mobile apps. This is way faster because of build artifact cachcing and you have beefier build machines.
This is wrong since at least 2015 FB's build system had the build artifact caching, and you can see in buck's git history that they've had remote execution just like Google (in fact, using the same RE api as bazel) since like 5 years ago.
I did spend way more time doing www than C++ though and Hack was specifically written to facilitate regex searches like being able to search for "SomeClassName::someFunctionName" to find usages as well as the other examples of prohibiting namespacing and type aliasing (in Hack). Google cs doesn't have that constraint.
> SrcFS and ObjFS aren't what solve this problem at Google
It's a mix. You can't just pull out one piece of the Google dev infra because it's all connected. In a P4 client, you'd list some paths that were "local" allowing local modifications. Forge, SrcFS, ObjFS, TAP and Sponge are all pieces of this puzzle.
> This is wrong since at least 2015 FB's build system had the build artifact caching
First, there's more to FB builds than infra C++, most notably iOS and Android, which are all built locally. It's why iOS/Android engineers have big, chunky machines like the iMac Pro or the trash can. If this has changed, it's a fairly recent change. Google builds mobile apps on Forge. There are literally racks of Mac Minis to build iOS. With this you can build artifact cachincg like you do with, say, Google3 Java or C++.
Second, "local" requires some further explanation. Typically, things are built on a devserver (although this was transitioning to on-demand VMs, for which the Google equivalent was CitC). But a devserver build required a full checkout and build with artifact caching. It could then be incremental until an hg pull forced a larger rebuild. There were sparse checkouts but I think the support was pretty limited and only worked on certain infra projects.
Either way, the whole FB C++ build experience was fairly primitive and even worse for iOS/Android.
Again, ios and Android have had the remote artifact caching you mention as so important since many years ago. Android has had remote execution for years (I think even before the infra c++ you mention).
> But a devserver build required a full checkout
This is not true with sapling, which had been used extensively for years.
Flatly untrue, on iOS much of the infrastructure was well ahead of what anyone else had aside from Google. It worked so well that a tooling team responsible for upgrading to a new version of Xcode soon after its released declared victory and took credit for the compiler upgrade, as this had been an issue in past years.
They hadn't realized that a compiler team had been quietly running their builds company-wide for nearly two years and fixing compiler bugs on the bleeding edge of clang/llvm the whole time, so by the time the compiler was branched for Xcode and released from Apple, the fixes were already complete, open-sourced, and merged fully upstream.
I skipped over that parenthetical in first read, but my God, it's not even wrong. I mean citc is excellent and all, but describing those as equivalent just makes me wonder if you held some non technical role and are just playing telephone from people who better understand things. Like I'm just imagining a conversation where you're like, how does Facebook handle this specific thing that's handled by a X at Google (or the reverse)? Getting an answer on one side that involves citc and the other that involves on demand and deciding that that specific problem represents the entirety of those two ecosystems.
I guess one of the difficulties I have here is that a lot of what you've described across a bunch of google-facebook comparison is just like so factually wrong that i can't ignore the problems with a direct reading of what you are saying. At a deeper level, I think I'd totally accept google's CITC and facebook's on-demand as like philosophically similar in that solutions to seemingly unrelated things fit into the structure enabled by those systems. Again though, I think crediting you with trying to discuss things in those terms would be too generous given that that equivalence is no more true of on-demand VMs than it is of devservers.
> I'm honestly not sure what sapling is
My rough understanding: arcanist + git evolved into arcanist + hg, hg became a frontend to cHg and then the whole edenfs/sapling stuff started to replace that to optimize sparse checkout workflows.
> - Buck should've been killed and replaced by Bazel. I can't speak to Buck2 but this seems like a pointless investment;
Bazel's open-source version remained significantly less capable than Buck for a long time; Buck migrated towards Skylark just like Blaze, and cleaned things up greatly.
> - Thrift should be killed and replaced with gRPC/Protobuf. Same deal;
In general, agree with the former, but that's because protobufs and especially flatbuffers were already used in many places for many many years.
> FB's code search is just grep
There was also a less popular semantic search that people didn't use nearly as much.
> There's no dependency declaration at all
Strictly speaking this is not true, it was just not usually needed.
> There is no way to say "if this file is modified, run this test".
You needed to update the target determinator, but I do think this was possible.
This is a solved problem, and even if it isn’t, there’s entire communities dedicated to it. There’s literally no reason why Facebook needs to be in the business of reinventing the wheel from scratch beyond some dude trying to show “impact”. It’s a distraction from core business problems.
Hosting a gigantic monorepo for 25K concurrent users is so far from a solved problem.
There are very fundamental differences between a mono-repo and a bunch of repos for each "service" or whatever. Lots of tradeoffs. I've worked both and I can see the reasons for huge monorepos. They make a lot of things that were previously hard much simpler... The tradeoff is your tooling needs to be able to scale with growth of the company. And for a company the size of FB, dedicating an entire team to improving the tooling for their monorepo is well worth it.
Saying that Meta "reinvented CVS" means you either know very little about source control or you are just being purposely misleading. Neither is a very good look.
It is a fork of Mercurial and there wasn't a community dedicated to making it scale to the scale Meta was reaching hence why they invented resources in making Mercurial scale.