Why’s that company so big? I could do that in a weekend (2016)
danluu.com
danluu.com
I think it's hard to argue (although some people clearly did) that making a search engine is easy. You need to index all of Internet, cache it maybe too, rank stuff, filter out spam, work over different languages, do something about images, parse a lot of (broken) HTML, maybe run some js to figure out what a dynamic page should show, then serve it up super quickly... google is faster than pinging a local (!) NFS server at my work.
But then I remember a comparison on HN about how booking.com is O(50) people and Airbnb is O(2000) total. Or I remember the total headcounts at Twitter, an app that lets you post very short messages.
Or LinkedIn, a vast company with a very subpar UX, basically an inferior clone of a dozen of other social networking sites, only kept afloat by its network/monopoly effect. I mean, searching for people by name barely works! You're better off using Google to look for people on LI that its own search box!
No doubt it's just my cognitive limitation, like I can't visualize the vastness of the Cosmos, but this is a close second.
100% agreed.
They really should have used Twitter as an example. Any competent full-stack developer could probably have a working clone of Twitter running in a weekend.
The challenge I think is scaling it to millions of daily active users.
But, even then, Twitter has been around for over 10 years. The scaling problem seems to have been solved. I would think that these days, Twitter engineering would just be a whack-a-mole of odd obscure bugs. I can't think of any useful new features it's gotten in the last two years. Why do they still need so many engineers?
But the point is that this product is not what makes the company money.
If you only consider this, sure it’s a easy to clone product, backoffice for the advertisers, algorithms to maximize engagement, external integrations, marketing, PRs, partnerships are the thing that brings money home.
Of course if there was no Twitter, I imagine you could earn some money/be profitable with just the public product and some basic backend for ads publishers. But you’ll probably encounter a revenue bottleneck and start researching fresh income which you can do by … becoming Twitter, the big company.
The obvious differences being the scale, not really having to worry about the complexities of running your own data centers and overlay networks or whatever, and the fact there is no advertising, no sales, so none of the clutter that comes from trying to optimize that as well as defeat your user's attempts to subvert being tracked, nor any tracking of what you're doing off-site.
The signup process is really nice, though. You don't need individual accounts on each site. They pull your identity automatically from your IC PKI client cert and create the account for you when you first use it.
I work from home now, so I'm not on there any more, but they called the Twitter clone "eChirp" and the Facebook one "Tapioca." They also clone Wikipedia and Stack Overflow, but those are just read-only mirrors that get updated once a day.
What's a tweet? We can haggle this up and down so cut me some slack. 200 bytes? 50 bytes? gzipped -90%, 5 bytes.
Holy shit what absurd amount of traffic could a 1gb fiber connection handle?
oh huh 200 million tweets a second. That's silly. How wrong could I be? -- then you run down the numbers and you realize that no matter how wrong you are, a twitter-clone-for-most-people oughta be cheap. The network effects and fuzzy human reasons are their moat, nothing else.
One of those fuzzy human reasons is how amazon EC2 eats your ass. Just gobbles it from stem to stern.
Where I see your engineering headcount blow up is in moderation, legal and advertising; which I think everyone skips over when doing their mental math. Campaign management, targeting, reporting, and optimization on such a large corpus of data is where I struggle to see even a team of 5 building that in a couple of weekends.
The text part is basically solved at this point, but not quite that easy. I hear dealing with blocks is one of the slow parts.
I hazard there's a verdant, halcyon prairie between the two zones. The common zone of maximum extraction, and the hypothetical 200$-a-month-operating-costs text only twitter that can service 5% of the entire planet.
OK, but you aren’t just pushing raw ASCII across the wire, you’re getting those results out of a database.
For a weekend clone, someone will probably just use an RDBMS (a very reasonable choice). I’ll be generous and say they can retrieve a given tweet in sub-msec time, say 10 usec. Suddenly, you can only send 100,000 tweets/sec, ignoring all other overhead (which of course you can’t).
I’m sure there is unoptimized code and some unnecessary bloat, but fundamentally there are limits on data retrieval that are very, very hard to get close to.
Think about TikTok: It’s essentially just an app that lets you scroll through videos, yet it became breathtakingly popular just due to the fact that they seem to know exactly which videos you may want to see.
And then how are you going to ensure you don't provide content rabbit holes that can be harmful to mental health. Then detect credible cases of people threatening suicide to get a welfare check run.
That's just one of the things you don't see on the surface, there are hundreds more functions like that, that you'll very quickly realise you need to do so you don't get fined out of existence and/or end up with execs going to jail.
I'm an AppSec engineer. I don't write a lot of code, and when I do, it's entirely scripts and occasionally modifying existing code (both front- and back-end, but more often back-end) to implement security fixes.
That said, I could probably hobble it together. It'd be ugly as sin with hand-written Web 1.0 style HTML and likely server-side rendering since I hate writing front-end code, but it really wouldn't be hard as long as you're keeping it simple - Just text posts (No video/pictures, though maybe that wouldn't be too hard to add?), no algorithm for engagement (just chronological ordered tweets), no analytics to determine what's trending. Obviously no mobile app.
That's not to say they are or aren't staffed right, just that the part we are conscious of daily isn't the entire thing.
One thing to ask yourself - is the company chasing growth? A company like AirBnB or Twitter could absolutely run on a much smaller team if they agreed to mostly abandon growth. They could focus on reducing tech debt, make architectural decisions that reduced flexibility in the name of reliability and lower maintenance, and then eventually lay off a huge portion of their staff.
Instead they have a growth thesis - that growth requires investment, having to favor velocity over unit economics, etc. The higher the growth target, the more inefficiencies you take on.
I can be critical of Google at times but I'll never actually properly come to terms with how Google's search index is still so fast.
Does anyone have a link to this?
BKNG: Employees: 23,100, Revenue(2022): 17.1B, Net Income(2022): 3.1B
ABNB: Employees: 6,811, Revenue(2022): 8.4B, Net Income(2022): 1.9B
Making a search engine that is good is actually a reasonable task as shown here on HN. It doesn't seem to take very many people, either. Google has set the bar really low.
By contrast, making a search engine that is monetizable and good seems to be an impossible task.
It's a long story as to why, but I actually ended up building an installer engine in Powershell some years later. I can now officially say all those people are morons, and you definitely should not in this day and age be rolling your own installer engine.
Microsoft promoted it strongly, and mysteriously they never come up with a decent installation support, even for basic desktop apps.
Answering the topic here: it is the business more than the product. With exceptions, as usual.
Windows installer came along and was a lot better. The InstallScript engine carried with it a lot of legacy tech-startup technical debt.
Fun fact: InstallShield became a general purpose packaging tool at some point which spit out *.msi packages (and had a bunch of extensions to Windows Installer engine via custom actions). So it was pretty awesome from that perspective.
In this day and age Windows Installer itself has a lot of legacy design cruft from the era it was designed for - the late 90's.
So you're not wrong, but based on the era InstallShield was more or less crappy depending.
Sorry, but I want to add one more thing: the upgrade logic of MSIs is completely wrong, you need a lot of tricks to do an upgrade.
Patches though... Fuck. I think I could still produce a document that explains everything you need to know about patching but it is a lot more complicated than regular upgrades.
At the same time, my career has moved to Cloud stuff now. Is my life better? Definitely not.
Well, when your customer is an enterprise who wants to deploy your product hands-off onto five thousand machines, and expects a clean uninstall when they want it removed later - yes, it's that hard. Writing installers really is it's own sub-speciality in software development.
The testing cycle was the biggest pain, this was right as VirtualBox had become a thing so at least starting with a fresh system was slightly easier. (The entire company had 4 engineers + some researchers).
I think I was two years post-college as were most of my other colleagues... I'm fairly certain that product never made it out of initial trial (I left the company before it finished).
That sounds horrible. Why didn't you use file based HSQLDB or H2 (or any other pure java sql DB)?
I can't remember exactly why HSQL/H2 wasn't used but I think it had to do with fault tolerance and redundancy and what we could do with postgres replication was preferred over what we would have otherwise used SQLite for.
I was always very impressed by it.
If the application is a single .exe or something that works fine when you unzip a .zip wherever, then the installer functionality needed is almost trivial.
My limited experience is that most complexity in installers comes from the system not helping making it comfortable. Like something not having a simple to use version number, but having to check for 5 registry keys instead, or some such.
However, if someone just claimed they could write a better installer for one specific program of theirs only that was better for that one specific program than an installer built with one of those general engines, that is not so unreasonable except for the in a weekend part.
I once worked for a DOS and Windows utility software company that wrote an installer for each product, and I was usually the guy that ended up doing the installer and it wasn’t difficult. If however they had decided to try to make an installer engine product for other companies to use I would have worked hard to not get assigned to that project.
Reference counting (which could sort of be rolled into the above item), to make sure you're not removing something someone else is using, if you install stuff to a shared location.
Silent installation. Yes, this is an afterthought for a lot of devs, but a lot of techs expect it for rolling out software.
Extensibility / Customization at install time. Also ties into the previous one (how to configure the install with no UI running).
Resume-after-reboot functionality, to update files which may have been locked.
Publishing accurate information into Add/Remove Programs (or "Programs and Features" as they call it these days).
Handling elevation in cases where the user isn't Admin (or post Vista - the installer is launched under the non-admin token).
"Repair" functionality, which requires maintaining the currently installed state and configuration.
Support for packing up and deploying redistributables (maybe even download support).
You start adding all that stuff up, and you realize - you can't bang that out in a weekend. It's a LOT of work actually.
https://groups.google.com/g/alt.games.half-life/c/15MUk17ppH...
* How would we install all these toolbars?
* How would we change them in flight based on characteristics we can get from elevated permissions during install time?
* How can we bundle adware with the installer so it works without an internet connection?
* How can we subvert the Microsoft Store signing algorithm to add arbitrary unsigned payloads at the end of the executable?
* How can we run unsanboxed internet-connected javascript with elevated administration privileges and direct access to Windows API? [0]
* INCOHERENT QUESTIONS ABOUT METRICS?!
What is "repair" for?
What exactly does it do, and what's the intended use case?
I rarely do much on Windows these days, but my impression was that the age of random applications dumping DLL and OCX files into system32 was fortunately long over, so programs stepping over each other should be almost nonexistent at this point.
Story time regarding DLL naming from the old days - Windows 95 would only keep one copy of a DLL in memory (the days when having 8mb of RAM was amazing). We had developed an application that worked perfectly, except on this one salesman's laptop - where it would only work occasionally. What I found was they had installed "Tiger Woods Golf" on their machine and it had a DLL with the exact same name as ours. If the salesman played the game before running our application, Windows said "Oh, I already have that DLL. No need to load it again." and our app would blow up.
Turns out many players still had floppies. Whoops.
https://learn.microsoft.com/en-us/windows/win32/api/errhandl...
I only put "correct" in scare quotes because it is kinda lame that this is still required. When was the last time anyone was happy that they got an "Abort, Retry or Cancel" dialog?
In the ‘80s and ‘90s one of the most popular PCs in Japan was the NEC PC-9800 series. These were x86 machines but were not IBM PC compatible. MS-DOS and Windows supported them, but didn’t hide all the differences from IBM PC compatibles from apps.
In particular, the boot drive was A: so a system with two hard drives and a floppy would usually have A: and B: be hard drives and C: be floppy.
That really confused a lot installers that assumed hard drives always started at C: and that A: and B: were always floppies if they existed.
* Getting real deep with detecting OS/instruction set and edge cases
* Constantly validating permissions on every directory and file
* Constantly verifying checksums on everything put in place
* Concurrency controls to make sure the user didn't launch the installer twice, or the system wasn't live running when being reinstalled.
* Dependency verification was it's own rats nest of problems
* Uninstalling
Easier problems: * Logging
* Status tracking (except for really large files things get weird...)
* Aborting/Cancelling installSo - do you want to log variables for debugging purposes once installers get to a certain level of complexity? Great. Now you're logging people's usernames and passwords, and you have to add some functionality to not do that.
Or homebrew which is really just dependency management + standardized directories for files which is what it should be. Installshield was a solution for an entirely artificial problem.
There is a reason Mac had such tiny market share for so long, especially among enterprises. Installation was among pros and cons, but I would advise you to come off your high horse and understand the rest of it. Some reasons technical, some not. For a long time enterprises shunned Mac for extremely pragmatic reasons.
I think the last installer I used was probably the Linux Mint/Ubuntu live CD installer?, But other than that I don't think I've seen an installer for a couple years. Albeit I'm in the Linux ecosystem now so maybe that's why
The model on Windows is that you ran installer program that installed the application. Installers had to do a lot more stuff than drop the application in the right place like change registry entries or write files in other place. InstallShield wrote the installation engine and build tool.
Windows moved to model of .msi packages with .exe wrapped. InstallShield turned into an .msi build tool. Macos has .pkg installer for things that can install .app. Linux has .dpkg or .rpm. All of them have the installer pre-installed.
https://en.m.wikipedia.org/wiki/InstallShield
MSFT released MSI 1.0 in 1999.
https://en.m.wikipedia.org/wiki/Windows_Installer
Post 2008 or so, there was a very good question of why anyone would use anything other than *.msi on windows. But most Installshield users by then were building those packages with InstallShield by then.
These days, lightweight apps and dev tools are still preferred to use *.msi or *.exe based formats instead of windows store apps. Partly because it is quick and dirty to do so, and cheap using things like WIX and NSIS. But InstallShield still has a legacy base among less sophisticated users; it is easier than WIX and more robust than NSIS and the like.
Windows Store package formats are superior in most ways.
Each type of entry in the config file had a module which knew:
- How to check if the entry was already done and in a good state
- How to uninstall cleanly the thing (not actually needed; this was for system imaging)
- Usage of variables that would dynamically configure the system
- Variables also could conditionalize what gets installed
At the end of the day it was somewhere between a true installer engine and Powershell DSC (DSC would come about a few years later). This installer is probably no longer in use, as you can imagine more robust ways of solving this need exist now.
It was a heck of a lot of fun to build and watch it working though!
There are reasonable discussions of this sort, but the example OP goes on about is kind of a straw man. Just because anyone can go on the internet and say anything ("I'll build google in a weekend!") doesn't mean it needs a whole rebuttal that takes it seriously.
edit: just realized this was 2016. a purer time.
The same logic applies to bugs on more and more obscure setups, or more and more niche features.
In my experience, there's a huge difference between "software that works," and "software that ships."
Most of my test harnesses are full-fat applications. They are written in a ship fashion, but they don't have the polish and TLC that I lavish on the ship stuff.
I can write a test harness app in a day or two. It will be a fully-functional, accessible, localizable app, with a non-trivial feature set.
However, if I write something to ship, it may take a couple of months (or longer).
An example is the app I'm writing now. I wrote the "heart" library in about a month and a half. It has a test harness that runs every feature, and the test harness was the largest part of the job.
However, the ship app is coming to beta, about a year later. It knocks the test harness into a cocked hat.
The real issue for the bloat is simple lack of engineering skill. Most "engineers" have very little clue of what goes on under the hood when their code is running, and really only rely on learned patterns to complete their job. So many times I have been given code for a service that when run locally doesn't work, with the explanation that you have to test it in beta stage - the reason this happens is because nobody who started writing that code has any idea of what goes on outside of the boilerplate that is provided for them to fill in.
10% idea, 90% execution. Writing code over a weekend is the “idea” only. The rest of the work takes years of disciplined execution.
The world is full of people that worked really, really hard on something and failed to achieve the specific goal they set out for.
Success is complicated.
For example, professional networking is a tremendous amount of work. But in meeting, bonding and potentially working with those people opens so many possibilities for serendipity / luck what have you. Sure there are those with Daddy's money, but more importantly the relationships, networks and resources that Daddy has are more important than just the cash.
Nepotism, sure, but that's how the society has pretty much always worked.
Google was a far superior product at the time they launched. They captured lightning in a bottle and participated in - and accelerated - a big inflection point of internet adoption; the web went from nerds to mainstream as the younger, web-native generation grew up.
The technology is relevant insofar as it gave them an insurmountable lead at the time they had it. Re-creating the initial search technology itself pales in comparison to re-creating the initial conditions that allowed Google to reach millions, and then billions, of internet users in a relatively short timeframe.
This is a fantastic point.
Stable Diffusion has made a state of the art software with less time and engineers than OpenAI
Let's stop talking about hypotheticals, the examples are many
Google has its flaws; but this is manifestly false _even_ if you restrict "Google" to just Search.
More specifically, new entrants (and especially so in new categories) regularly get the luxury of extended time to iron out issues, create a better UX, find all the edge cases, invest in internationalization, distribute to multiple platforms, integrate with other systems, address security issues, and more. These tend to be exactly the kind of things that make recreating the incumbent products hard.
If OpenAI continues on its current trajectory, it can easily end up with a team equal to or larger in size than all of Google working on search.
Both are fantastic, but not the same.
Based on the theoretical foundation from Google, the framework built by Facebook, hardware produced by NVIDIA and cloud infrastructure run by Microsoft? This doesn't make a great example.
A massive social campaign was ran to make twitter seem like a hateful place which advertisers should stay away from.
If that hasn't happened, for the most part twitter runs just as well as it did prior.
Twitter lost advertising because its a bad outlook, not because of technical issues.
Twitter still works. Sure, it has issues, but still works. It definitely was a bloated company, and if Elon could keep his autism under reigns, he could have came out with reducing the company, keeping the same advertising income, and being more cash positive.
That's all your imagination - we can make and support any claim that way. The actual facts we have point the other way.
Let's stop making Musk look like a victim. Nobody blames autism when Tesla's sales increase.
Well, they had 3x or 4x more income and market value than Twitter with fewer employees. That's quite a lot to show for it.
Elon hasn't realized any losses yet, while on paper the company may be worth less that game isn't over for him.
He hasn't done himself many favors by stirring up controversy with his own statements, but it isn't surprising that revenue for an ad-based business takes a hit when the brand is raked through the coals by a large enough group of people making noise. That could be one step towards it failing, or it could be one of many dips while the company is not rebranding itself and changing their business model.
Hindsight is 20/20. And there's plenty more to product than the build.
Put another way, and I say this often: "Making it look easy is very very hard."
If only more fools understood this, there'd be less fools.
Luck does happen, sometimes people miss seemingly obvious ideas.
Almost 99% of the cases this is true. The question is whether this complexity is actually valuable or just accidental one accrued from path dependency.
Flunkies - people hired to make their boss look important, because said boss has a lot of staff in their department. Their purpose is ultimately managerial infighting, which is zero-sum.
Duct tapers - people whose job is to inefficiently work around problems caused by solvable social issues, e.g. a person whose job is printing out emails because $EXEC is old and refuses to learn how computers work.
Box-tickers - people whose job isn't to do something, but to be seen doing something, so that company can claim they're trying to do X when they don't give a shit about actually accomplishing X.
I'm skipping the others, because I don't understand them well enough to explain them in my own words.
Of course, someone (like Elon) might believe that they have such magic wand, but Twitter is losing more money than ever so probably it was not the case...
Turns out painting is easy if you don't care if it looks like crap when you're done. But if you want it to look good, you'll spend 90% of your time on prep: Moving out or covering the furniture, laying down and securing a tarp on the floor, taking down all the pictures, scrubbing the walls clean, patching holes in the drywall and retexturing, and then carefully masking around the edges of the walls with tape.
90% of the remainder of the time is stirring your paint and carefully painting around all the edges with a trim brush, going extra slow wherever there's a moulding edge that couldn't be masked. There are tricky roller guides that make this faster, but it's still easy to make a mess that requires laborious correction. If you screw up these edges they will ruin the appearance of the room.
The next bit of the job is rolling on the big areas of paint in the middle of the wall, which is what everybody sees first. That only takes 5 minutes.
Then finally you spend a couple of hours cleaning everything up and replacing the furniture.
How hard could it be, indeed.
The difference between being able and doing is infinite.
If you did it today, you'd probably want a fairly literal-minded search engine front-ended by a large language model to make it more user-friendly. That might be easier than trying to make the core search engine appear semi-intelligent.
The principals were then re-hired by Google.
That is shocking!
However, I'd like to point out that there are rare cases when some newcomers can replace existing businesses because they found some technology that simply replaces years of work.
"Microsoft could whip up paint.net in 1 day."
It takes a team college a few hours to analyze a bug or to write a new feature or upgrade something for security etc.
A week goes by so fast.
The last 20% are just what they are.
The software aspects make a big difference too, but it is actually quite easy to re-implement software. Something like git for example was built rapidly (3 months?) after a few years of using other VCS to learn what it needed to do.
body { max-width: 80ch; margin: 2rem auto; font-size: 18px; line-height: 1.5; }
I've built a few products with a skeleton team. You can move really quickly on a lot of issues if you have the right people. And let's face it, a lot of successful companies are not exactly rocket science in terms of what they come up with. Slack: middle of the road chat thingy. Their UX is nice and cute but otherwise there isn't a whole lot of special stuff going on. There are a gazillion alternatives that do more or less the same thing. Airbnb. Cute idea but it's a glorified market place. Twitter has some interesting tech related to scaling search but otherwise it's basically a fast version of mastodon, which is community supported. And they weren't all that fast or scalable initially (fail whale?!).
There are a lot of tech companies where the tech is pretty middle of the road. A lot of these companies actually established themselves with much smaller teams than they currently have. My theory is that they actually could not have done it with a bigger team. You need a small and focused team to get results. That's why startups are a thing. Large companies with way more resources struggle to do the same things that small, nimble teams of less than 10 people accomplish.
A lot of these big companies established themselves before they bloated their teams; not after. Google was a handful of graduate students that built a nice search engine and ran circles around their competition. They were smart, not big.
To move quickly, you need a good team that is not burdened by design by committee style management. Google seems paralyzed lately actually. I'm not sure they'd be able to build another search engine. Plenty of talent and money but their ability to focus all that on anything has gone down the drain.
> Businesses should keep adding engineers to work on optimization until the cost of adding an engineer equals the revenue gain plus the cost savings at the margin. This is often many more engineers than people realize.
But it's well written, do read it.