It's hard for me to imagine someone grokking a 10M+ line codebase without external help, but I've never tried it. I do agree with the assertion that most codebases are not as _special_ as they like to think.
10M is a pretty massive codebase like the entire linux kernel with all drivers is somewhere in that size. Most corporate systems aren’t that big and even for Linux you wouldn’t need to understand all drivers to understand the core kernel, I suspect the core kernel is maybe max 1M.
https://engineering.fb.com/2022/10/24/android/android-java-k...
Regardless it’s a bit concerning that it takes 10M loc for a messaging app.
"One former Tesla engineer, who spoke on the condition of anonymity to candidly describe the matter but was not involved, said Tesla engineers would have trouble capably assessing Twitter’s code. Distributed systems, the large-scale and spread-out network that Twitter is composed of, are not the automaker’s specialty, the person said."
:P
It’s certainly true that it helped I was
familiar with the domain
Technical prowess and domain knowledge are excellent assets, obviously, but in my experience they're often not enough.The big tangled enterprise codebases I've dealt with (insurance companies, fintech, construction, etc) involved absolute metric tons of undocumented domain knowledge and lots of company-specific "tribal knowledge." Some tribal knowledge was embedded in the code in undocumented or semi-documented form, and much existed outside the codebase entirely... all kinds of custom infrastructure, etc.
I don't care how sharp and domain-familiar a team is. That sort of situation is not easily tameable.
Everything seems so incredibly trivial until you reach something that isn't.
For example you could look at my team's mobile codebases and probably figure out what was going on, but understanding all the services we consumed, and what they consumed, etc. (given the deep mix of micro services and macroservices) would make understanding the why of the entire system impossible.
A pretty trivial change for someone whose steeped in the codebase, likely impossible without a few weeks (or even months) of effort for anyone else. Of course, this all becomes exponentially easier if you have an author of the code to point you in the right direction.
What if a large portion of the codebase is, for example, shader code? I chose this example because coding for the GPU isn't the same as coding for the CPU. Do you think that's a scenario in which you'd require more study, or are you confident your experience would spill over into this new domain, no study required?
This operation is not just interviewing people, it’s a kind of knowledge transfer and finding the people who are the best at explaining how the code works and can answer deep technical questions. This is how Elon works generally (if you look at the SpaceX interview, you can see that he just goes to people and asks them questions about the parts of the rocket while he’s doing the interview).
The codebase often isn't the issue, it's the use cases and the reasons why it evolved into the form it did.
But to address your question here is my recollection of what happened, now more than 20 years ago fwiw. We were given the code and I spent maybe 6-7 days, all day, reading it and analyzing it with this tool I had, called Source Navigator [1]. Then we spent 1 full work week at the other companies HQ, mostly in meetings asking questions on different modules and classes. Then when we returned to our offices it took me another 2 weeks of work to get the system setup and deploy a few components inside our own middleware system. I was the primary c++ expert, there was another business analyst and I had a more junior developer who worked with me. So in comparison to Twitter certainly a much much smaller scale situation. The team that had written the system was around 15 people.
We definitely used the code and I don’t recall it being much a problem at all that other people had written it. Plenty of open source projects have random contributors show up and work fine in their code base.
I think a lot of SWEs have pretty big egos and tend to overestimate how special or unique their particular projects are based on my own professional experience. This particular situation was an example but there have been plenty others. When I fix or find bugs in other’s code sometimes they are surprised which for me is always surprising. Why are you so surprised I can debug your code?
[1] https://sourcenav.sourceforge.net/
(I’d be curious if people have a favorite more modern version of a tool like Source Navigator)
It is about coming up with BS false positives, pointing them out and saying this code is crap.
Of course if someone is professional and understands there was different context and all you have is code and writes down false positives and discusses them it is OK.
Yes you can, it's called an audit and there is nothing wrong with that. The company you work for should have regular security audits for instance, ideally done by a third party rather than internally to eliminate bias. This isn't a "code review".
A quick interview and a little demonstration of contribution can help assess these things significantly, you don't have to understand the codebase that much to do it.