DéjàVu: a map of code duplicates on GitHub
blog.acolyer.org
blog.acolyer.org
EDIT: ohloh became openhub and now the code search is discontinued. So there is the nonfunctional GitHub search and an open niche for other projects...
Indexing and retrieving code at scale is actually a really challenging problem due to the fact that there is a lot of code, on a lot of branches, in a lot forked repos. With GitSense, it doesn't even try to determine the authoritative source (repo/branch), since I personally think this is a lost cause, given current AI technology.
With GitSense, everything is context driven, which is how you can reasonably remove duplication. To search, you have to define what branches/repositories to consider, which can a be a few to a few thousand. Note, once a search context has been defined, it can be reused, so this isn't something you have to create every time, if you want to search.
I sort of envision a Yahoo type (the first incarnation) approach to searching for code. The basic idea is provide a curated search experience, where domain experts can share what they believe to be relevant branches to consider, for a given problem.
Without some human intervention, I think duplication is a given and as you point out, can lead to useless results.
There are even scenarios where I’ve seen Java projects check in their dependencies (in conservative industries) so this erodes the value of the numbers greatly.
We have a pretty big issue with duplication of effort. I was hoping for an article that would be a wake up call but instead it’s just measurement artifacts.
If you all grab it from the same repo (possibly through some sort of 'repo link' mechanism that keeps a local clone but is clearly marked as a clone, of course) then when a bug is fixed, it's fixed.
If you copy and paste it, suddenly every fix has to be done once per project using the library.
Nobody is supposed to modify those files. They’re just a local replica.
And no, I refuse to believe such requirement is hard to solve at all. I refuse to believe AWS doesn't create some container-like environment behind the scene for launching and running a function.
See https://forums.aws.amazon.com/thread.jspa?messageID=791221&t...
there was at least one point in time where that was the recommended strategy
This is pretty much what the golang world does (though now there are some tools that do a better job).
I have already been in a situation where a dependency version that I was locked to was simply removed from the official registry.
The right solution is to have your own registry or backup the archives of dependencies that you are using.
I think it is better to commit the archives rather than the whole node_modules as it does not produce a mess.
1. http://blog.npmjs.org/post/141577284765/kik-left-pad-and-npm
Other than a local npm mirror, what do you think is the best way to store artifacts for new builds? Vendoring code is not a bad solution if it permits new builds to occur even in the event of a package manager outage. I see vendoring code as the simplest solution for small companies who don't want the management overhead of running a local mirror of package managers.
That said, I'm very much open to learning a new technique for handling this problem.
Even with a good deduping/compressing filesystem, the way git history is stored means that they're probably missing out on a ton of savings here. Eh, it's probably not worth the complexity/deviation from standard Git tooling.
You can read about their architecture in their discussions of Spokes, which replicates repository networks (the original repository and its forks) across data centers. eg: https://githubengineering.com/stretching-spokes/
Trying to put all of GitHub's object files in a single packfile - even just putting them on a single server - would be impossible.
But even on hosting providers with a bespoke implementation - that do not use core git to manage Git repositories - this would be challenging. We have a custom Git server implementation in Visual Studio Team Services, but it still makes sense to shard object storage with the repository: you have to worry about scalability and performance, but also things like data sovereignty. We can't just put a user's git repository in some global SQL Azure database that contains all the repositories in Visual Studio Team Services: the repositories need to be geographically located with the VSTS account they created.
Puting whole github in single packfile is obviously impractical, but having whole github on some bespoke Venti/IPFS-style content addressable object store is not.
If they are, then it likely just a direct implementation of git the technology. you can see how git stores data here: https://git-scm.com/book/en/v2/Git-Internals-Git-Objects
There is the concept of submodules which allows for multiple repos while maintaining the checksum mechanics that allows sharing the same bit of information between branches and across commits: https://git-scm.com/book/en/v2/Git-Tools-Submodules
The trick is that git maintains an abstract file system (ie a graph) across commits. The graph consist of pointers to content without having to create a clone of the actual content for every new version of said graph....it gets a little dizzy to explain but it is really not too complicated :)
This should definitely be taken as a lesson though: JS needs a better deployment solution. That, or better education on the current solution(s).
FWIW, and I know it's not much, I really don't use Node unless I have to (or javascript outside the basics, JQuery & LoDash for that matter); I was turned off from it when I was told to download and install Node via a copy-paste from their website of some short command-line wget script. That's shoddy at best; so the current state of affairs can be linked back to early practices. It's nice that Node has been cleaning up their act, but it's still kinda a crap fest; and now that is the standard that they've provided for their community.
To wrap up this meandering train of thought. The paper actually addresses this nicely, because when 70% of the fluff and cruft in JS repos is node_modules, you end up with hidden dependencies, which is how things like left-pad happen. With pip I know exactly what all of my dependencies are (explicit dependencies); with Node, you're required to dig into the node_modules for every known dependency of that initial dependency (implicit dependencies).
I would have thought that JS has fewer dupes because of NPM
My impression is that most public projects on GitHub are only of interest to the author, and maybe a small handful of people. I, for example, have over 100 non-forked public repos and, except for 3 or 4 projects, nobody even looks at most of them, much less clones them and uses them. Even the ~4 that do get attention, it's usually not because they're using the code itself - it's because they're doing something similar and want to see how I did it.
On the other hand, I only have anecdotal evidence to back up that claim, so who knows.
Welcome to the future of copyright trolls.
Are they even getting it right or do they all have the same bugs? Are there no existing libraries? Are the downsides worse? Can we fix that? Should this functionality live in the core language (did we miss a feature).