GitHub has completed its acquisition of NPM
github.blog
github.blog
Years ago, they introduced "orgs" which they sat there and explained to me with slides and pictures and concepts and business bullshit for an hour. I did not understand a thing they'd said. Finally, they were like "We're selling private namespace in the npm registry for blessed packages for groups or businesses." I understood that. If they'd just said that up front....
They had some great people, some very smart folks like CJ, but they completely biffed every business decision they ever made, and when you'd go in and talk to the leadership, they were always acting as if they had some sort of PTSD from the community. I mean, people were putting spam packages in NPM just to get SEO on some outside webpage through the default NPM package webpages. People were squatting and stealing package names. Leftpad... the community management here is nightmarishly hard, and I was never convinced they'd ever make money on it. MS doesn't NEED to make money on it. They can just pump in cash and have a brilliant tool for reaching UX developers around the world, regardless of whether they use Windows or not.
I feel like the GitHub group at Microsoft is now some sort of orphanage for mistreated developer tool startups. GitHub had similar management issues: they refused to build enterprise features at all for years unless they were useful to regular GitHub.com. And there were other people issues at the top for years. Chris seemed more interested in working with the Obama administration on digital learning initiatives than with running GitHub, for example.
It's not NPM's fault (well, other than wrt. the leftpad thing), it's all about the "community". The Javascript open source community is a dumpster fire.
The JS community is unstable, but in a way that produces useful things rapidly. Stuff just gets thrown out there and the good parts stick. Certain dampening factors protect you from the most dangerous effects of this (unless you live your life at the bleeding edge) like the jet's computer controlled micro-interventions protecting the pilot form how jittery the airframe actually is. Occasionally this goes wrong and the protection fails. "Leftpad" was one such matter that caused thing to crash and burn temporarily. Luckily unlike an expensive jet hitting the ground, the damage from such incidents is relatively easy to repair after the fact.
The throw-it-out-and-see-if-it-works part isn't the only issue with the JS community but it is a significant one and does relate to some of the others, though one that does provide some benefit in the form of forward momentum.
That the JS community is huge and overwhelmingly newer to the industry than (for example) me does not excuse bad behaviour. What it _does_ excuse is a lack of awareness of history and historical context. And the solution to that is education.
It pains me to see so many wheel-reinventions, both technical and social. But the flip side is that people who don't know something's not supposed to be possible will _sometimes_ manage to do it anyway.
NPM, on the other hand, seemed very much to be all about package management being a simple problem, and learned the hard way that there's a reason why other systems have more complexity.
$ time rm -rf node_modules/
real 1m2.969s
user 0m0.409s
sys 0m15.853sIn an existing UI project repo... ci (which clears node_modules) then installs from lock...
added 1880 packages in 23.283s
real 0m24.073s
user 0m0.000s
sys 0m0.135s
Still slower than I'd like... but I'm pretty judicious in terms of what I let come in regarding dependencies. That's react, redux, react-redux, material-ui, parcel (for webpack, babel, etc) and a few other dependencies.For one of the API packages ci over an existing install...
added 1069 packages in 12.911s
real 0m13.708s
user 0m0.045s
sys 0m0.076s
So either you're including the kitchen sink, or you're running on a really slow drive.Opening (or deleting) an empty file is about 2.5x slower on OSX than on a Linux running in VirtualBox on that same Mac: https://discuss.rubyonrails.org/t/why-is-rails-boot-so-slow-...
Was using a hackintosh and rmbp until about 2 years ago, stopped using mac at work, and in october switched to a new desktop and jumped to linux. Been back in windows + wsl2 for a couple months now.
Back using mac, and most of my windows until a couple months ago, was still mostly linux via VM.
Guess I never realized how slow macos's file system was for deleting files.
edit: Also, for those curious, WSL2 files in Windows is slow, and windows files in wsl are slow... each are fast in their own sandbox.
$ npm list --depth=0
<removed for opsec>
├── @babel/plugin-proposal-class-properties@7.8.3
├── @babel/plugin-proposal-decorators@7.8.3
├── @graphql-codegen/cli@1.13.2 extraneous
├── @graphql-codegen/core@1.13.2
├── @graphql-codegen/typescript@1.13.2
├── @graphql-codegen/typescript-graphql-request@1.13.2
├── @graphql-codegen/typescript-operations@1.13.2
├── @graphql-toolkit/core@0.10.3
├── @graphql-toolkit/url-loader@0.10.3
├── @types/lodash@4.14.149
├── @types/node@13.11.1
├── @types/react@16.9.34
├── @types/reactstrap@8.4.2
├── @zeit/next-css@1.0.1
├── @zeit/next-sass@1.0.1
├── babel-plugin-module-resolver@4.0.0
├── bootstrap@4.4.1
├── dotenv@8.2.0
├── express@4.17.1
├── UNMET PEER DEPENDENCY graphql@15.0.0
├── graphql-request@1.8.2
├── graphql-tag@2.10.3
├── helmet@3.22.0
├── isomorphic-unfetch@3.0.0
├── UNMET PEER DEPENDENCY jquery@1.9.1 - 3
├── lodash@4.17.15
├── mobx@5.15.4
├── mobx-react@6.1.8
├── next@9.3.4
├── next-fonts@1.0.3
├── node-sass@4.13.1
├── nodemon@2.0.2
├── react@16.13.1
├── react-dom@16.13.1
├── reactstrap@8.4.1
├── styled-components@5.0.1
├── styled-icons@10.2.1
├── ts-node@8.8.2
└── typescript@3.8.3That seems .. not so bad really? Microsoft gets to buy them cheaply, and they don't get obliterated or acquihired, and they don't seem to have Yahoo'd them into slow death either.
Being the canonical registry for a language (Rubygems) or technology (DockerHub) tends to be a huge expense.
The main expenses are cloud costs (bandwidth and storage) and security (defense and curation).
I've not seen examples of organizations turning this into a great business by itself. For example Rubygems is sponsored by RubyCentral http://rubycentral.org/ who organize the annual RubyConf and RailsConf software conferences.
Please note that running a non-canonical registry is a good business. JFrog does well with Artifactory https://jfrog.com/artifactory/ and we have the GitLab Package Registry https://docs.gitlab.com/ee/user/packages/ that includes a dependency proxy and we're working on a dependency firewall.
This is an example of the security curation you note earlier. People reviewing flags alongside the code, remediation during the delay, highlighting package authors you like are some of why “Being the canonical registry … tends to be a huge expense.” How many people build and run GitLab Package Registry?
The security curation examples I mentioned are intended to use automated signals so there wouldn’t be an expense for labor. In fact we intend to run them separately on every self managed GitLab installation.
As a workaround, you can add your access token as an environment variable and use that to publish/install packages via CI/CD.
And if you are interested in contributing to GitLab that issue is a great way get started.
Hosting a package registry on AWS for example, unless you figure out a way of seriously reducing the amount of traffic (which seems to be against working towards making a popular registry), is suicidal because of the bandwidth costs.
See also, CocoaPods. But we're really talking about two things here as if they're the same, and I think it's because the language we use for them isn't clear: package managers (like CocoaPods and Composer) and full-blown package registries (like RubyGems and NPM). I mean, the difference between them is slight; and there's overlap. They both resolve package dependencies and sort out problems, for example. But they're such very, very different things!
It's not really fair to compare these two approaches and say things about Composer like "this allows them to serve the whole community with a team of 2-4 people": they're just not directly comparable. Those 2-4 people + the entire GitHub staff is what supports the whole community.
Now - don't get me wrong: I think it's fine that GitHub does this! I'm not saying that CocoaPods and Composer are making the wrong choice here at all. In fact, I'm proud to work for a company that provides this service to everyone. But it hand-waves away a lot of stuff as just an "architecture" decision, when in reality it's a decision not to do the hardest part - which is hosting, because it requires gobs of money and dedicated staff and people on-call 24/7.
Framing it as just a different "architecture" implies that registries like NPM are making the wrong choices by hosting their own downloads, which I think is not nearly that simple. Because what happens if GitHub or Bitbucket decided to not allow this type of usage? It would be devastating for the community - so I think it's very appropriate for package managers to make a decision about whether or not they also want to be a registry. They're considering what could happen down the road and what that would mean.
CJ's comments on all of this, re: entropic, are really good: https://www.youtube.com/watch?v=MO8hZlgK5zc
I also kind of wonder what is the real value of a centralized repository versus just directly referencing git repos. I haven't used this gpk[0] project yet, but it looks like an interesting alternative, on paper.
import * as E from 'https://cdn.jsdelivr.net/npm/fp-ts@2.5/lib/Either.js';
Whenever you `--reload` you'll get the latest 2.5.x release, no?This should work out recursively for all the deps if they follow the same pattern, no?
At this point they've just replaced `npm` with `pika` and expect people to think it's an improvement to not have a centralized file containing package dependencies and the last resolved versions of packages. Criticisms:
1. Updating dependencies becomes much more of a diff (every file that uses it needs to be updated). The alternative solution to this is creating a single `packages.ts` file that imports all the external deps, and all your modules import from there. Which just `npm` with more steps.
2. No lock file. I saw them present this at tsConf and I believe the original idea was that this wasn't needed because the URLs are the lock, but as you point out in practice people use URL's that aren't unique identifiers. This means no way to guarantee that all developers or even users of the project have the same dependency state. This will be a massive pain for debugging and maintaining packages. (Edit: see below, you can lock deps (unclear exactly how) - but only if you pay them!)
Looking at the Pika home page, they tell me I can build and release without a bundler.... so you're telling me that they expect people to put out production software that downloads it's dependencies at runtime and thus:
- Is unstable because if they use the semver-URL scheme you mentioned and a patch version comes out that breaks the website, all users are instantly broken instead of the internal build being broken.
- Is unstable because if a dependency/host goes offline, all users are broken instead of an internal build being broken (think leftpad but much worse as your users are instantly impacted, likely before you even know about it)
- Is insecure because if a host is malicious, they can choose to supply different packages for a small subset of the requests, such as those coming from govt. requests against political targets, hosted build machines, etc. and nobody will have any way of knowing because there's no lock file/integrity hashes.
I further see on the Pika CDN page that they share packages across websites. This is seems to be a massive security flaw, as websites are able to modify these packages and now those modifications will apply to all websites using the package. It's prototype pollution-as a-feature!
Oh, and at the bottom of the page:
> Want to get more out of Pika CDN? The CDN will always be free, but you can also access paid, production features like:
> Granular Semver Matching & Version Pinning
> ...
So base features that are the core of working with npm/yarn/any modern package manager are a paid feature in Pika.
I'll pass.
<< - Is unstable because if a dependency/host goes offline, all users are broken instead of an internal build being broken (think leftpad but much worse as your users are instantly impacted, likely before you even know about it) >>
So you don't recall Leftpad?
<< - Is insecure because if a host is malicious, they can choose to supply different packages for a small subset of the requests, such as those coming from govt. requests against political targets, hosted build machines, etc. and nobody will have any way of knowing because there's no lock file/integrity hashes. >>
Are you kidding me?
https://www.zdnet.com/article/microsoft-spots-malicious-npm-...
They are also working to support importmaps which could become part of this story.
In any case, competent and smart people are working to make it usable. It is still a work in progress but I love the direction they are taking. It will be great to be able to run safe javascript where I provide it with the permissions it gets.
NPM is both a massive repository as well as a package manager.
Deno will (soon?) have a package manager, but it won't be tied to one central repository run by a private company, now part of a massive corporation (Microsoft.)
So let's not lump package manager and repository: you can have a package manager that pulls in all the exact or latest dependencies at build time, but it does not have to be tied to one central repository owned and administered by one company.
Centralize-able yet decentralized.
So this problem it’s pretending to solve isn’t actually a problem. And the solution introduces more problems (see comment on sibiling)
I would like to see a package repo for deno, if only to ease publishing/finding modules.
In fact Ryan Dahl said one of their biggest regrets was blessing npm in node. I'd be surprised if Deno jumped into blessing a package registry any time soon.
The acquisition of GitHub was absolutely intended to capture a tool/ecosystem that developers liked using and benefit from that positive sentiment. That's why Microsoft has been so cautious about branding GitHub as a Microsoft property out the gate. It's trying to ease devs into the idea that the company is something devs can like, and I wouldn't be surprised if this psychological strategy is at work with the npm acquisition, too.
You can’t change people’s minds on emotional baggage like this. You just wait for them to die off/retire and target the new blood.
About all Microsoft has dared to change (at least transparently) since obtaining GitHub has been the pricing page, and objectively speaking right now... it's a hot confusing mess.
A mascot featuring tentacles is all too fitting here. I don't know a lot about NPM or JS but it's being described as a dumpster fire which is also all too fitting. And I'm able smirk as I write that with no animosity or skin in the game.
I believe you can. They changed mine.
The quoted statement had nothing to do with Windows 10 and was specifically do with Microsoft's digital services like Hotmail & OneDrive, which the statement has now been updated to clarify:
"Finally, we will retain, access, transfer, disclose, and preserve personal data, including your content (such as the content of your emails in Outlook.com, or files in private folders on OneDrive), when we have a good faith belief that doing so is necessary to do any of the following:"
Microsoft doesn't care about your local files, but you agree with the generic "to comply with laws and protect our systems" statement if you decide to use their cloud services.
So...appeal to devs, something something, money??
I really like how MS are improving our tools and embracing open source, I really do. But I've never quite understood how the return on investment in these things justify the cost. I just struggle to the an obv big picture here.
I.e. is it incorrect to think of the GH acquisition as mostly an azure marketing expense?
Now most software is written and deployed on the cloud, but not their cloud. But they could make the dev experience compelling and easy - write your code in VS Code, which automaticallgy integrates with Github. Github is mostly free until you're a large company so why not use it. Of course you need to run CI and Github makes that easy so go ahead and add a single file to configure that.
Now you have a build artifact ready to deploy on github. Would you like to click a single button and have that deployed to Azure? They'll also throw in monitoring if you do it. Azure bills is the pot of gold at the end of the developer experience rainbow.
But this is just speculation. I don't understand business very well.
Microsoft, "the people" were never bad. They did a lot of cool things, they had great coders and scientists. The upper management did bad things, and made poor decisions, guys like Ballmer and business strategists.
With the new management things changed a lot and I hope Microsoft will keep up at doing good things.
Not all things are rosy: desktop development with MS tools suck, they changed framework after framework, Windows Forms and WPF aren't cross platform, C++ desktop programming is still Windows only, and not Visual at all. I don't get why the naming: "Visual C++" since QT or C+= Builder are much more "Visual" than MFC and Winapi.
But I guess desktop doesn't bring MS too much money so they don't care enough about it. If that's the case, I don't get why they don't promote and support a third party, cross platform development framework like Uno or Avalonia.
While Upper Management has lessened some of their excesses, they still tend to be very aggressive in some area's, and time will tell which 1/2 of the company will win. The Open Collaboration group, or the "We want to control the world" group, there is still an internal struggle there
Many believe the "We love Open Collaboration" is just a facade and the "real Microsoft" will revel itself in a few years
Looking at the massive growth in frontend technologies within the last decade, it’s easy to see why Microsoft wants to be a little ahead of the curve this time.
Getting to decide what goes into npm (the client tool) and what does not allows them to focus the ecosystem onto brands that they own and encourages people to buy their proprietary software and services.
It also provides them the option of cutting off third-party tools (e.g. yarn) in the future if they deem it beneficial.
My prediction is that they will prioritize the platform features that can only be accessed by first-party, branded tools, like the forthcoming GitHub mobile app. Eventually they will stop maintaining support for the APIs that the other tools use, and it'll be Microsoft tools from end-to-end. Pushing to GitHub and publishing to NPM from right within VS Code, et c.
Of course, using them on Windows will always work best, and deploying to Azure will always be easiest.
If anything, I expect basic Visual Studio internals will eventually get open sourced as cheaper to maintain that way than Roslyn rebuild-all-the-things. And VS Code will adopt them and continue to cannibalize VS mindshare.
I doubt Microsoft would intentionally degrade developer experience, like "cutting off third-party tools". Rather, they're seeking to gain market advantage from the tight integration of services. It would make sense for them to encourage third-party tools to play well in that "Microsoft ecosystem".
> Pushing to GitHub and publishing to NPM from right within VS Code
That's exactly what I picture coming soon, if not here already. Also: develop in VS Code, click to build, push to GitHub, deploy to Azure.
Can anyone point to a technology company that has a larger attack surface or a more fundamentally insecure product?
Maybe to be able to be "legally compelled" to distribute backdoors onto linux servers / developer machines running npm in exchange for being awarded the JEDI contract.
The old model of business was like selling pies to dad. Dad might buy or not. Now, you give a free candy to the kid, and dad will buy the pie, too. :)
What's preventing the dream of decentralization from taking off? We have the technology.
And around the time when home connectivity became good enough that people considered home hosting was also around the time Slashdot was created.
I had a members.aol.com/benibela site or something
Examples: Let's Encrypt (Certs), Internet Archive (Culture), Quad9 (DNS), Wikipedia (Knowledge), OpenStreetMap (GIS), Python Packaging Authority ["PyPi"] (as part of the Python Software Foundation)
EDIT: Seriously, start non-profits whenever considering implementing technology infrastructure you're unlikely to want to extract a profit from and are seeking long term oversight and governance.
Generally speaking, much like for profits, it is the people that run it that decided its culture, not a legal structure.
The average developer isn't interested in showcasing their social/political views, starting a revolution or building the future of the internet. They just want to get the job done as quickly and effectively as possible and go home.
All being equal, though, we should try to avoid technical systems with unnecessary trust relationships and critical concentrated dependencies, though.
And if we must have a concentrated dependency, we should do our best to pick very trustworthy ones-- both based on their track record and their likely future interests relating to organizational structures.
If you don't trust the code itself – it doesn't matter where it is hosted. Validation/audits etc. are essential in any case.
If you don't trust the reliability/uptime of the central service — it's trivial to create a private mirror containing the packages you want. Every internal build system I have seen does this already.
NPM's had its share of catastrophes and controversies that have made things tougher and caused harm to downstream users. It has also saved lots of effort.
Focusing on e.g. mirroring misses the point, IMO: One's dependency on something fundamental like package infrastructure isn't to survive one deployment or minimum sustaining, but to be an ongoing part of your technology stack. Yes, you can move on, but it'll be costly. If the component decides to start to suck, you're going to feel pain.
I think this is probably a good thing in general, and should maybe lead to some interesting enhancements, and maybe even finally solve the distribution of binary modules at a better level.
The post if I remember was mostly ignored, but received a few downvotes and maybe a couple of negative comments.
Based on that, it seems that what's preventing decentralization from taking off is ignorance and apathy.
A day or two later the guy announced npm, Inc. if I remember.
There are actually a lot more developers that have accepted a federated services worldview than a peer-based fully distributed one. But there are package registry projects along both of those lines.
https://github.com/orbs-network/decentralized-npm
https://blog.aragon.one/using-apm-to-replace-npm-and-other-c...
https://github.com/entropic-dev/entyropic
But again, ignorance, apathy, and the status quo remain the most popular options.
Time and money.
I have recently started getting into JS programming. I have thus far avoided NPM, because I've been trying to use CDNs for all my external dependencies.
My thinking is that it saves me bandwidth costs and potentially saves my user's bandwidth as well if they get a cache hit.
I get the downsides are that I don't control the CDN and they could go offline, but honestly I expect I am much more likely to go down from some mistake in my own deployment rather than a well known CDN being offline.
I am wondering if I am missing something though, because absolutely every JS package I read about suggests you use NPM (some also link a CDN, many don't). Should I be using NPM to manage my JS dependencies instead of using CDNs?
And even ignoring my user's bandwidth, it would still save me significant bandwidth (depending on the size of my website).
I guess eventually your site might grow such that your dependencies are not a significant portion of your total download size, but I am not currently there.
There are advantages to having all needed assets locally. The main point for me is to minimize external dependencies during runtime - fewer points of failure. Also, vendor libraries can be a single minified bundle served from the same domain. In production they can be moved to a CDN, i.e., CloudFlare.
Using NPM makes sense once you start having more than a few dependencies, or a build step.
On the other hand, if you can get by with library CDNs and don't feel the need for NPM - I'd say that sounds fine, to keep it simple and practical.
CDNs are far less reliable than my own site, and if my own site is down it's not much help that the CDN is up. Pull in two libraries from CDNs and suddenly you have three points of failure instead of one. Their traffic spikes become your traffic spikes, their downtime is your downtime. And for what? The possibility that maybe the user had that one tiny js library cached, or that the CDN has a node 100ms closer? Not worth it.
See Nat's original words about the acquisition: https://github.blog/2020-03-16-npm-is-joining-github/
GitHub Teams get 2GB of free private package storage (unlimated for public packages) https://github.com/features/packages
He did mention that they were using their enterprise customers to subsidize the cost to allow them to offer Teams for free, so maybe if they get enough enterprise customers using GitHub Packages we might see unlimited free private packages for individuals/teams. I don't see a ton of value for GitHub/MS in that situation, but maybe.
To be sure a lot of this may be about talent as well. They may need to be liked in the community otherwise they have a hard time hiring top talent. They are buying a community since they probably couldn't build it themselves.
I'm personally not worried about the acquisition (I think it's a net positive for the community), but there are plenty of reasons to discuss this interaction.
I hear you, I think MS is prone to mismanage the nicest things we have. And yes part of this will be some PR mess up.
We definitely cannot trust a large cash reserve to guarantee the survival of these technologies, it's just not how businesses work.
MS was in the business of selling operating systems and desktop software. Not only they realized that open source doesn't threaten that territory (remember year of the Linux desktop?), but they've changed the revenue model and are making much more money now from selling services than by selling software.
>Because of how other large companies like Google have smothered open source projects?
Why should we judge what one company does based on what other company does?
I am sure no company does something for "good of humanity" but to earn money. What matters is if in the process of making money, they also do good things or bad things.
Microsoft has stopped doing bad things and started doing good things. I only hope more companies will follow.
I could imagine other potential acquirers just wanting to do an acquihire, and deprecating the service a few months or years later (when the technical debt is too high to continue running the code in maintenance mode).
Docker Hub feels a bit neglected - it could be aliased to docker.pkg.github.com and that'd be a huge improvement
MS gobbling this one up would not be a big surprise.