No, you cannot trust third party code without reading it first
unixsheikh.com
unixsheikh.com
You do need a model for evaluating whether or not to trust a dependency, and for understanding if and when to extend that trust to a subsequent version/update.
I’m going to suggest that if your model is to fully understand the dependency, e.g., by reading every line of code, then you should probably stop having any dependencies and develop everything yourself —- just to save time. (But let’s face it, you should probably just find a new line of work).
It sounds glib, but I’m serious.
You can literally only run software only you’ve written, only on hardware you’ve designed and manufactured (and personally delivered), or you are trusting someone else, somewhere.
What is your basis for this? Actually?
If you’re not sure, but have an intuition, that’s not too bad… you’re in good company and like practically everyone else.
IDK, spend a few minutes thinking about what your model of trust is, and why even have one. Really, if the best you can come up with is to not trust anything you can’t personally verify, software development is not the best for you.
But the issue I see is that larger software installations have larger attack surfaces. And the typical python, js, golang (eg.) projects are just exploding in dependencies.
I think the python project I run at work is now over 100 deps on PyPI. My simplistic golang side project is around 50 or so.
It's a front end for gathering metrics about webservers and the like.
And sure enough sitting in site-packages were these java jar files.
But in terms of Getting Things Done, I also did `pnpm add date-fns` this afternoon and have never reviewed the code for `date-fns`, because it seems to do what it says on the tin and is generally well-regarded. There’s a balance to be obtained, and you have to trust someone, because you’re not going to read the source code to clang or gcc.
So in general, I agree with you: the article here is horrible advice.
I'll add my voice to the chorus: The article is horrible advice.
Nobody reviews third party source code for security issues (well, almost nobody). It would not be a productive use of time. There are almost certainly many other less expensive things you could do with that resource to improve security.
> Of course you cannot do that with everything, you cannot read all the source code for the kernel of the operating system you're running, you cannot read all the code that makes up the compiler or interpreter you're using, but that is not the point at all, of course some level of trust is always required. The point is that you need to do it when you're dealing with code you are writing and importing!
OK, so we're drawing an arbitrary line of what we read and what we don't.
A basic web application at this point is going to import at least 10k lines of other people's code. A React app probably uses millions.
There's just no way around this. We all trust code we haven't read and have to continue doing it.
Reading code doesn't even guarantee finding security flaws. Most of us aren't security researchers, and none of us can understand an entire project's source code the first time we read it.
Wrong. The way around it is to write small amounts of code. We should be searching for local minima instead of convenience. Tightly integrated, batteries included programs like redbean show one way. Clojure, with its emphasis on terse, composable objects shows another. Those who write web clients using only HTML, CSS and JavaScript show yet another. We require a reusable set of small primitives that can be examined, mastered, and recombined to suit our purpose. We do not require...a great deal of what we now have.
So at a minimum, you'd still need to read and understand many thousands of lines of some of the most complex code in common use.
It's not realistic and it's never going to be common practice.
I find often it’s not the content of the article that’s relevant. It’s generally not as there’s a lot of trash out there. But they do raise good discussion topic that lead to active conversation. I find that is the real value for me personally.
If compared to other industries: We all drive a car we never fully inspected ourselves, fly a plane we have't looket at ourselves and eat food we didn't see where it grew. How can we adapt the supply chain in software engineering that we can have the same level of trust (or trusted parties)?
I agree: we should have this much regulation for software security. The stakes are extremely high.
Imo the line is not arbritary. You are including an library/framework for a purpose. The code path for this purpose should be explored. Recently I allowed a friend of mine to DDOS my server. He used unvisited software. Now my server suffers from thousand wild guesses daily. Previously, I had ten visitors a day. Therefore I conclude that the software leaked my server. And his learning: Don't trust any software; Exspecially if the software is for malicious purposes in the first place..
> A basic web application at this point is going to import at least 10k lines of other people's code. A React app probably uses millions.
I think it is a valid assumption to increase the trust in frameworks with a giant user base. React is from Facebook, isn't it? So I would have read the path _and_ random internals to picture the trustworthiness.
> There's just no way around this. We all trust code we haven't read and have to continue doing it.
It shouldn't be black-and-white thinking. When introducing a _new_ dependency I see myself responsible for estimating the trustworthiness. But yes, with an increasing amount of dependencies the process of updating such needs more effort as well. But then, we can postpone such updates, if the application under development is safety critical...
> Reading code doesn't even guarantee finding security flaws. Most of us aren't security researchers, and none of us can understand an entire project's source code the first time we read it.
Well, one should be able to judge about the workings imo. Otherwise maintainability can get painful in the long run. So better grasp it before deciding to build upon it.
Another comment stated that we are trusting cars we drive, etc. etc. This is about the users running our code. A proper craftsman will inspect its material, before using it in quality products. When building a throwable tool for his work, he doesn't spent to much time worrying, of course.
Yes, but it uses a massive number of libraries that were not developed at Facebook.
> But yes, with an increasing amount of dependencies the process of updating such needs more effort as well. But then, we can postpone such updates, if the application under development is safety critical...
Postponing dependency updates is very, very bad for security. That is not a solution to supply-chain attacks.
> Well, one should be able to judge about the workings imo. Otherwise maintainability can get painful in the long run.
The whole point of APIs is that we do not need to understand the inner workers of code that we're calling. How many of us use bcrypt and couldn't tell you anything about the underlying algorithm?
You are right; This underlines the thought and care some distributions put into their package management system...
> The whole point of APIs is that we do not need to understand the inner workers of code that we're calling. How many of us use bcrypt and couldn't tell you anything about the underlying algorithm?
That is right, but code written by unknown developers can be a huge risk. Of course you are not assumed to read up upon any dependency, but from external sources. A quick glimps on the imports and dependencies goes a long way, I think. In the end your team is responsible for security issues, even if they appear in an external dependency. Companies with direct customer sales are spending tons of money for mitigation strategies. Maybe a chunk of this money should be spend on validating this beforehand.
Minimizing the number of dependencies helps a lot too.
But don't blindly npm or pip install something unless you trust the developers. npx/pipx are even worse. All it takes is a one typo-squatter to steal your ssh keys and maybe even saved browser passwords or cookies.
If some big org is using the same code I am then if there's a bug a patch will be written and released.
This will be a good thing for large companies, create a two-tiered world for open source, and create another set of gatekeepers. And if I were the author of a popular open library or tool, I would be thinking about how to get a piece of that action to support future development.
Will you be able to identify and avoid issues with 3rd party libraries by reading their code, and the code of all the other libraries that they depend on?
Do you know what all vulnerabilities exist in cyberspace?
I mean that a static code analysis tool can take you much further than reading 3rd party code manually... and that is still going to fall short, but that is as good as it gets
Really, reading it is only loosely coupled to the trust model.
Reading code you use is a good habit. Reading it CAN improve its trustability, if you're familiar with both the language and the specific application space that the code operates in. Otherwise it's simply false that reading the code will help. In cases like that, it is better to use different heuristics for trust.
This 'overshooting the mark' behaviour is sadly very typical of a certain pocket of the Unix community, that have far too strong a conviction to their beliefs. I believe it severely damages the credibility of otherwise knowledgeable sources.
When that is the working environment, he's right. Trust no one.
What you want is to reduce your risk so if you can present evidence that Google or some other high profile has done the due diligence then that will be good enough to indemnify you.
Must regular developers don’t have time to read the code anyway. You need to use the slipstream of other companies to protect you.