Code scanning for security vulnerabilities now available
github.blog
github.blog
In theory automatic vulnerability scans sounds great, but having every repo ping you with not-actually-an-issue becomes a chore very quickly. So far the vast majority of vulnerabilities I've seen are actually noise/not applicable. If this code checker is actually good, unlike all of the previous ones, that's another thing and might actually be a game changer.
Prominent open source authors have often suggested ways that GIthub can help but seem to be ignored, e.g. allowing to add friction to opening random issues would benefit open source greatly. At some point many beginner devs migrated from StackOverflow to Github because their really bad question were being closed there, and now they just overwhelm open source authors.
[1] https://twitter.com/sindresorhus/status/1123986529498664961
[2] https://twitter.com/FPresencia/status/1311551520689713152
There's also not a severity indicator - some minor issue not encountered in normal use is just as noisy as an extremely important issue that affects every user.
Other tools like Jira and Gerrit are far better at this.
From that limited experience, I'd say that false positives are less of a problem with Semmle's checkers than with other security-focused static analysis tools. This is partly due to Semmle checkers being much more customizable; Semmle has developed a declarative query language called CodeQL* which its checkers' built-in and user-provided rules are written in. Microsoft's security development lifecycle has a lot of mandates which are captured by custom CodeQL rules precisely enough to match their intent.
You can see some examples of how Microsoft uses Semmle here: https://msrc-blog.microsoft.com/2018/08/16/vulnerability-hun...
Most of the time what I see is
"A dependency of a dependency of a dependency of Webpack is vulnerable to a Regular expression Denial of Service attack" or prototype pollution or something like that.
In general, it's quite strange to me that vulnerabilities in `devDependencies` are considered less important than those in `dependencies`. These dependencies are generally for tools that are run within your company network, and contrary to what people insist, that seems quite risky to me.
I imagine those cases could be detected using static analysis but the current vulnerability reporting tools that do not even look at actual code are completely useless in that regard.
It needlessly accumulates risk when you prioritize explaining why things don't need to be fixed rather than just lowering the bar to fixing them
At the very least, running and pruning scans should happen on projects so that at least we can have the conversation. It's like PCI (as an example, not an ideal); PCI isn't perfect, but at least it encourages a conversation about security. Today we're at the point where almost every single organization at least discusses security; but I remember when PCI first came out. I can't tell you have many times people used to ask why it was problematic to store passwords in plain text.
It this the best step forward, probably not. Is it a step forward, absolutely.
Neither is proof of a vulnerability. Both I will gladly, always 100% validate.
I agree with the parents point, if there is value in such tests, it's so miniscule that it dwarfs the added effort to filter out all the false positives if the developers aren't total newbies. And if they are, code reviews by more experienced developers should be a given, making it redundant as well.
I'm glad it's not just me seeing this - My repos aren't even that popular and some of the issues just seem to be "help me build my project..."
Unfortunately all tools have either false positives or false negatives, and in practice often both. Tools can (and should) take steps to minimize them or their impact. Nothing makes you use any specific tools; if you don't like what a tool does, don't use it.
> Specially troublesome are when e.g. the "vulnerability" (if it's even one) is in a devDependency that is not deployed to production.
This particular GitHub tool is for analyzing source code, and would not not normally analyze your dependencies. So that doesn't seem relevant in this case.
Of course, someone could mindlessly use this tool (or look at its results) and complain. That's easy, just ask for funding to fix the problem, or at least a pull request that fixes it.
- C/C++
- C#
- Go
- Java
- JavaScript/TypeScript
- Python
https://docs.github.com/en/free-pro-team@latest/github/findi...They stated in the comments here that they are actively working on it, and I would personally appreciate it in addition to the number of other static analysis tools I use.
> A lede is the introductory section in journalism and thus to bury the lede refers to hiding the most important and relevant pieces of a story within other distracting information. The spelling of lede is allegedly so as to not confuse it with lead (/led/) which referred to the strip of metal that would separate lines of type. Both spellings, however, can be found in instances of the phrase.
[1] https://www.merriam-webster.com/words-at-play/bury-the-lede-...
Lede is just lead spelled incorrectly with the same meaning.
https://books.google.com/ngrams/graph?year_end=2019&year_sta...
> Contact Sales to learn more
Looks like they expect you to upgrade to their Enterprise plan. That’s a 5x increase in subscription costs.
If you're in an organisation that's are already using GitHub and this scanning capability is as good as that of CheckMarx, Snyk, etc. then it would be a no brainer to upgrade your Enterprise GitHub plan (if you're not already on Enterprise).
So, I guess, "well done"? (it hurts a little though, I'm too in the camp of wanting to see the pricing beforehand)
One of my ancestors had a company selling commodities. He wouldn’t answer the phone until the customer had called three times and left messages. He said this was a filter to identify the customers who really needed his product.
The logic is sound and on the surface clever. But would he have made more money servicing all customers? Or perhaps marketing?
My dad runs a shop restoring classic cars, mostly as a hobby. The number of minutes you can spend on the phone discussing some potential client's great-grandfather's old clunker and their dreams / aspirations of getting it fixed is nearly limitless, but the number of clients willing to actually pay and wait for it is a tiny fraction.
Similarly I looked into having a hair transplant done and found most places near me actually charge for consultations! It probably makes sense, though. The market of "guys that want to be less bald" is gigantic, but the ones serious enough to plop down tens of thousands of dollars for it is much smaller. Narrowing that group down made for a consultation that was much more personal and focused than it'd have to be if they offered it for free knowing 95%+ will never be seen again.
We could argue that they would get more money, but at the end of the day they probably decided that a single bigger fish is definitely worth losing business from a hundred smaller ones. I guess it's one of those "nice problems to have", if you grow so much that you have it.
My impression is that they haven’t picked pricing yet.
It frustrates me when the price answer is “contact sales and let’s talk about it.”
These marginal services really depend on the price, I think.
So the problem is that they are trying to figure out the “value” instead of giving it out for free.
What’s interesting about gitlab’s sast is that it’s really mostly just curating open source tools and using their existing CI stack. So it’s possible to just look at their code and set up, but a hassle.
I’m hoping GH just gives it away for free. Or does some sort of sonarqube method where you layer on indemnity or something.
The problem is that per user pricing sucks because most users don’t care. The PM or maybe one or two are interested, but paying for 400 users with licenses to please 5 audit people doesn’t make sense when it’s really just being run by a single Jenkins robot (or equivalent).
Where "oracle" means predictor and not the database/java folks.
Here's my mvp:
public static (bool IsOwned, string OwnedBy) GetAttribution(string CodeSnippet)
=> string.IsNullOrWhiteSpace(CodeSnippet)
? (false, null)
: (true, "Oracle Corp.");
Investors please :)Will never happen. Can you imagine what would occur if Github started harrassing private repo owners for including GPL licensed code? Or automatically making them public? All their customers would bolt immediately. Copyright infringement is Github's bread and butter.
Stack Overflow attribution is equally unlikely. The same group that says Oracle's claim that Java's API is not original and unworthy of copyright protection, cannot then turn around and claim 30 lines of code from SO is original and deserves recognition.
I would bet most stack overflow answers don't qualify as copyrightable, at least in the US. Though I think automatically finding copied stuff would be very useful. Any time I copy something small I try to include a link to where I got it from, if someone has to troubleshoot my code it may help them to see where it came from.
Why would Github do that? It's perfectly fine to not distribute code you received under the GPL. The license just says that if you do distribute the code in binary, you must provide recipients with the source as well.
In the past, I've done a couple one-off scans at https://lgtm.com/
GH Actions are:
- complex enough to develop that third-party solutions are attractive, especially for simple-seeming tasks (I have just been through this)
- but yay there's a "marketplace" for actions!
- the code in the marketplace is the wild-west, but you'll find something that seems like it'll do what you want
- you'd better hope the code that was committed to be executed is actually the compiled source code (if, eg, it's based on the official Typescript example and committing some [com/trans]piled JS blob that is what actually gets executed)
- it has access to your code, maybe including write access
- or it could do damn near anything else
Example:
steps:
- uses: actions/checkout@v2
- uses: sfdx-actions/setup-sfdx@v1
with:
sfdx-auth-url: ${{ secrets.SFDX_AUTH_URL }}
or steps:
- uses: actions/checkout@v2
- name: Run deployment script
env:
HEROKU_API_TOKEN: ${{ secrets.HEROKU_API_TOKEN }}
The first is for an Action, the second is for a normal workflow script.CodeQL code scanning automatically detects code written in the supported languages
C/C++
C#
Go
Java
JavaScript/TypeScript
Python
Source: https://docs.github.com/en/free-pro-team@latest/github/findi...OpenTelemtry for example doesn't include Ruby in its initial beta program announcement.
".NET, Java, JavaScript, Python, Go, and Erlang!"
Well it has always been like that. Amazon and Google has always had a thing about Ruby, and it is a minority market so I am not surprised and it doesn't make sense from Business perspective. But Github is a heavy Ruby users so I would have thought Ruby would be a first class citizen. I wonder if it has something to do with the language complexity.
Edit: From Github.
https://news.ycombinator.com/item?id=23094160
We (GitHub) absolutely plan to expand the list of languages CodeQL supports, and Ruby is a language we'd love to add (we're heavy users of it internally). In the meantime, because code scanning is extensible you can plug in third party analysis engines to scan the languages that CodeQL doesn't support.
They have been part of GitHub for barely a year so it's not too surprising, especially given they are continuing to support the product for the enterprise customers they had previously not just GitHub.
[1] https://techcrunch.com/2019/09/18/github-acquires-code-analy...
We're adding Ruby support to CodeQL (the scanning engine used in code scanning by default). It's our top requested language, and one we use extensively internally. Adding each new language to CodeQL takes about 6-9 months and needs a team to maintain it in perpetuity, which is why we don't have it yet, but we're starting that work now.
The other languages we hear the most demand for CodeQL support on are PHP, Kotlin and Swift. We'll get to all of those - it will just take a little time.
In the meantime, all of the code scanning experiences are extensible, so you can use other scanning engines with it, like Brakeman for Ruby.
I really think that PHP should be on that list.
So it looks like it supports C/C++/C#/Go/Java/JS/Python/TS
> This open-source repository contains the extractor, CodeQL libraries, and queries that power Go support in LGTM and the other CodeQL products that GitHub makes available to its customers worldwide.
[1]: https://github.com/github/codeql-go#go-analysis-support-for-...
My experience with their solution has been:
- 'unusual' repo structures throw it off, silently doing essentially nothing.
- noise is real
- It caught some genuine vulnerabilities
I'm not sure if github also offers the secret detection, license scanning and dependency scanning that gitlab does.
This new GitHub feature will scan your code on potential vulnerabilities like SQL injection.
It's different.
- Dependabot looks for vulnerabilities in your dependencies, and creates pull requests to update you to fixed versions.
- Code scanning looks for vulnerabilities in your own code. So, for example, if you have written code that takes user input and creates a database instruction from it without escaping it, it will flag that you are introducing an SQL injection vulnerability.
(As an aside, we could definitely improve Dependabot to treat devDependencies differently. You do need to care about vulnerabilities in your devDependencies in _some_ cases (code exfiltration is the obvious one) but not in many - we should to get smarter about distinguishing between those cases.)
I'm not sure, how easy it will be to filter the devDependencies. Maybe scan for config files (webpack, babel...) would be a good alternative to manually tag each npm package.
Hey totally off topic sorry. Can you get someone to turn off pull-requests for the unofficial mirrors that you guys created for some open source projects? Users are being mislead into thinking opening PRs there is productive, but they're not monitored by our project and we don't own the repo anyway. https://github.com/wine-mirror/wine/pulls
I've tried contacting your support, but they just tell me that they don't own the repo, which is obviously false[1]. I don't know who to reach out to.
[1] https://github.community/t/how-is-the-mirrored-from-annotati...
The link has a section named Audit Process, which is interesting.
I hope this GitHub feature bring their valuation to the ground and their investors to the reality.
I’m fairly aghast at the state of things, and every bit helps.
Thanks, GH!
Obligatory xkcd: https://xkcd.com/2347/
This sort of project (ought) to level that playing field out. It ought to be clearer to us all who depends on what, and that clarity ought to give true leverage to the valued authors, or possibly point out the OpenSSl like risks we are taking.
(In fact regulators will one day wake up to just this risk, and then this will be a really valuable area to be in)
I don't have experience of the security fatigue and stuff that other people seem to be talking about. Maybe I just write better code, or use fewer, and fewer problematic, dependencies? ¯\_(ツ)_/¯
Anyway I think this is a really cool feature and I'd love to see more of these sort of value added and free features on top of public repos. Is there a place where like you can create your own like a marketplace or something?
Github giving away yet another enterprise class property away for free.
The future of security is very dire and available means like this are essential to protect democracy and freedom.
A lot of those noisy repos do actually run cog wheels of critical infrastructure.
All praise our source code OVERLORD.