GitHub Codescanning
github.com
github.com
This is the only static analysis tool I've really been interested in over the past few years; it's crazy effective from everything I've seen, and the queries are easy to write. Can't wait to play around with this beta on my own code.
"A member of our sales team will reach out to discuss details" is a great euphemism for "be ready to pay quite a few thousand bucks per year for this feature".
A small company buying a software might actually have a bunch of employees using it.
Nobody wants to spend hundreds of thousands of dollars a year -if not millions- on something that's barely used. Better spend on anything else that is tangible.
That said, Ballmer did get one thing right: https://www.youtube.com/watch?v=Vhh_GeBPOhs
Also, it might vary per account in an unpredictable way, since it's a heavyweight operation that'd cost more compute-hours when run against larger repos, but also involves static-analysis tasks that don't necessarily scale linearly with LOC. So they might just not yet have a predictive model yet for the baseline cost to them, but instead are doing a trial run for each interested user, observing the cost of the workload on their repo, and then extrapolating that out as the cost to run that workload persistently.
Either way, GitHub has been strongly on the side of "clear pricing" so far, so I doubt they would plan to leave things this way. But it's hard to get enough data for a feature like this, when you know each run you do "just to model the curve" is costing you real money.
Sometimes it is also possible to brand the exact same service in different ways.
If you have a SaaS that is fully OSHA compliant on all tiers, it may be worth it to not mention the OSHA compliance on lower tiers, but only offer it on the Enterprise tier for example.
Customer: Sorry, I just don't think it's worth the price...
Sales: Oh really? Because we've identified 42 critical vulnerabilities in your code base
Customer: Do you have an installment plan?
Customer: We know the software runs on java 7. We don't need to pay you 100k to find out.
But I work with a SaaS product and the up front pricing is sort of irrelevant if only because the customers want a highly customized product and frankly you really need to know their business / work with them to find what they want to figure out what the price for implementation and customization is.
In our case it isn't a scam or a ploy for more money, it's just the nature of the beast / industry / and every single time it is what customers end up asking for (lots of customization and etc).
Now maybe that is the case for Github Code scanning, but also maybe it is very much a product that depends on how exactly you want it to work for you and that changes the pricing ...
There are semi-standard XML formats that can be used to provide file position, severity and message, which could easily be produced from CI actions and would give devs and reviewers a great view of failing tests, compiler and linter warnings, ... With links to the files etc.
Instead we are still stuck with scanning text logs to figure out what failed.
I always assumed this is not implemented to eventually upsell something automated.
This is still a great feature though, which will probably prevent a large amount of bugs/vulnerabilities, assuming they can minimize false positives.
To give credit where it is due, I'd also note that most of Githubs new features since the acquisition were already present in Gitlab [1]. Github will be able to commit way more resources to make it polished, though.
[1] https://about.gitlab.com/stages-devops-lifecycle/secure/
Edit: apparently both Gitlab and Github have at least a limited version of this now, although Gitlabs implementation seems much nicer. See below.
The way of reporting seems somewhat odd, and the UI seems a bit limited, but it's a start!
Edit: https://help.github.com/en/actions/reference/workflow-comman...
It can automatically parse most XML formats, JUnit for example.
GitHub Codescanning functionality is best compared to what GitLab has in GitLab SAST https://docs.gitlab.com/ee/user/application_security/sast/ and Secret Detection https://docs.gitlab.com/ee/user/application_security/sast/#s...
This looks perfect.
For now, at least, it seems like this is one reason not to switch to GitHub actions yet.
https://docs.oasis-open.org/sarif/sarif/v2.0/csprd01/sarif-v...
It's really hard to do any kind of static analysis on something like Ruby or Perl where 1) you need a ton of context just to parse it properly, and 2) tracing calls is a nightmare. Given that, I'm completely unsurprised they haven't supported it yet.
Python is very dynamic. I've used and worked on linting tools for python, tried commercial static analyzers, and they do a pretty good job in my opinion in spite of the language being dynamic. Not perfect but miles above anything I would have expected.
Is it because Rust is bullet proof, or too young to be considered for now?
I also use codeanywhere for my personal use and whenever applicable I like to use codesandbox.io when it's JS-ish.
https://docs.oasis-open.org/sarif/sarif/v2.0/csprd01/sarif-v...
We use SARIF as the input format so third party code analysis engines can easily integrate with code scanning. Their results can then be shown in the same way that scans using our own CodeQL analysis engine are displayed.
Docs on how we translate each SARIF property into the code scanning display are below:
https://help.github.com/en/github/finding-security-vulnerabi...
(The beta notice on that page is very relevant here - we wanted to build extensibility options into code scanning from its inception, but whilst it is in beta the API won't be 100% stable. We'll do our best to avoid any unnecessary churn.)
We handle that with secret scanning - code scanning focusses on static analysis of your code to find vulnerabilities in your code, rather than committed secrets.
We have a partnership with AWS (and many other token issuers) that handles this really nicely. If anything that looks like an AWS credential is committed to a public repo we send it over to AWS - if it's a real token they notify the token's owner (and in some cases automatically revoke the key).
There's full details at https://help.github.com/en/github/administering-a-repository....
> We have a partnership with AWS (and many other token issuers) that handles this really nicely. If anything that looks like an AWS credential is committed to a public repo we send it over to AWS - if it's a real token they notify the token's owner (and in some cases automatically revoke the key).
So if something looks like a token from AWS or another token issuer, you automatically send the token to providers to check to see if it is "legit"? Is this something that is opt-in, or done automatically?
I'm assuming AWS does not give a "yeah looks like it", or "nah" response -- but rather "thanks, we will look into it" and then if it's a real one the rest is directly with their customer.
That way no sensitive information would leak between the providers
If you accidentally upload a key, but then immediately notice and force push, you're already too late since GitHub took the initiative to share that. I get that the user would be at fault here ultimately, but that doesn't mean that GitHub should be working against the user in sharing that.
What if it isn't an AWS token, but instead an encryption key or SSH key that you have blocked off to the public so you're not too worried about it but you're a warehouse worker protesting COVID-19 treatment. Now Jeff Bezos will be looking for dirt on you like you're Michael Sanchez.
If they made the detection information public then it would at least provide some transparency to see what they determine to be AWS-specific.
How do you "block off to the public" something committed to a public GitHub repository? The OP specifically said this was for public repositories.
If GitHub weren't doing this, I imagine the AWS security people would be crawling GitHub on their own, to cut down on security incidents. This push mechanism just makes it more efficient for both GitHub and AWS.
If Amazon is looking for dirt on you, and you have public repositories, you can bet they'll be looking deeper into your repositories than a quick credential scan.
Pretty simple actually. If your AWS services are internal only, then you may not be in a rush to change keys to a service that isn't exposed.
In practice, it pretty much does - bad actors continuously scrape the GitHub firehose looking for AWS secrets, and then automatically spin up EC2 instances to mine cryptocurrency. GitHub's token scanning just ensures that AWS sees the tokens too.
If you don't believe me, keep this website open for a few hours - it's a realtime stream of secrets scraped from GitHub: https://shhgit.darkport.co.uk/
Please also note that we only automatically send details from _public_ repos to our secret scanning partners.