Sure, just enumerate any and all possible types of sensitive data, the format they may be in, regex / matching functions to account for them (supported across 20+ programming languages) and I'm sure Github will have that done asap.
Alternatively, don't commit passwords/API-keys/sensitive-info to your repo.
This is, of course, the right answer.
However, it's frustrating that several frameworks make this very easy to get wrong. Anything that has an application.yml, database.yml, or similar configuration file that normally lives within the same directory as the source code, and which is intended to contain credentials, means that lots of people will make that mistake.
It's one of the fundamental errors that you see so often in web frameworks like Rails, this whole idea of mixing application code and configuration files into one big tree. I don't know how this practice caught on, but it has, and it so frequently causes mistakes, both big ones like accidentally publishing credentials, and simply frustrating ones like confusing the difference between application code and configuration in ways that make it more difficult to have multiple local configurations and updating the application code independently.
False dichotomy. It doesn't have to be "everything" or "nothing". An 80% solution here is better than nothing.
I still find it useful that gmail warns me before I send an email without an attachment if I've written "I've attached" in an email. Can gmail detect with 100% accuracy if I intended to send an attachment? Of course not. But the 80% solution here still does me and lots of other people a lot of good.
There will never be a 100%-fool-proof "Did you mean to commit this sensitive unicode string?" but getting in front of that with a, "ok, I've checked my code, ran my tests, pruned the sensitive data, is there anything I'm missing?" will go a long way both in present and future times.
There's an issue with your example. Google looks for the substr 'attach', but does it know that a file with a string 20 chars from the newline 50 lines deep with two single quotes is actually your root password? There's a world of difference between keying off of a word/phrase and understanding the context of a larger document+metadata. Trusting computers to do the latter will result in bad times for all, while learning for yourself can help spread the knowledge to those that are both technically inclined or not (infosec is everyone's issue!).
No one's suggesting anything like that. It's just a basic protection that Github could offer its users, because its users are humans, and humans make mistakes.
No matter how clever you get with your pattern matching, you're going to have to always play catch-up with Web Framework N+1's format / weird-ass package manager. The dual approach is the only sensible approach because it expands your coverage. The important thing to this is to internalize the knowledge learned from 3rd parties and integrate that into your native process / tools, but not everyone will do that.
It'd be grand if we could say, "yeah, they should take care of my security for me because I'm paying them" but reality is a bitch. It doesn't matter what they were 'supposed' to do if there was an infosec leak or attack, you can't ctrl-z that (set of) event(s) and that info is out there, so it must be fixed (re-roll credentials, regen keys/certs, etc) and accounted for next time. There really is no 'getting ahead' but 'being less behind' will at least help mitigate getting eaten from the herd ;-)
It's kind of amusing you are arguing that it's impossible to play this game, even though that's exactly what the perpetrators are doing, they are automatically detecting API keys and harvesting the code... maybe their script is hosted on github?
Any time you want to show me a 100% future-proof algorithm for sensitive-info detection that works across any/all code on github, I'd be happy to toss my hat in and say, "I was wrong", until then, people will never ever beat 0days they don't know exist (0day being more than just a SW exploit). Just do.not.commit.sensitive.info.to.github. Period. That is the only sure way to not mess it up. Software only executes what is in the code, regardless of how nonsensical it is (aka, your code will not save you from messing up, something something something, PEBKAC)
I'm not arguing it's "impossible to play this game" It's 100% possible to play it when and how you'd like. I'm discussing the rate of "did I win (read: not get pwnd)?" It's cat and mouse of automation for vulnerable/sensitive info ... but all of that is rendered moot if you ... wait for it ... don't commit it to Github which would mean it wouldn't get to Github's API which means it wouldn't appear in 3rd party services sucking the firehose from Github cloud-y silicon teet.
And expanding on this ... committing your passwords and sensitive info to your code-repo is so misguided it's actually funny. What happens if you have an employee and they go off the deep end? Whoops, gotta rotate all those passwords/credentials/un-fuck every branch/resync dev's machines/etc. Keeping that sensitive info in a private, self-hosted, well-maintained internal repo (with strong ACLs, especially wrt server/hosting environments) will go significantly further for your team's security than submitting a feature request to a company to stop you from making arbitrary mistakes every so often.
Yes, Github can/should help, but developers should not think they're owed it just because they constantly check in sensitive info to a website, that's all.
(I'm just applying your point to another kind of 'security' that is provided to you).
You are basically arguing the "We should all live in the woods, and hunt and kill our own food, because relying on other people is fraught with danger" line.
Or we could accept mistakes are made, and provide warnings/undos/etc. Kind of like how cars have airbags, even though they are rendered moot by just "not crashing your car"
The point were you define it is sensitive, but the place where you use it are not.