Grep.app – Search across a half million Git repos
grep.app
grep.app
This was my first search to show AWS API Keys. [1] and my second to show google API keys. [2] Unfortunately, there are many more of these types of searches - and far too many results.
[1] https://grep.app/search?q=AKIA%5B0-9A-Z%5D%7B16%7D®exp=tr...
[2] https://grep.app/search?q=AIza%5B0-9A-Za-z%5C%5C-_%5D%7B35%7...
There are many people who do these types of searches - regularly. If there is any key material checked into gh, consider it immediately compromised, even if the repo is removed.
Thanks for keeping an eye out.
Can't find a single match for an actual API key.
https://grep.app/search?current=2&q=api_key%20%3D%20%22%5Ba-...
As others have posted, Github implemented an alerting system for keys/passwords/etc. that have been checked into the codebase. My guess is most of these will be invalidated by now, either by the developer requesting a new key (and not publishing it), or the provider invalidating it for them.
Sourcegraph is much better than this site.
https://news.ycombinator.com/item?id=22396824 (614 points | Feb 23, 2020 | 155 comments)
The technology behind GitHub’s new code search - https://news.ycombinator.com/item?id=34680903 - Feb 2023 (168 comments)
Most of my GitHub searches don't actually revolve around actual code-code. I search for text strings in prose that turn up in code searches. So the file types may be text or MarkDown instead.
Perhaps grep.app only indexes and presents actual code-code and filters out all the prose I'd be interested in. Sad, because there are actually a few instances where my searches involve code with many intricate special characters, and GitHub's current code search ignores all those.
- https://news.ycombinator.com/item?id=28424845 (September 5, 2021 — 54 points, 6 comments)
- https://news.ycombinator.com/item?id=22396824 (February 23, 2020 — 614 points, 155 comments)
The current GitHub code search is based on Elasticsearch and indexes more than 100M repositories. Its tokenization is based on whitespace, case changes (like CamelCase) and punctuation (like kebab-case) and strips out non-letter characters like < or { which is why it can't do exact match search.
Our new code search, currently in beta and indexing about 45 million repositories, uses an search engine we built in house that indexes content using a technique we call sparse ngrams. This allows us to execute searches faster than a trigrams index, while also being smaller than a positional trigram index. My teammate discussed some of the technology behind it in our blog post that was published yesterday: https://github.blog/2023-02-06-the-technology-behind-githubs...