511 karma · joined February 7, 2018
I suspect that the actual execution time would be far faster with the A* implementation than any of the transformers.
For textual data types, 'file' gets confused often, or doesn't give a precise type. GitHub's 'linguist' [1] tool does much better here, but is structured in such a way that it is difficult to call it on an arbitrary file or bytestring that doesn't reside in a git repo.
I'd love to have a classification tool that can more granularly classify textual files! It may not be Magika _today_ since it only supports 116-something types. For this use case, an ML-based approach will be more successful than an approach based solely on handwritten heuristic rules. I'm excited to see where this goes.
Sometimes a file has no extension. Other times the extension is a lie. Still other times, you may be dealing with an unnamed bytestring and wish to know what kind of content it is.
This last case happens quite a lot in Nosey Parker [1], a detector of secrets in textual data. There, it is possible to come across unnamed files in Git history, and it would be useful to the user to still indicate what type of file it seems to be.
I added file type detection based on libmagic to Nosey Parker a while back, but it's not compiled in by default because libmagic is slow and complicates the build process. Also, libmagic is implemented as a large C library whose primary job is parsing, which makes the security side of me jittery.
I will likely add enabled-by-default filetype detection to Nosey Parker using Magika's ONNX model.
(1) disable creating new `code` objects directly from Python. This probably would break lots of things.
(2) Add a bytecode verification mechanism that would reject `code` objects whose bytecode would result in memory errors when executed. This could be a lot of implementation work; I'm not sure.
Interesting and largely unknown trivia: it's possible to invoke memory errors in the underlying C interpreter from pure Python code — no libraries and no imports needed!
One way of doing this is by creating new `code` objects with crafted bytecode. There is no bytecode verifier in Python to make sure, say, referenced stack variables in the VM are valid...
Some high-profile cases where credentials were leaked on public GitHub: Uber in 2014 and 2021 [1, 2] and Twitch in 2021.
> Search is still possible using google and other methods
Yes, you can search with Google or other sources. But the thing is, pretty much only GitHub has easy access to all the code present there, readily available to search within minutes of pushing.
You could try mirroring GitHub yourself, but you'd need enormous disk space and bandwidth, and would quickly hit rate limits. You also wouldn't be able to do fast full regex search like you can with GitHub's search, as you don't have their search infrastructure.
Hackers are aware of this and do make use of GitHub's search to identify possible leaked credentials. I recently experimented with uploading a dummy AWS credential pair to a public Git repo, and saw that numerous IPs started trying those credentials less than 5 minutes after I pushed.
> any theoretical gains from forcing login are dangerous to count on as vulnerable projects are still vulnerable
Indeed, requiring login is not going to _solve_ the problem. But it could lessen the impact of leaked credentials at large scale, by making it more difficult for automated systems to harvest them.
Perhaps more significantly, requiring login could give better audit trails in incidence response situations, as the logs would indicate which accounts were searching for secrets.
> I think the solution to projects with hard coded credentials are to make the credentials easier to find and exploit so they fixed more quickly after being created
Yes, this can help! There are several other companies that specialize in secret detection. I've also written Nosey Parker, a fast regex-based detection tool that has higher-precision rules than similar tools [4]. GitHub also has its own offering in Advanced Security to address this problem.
[1] https://www.reuters.com/article/uk-uber-tech-lyft-hacking-ex... [2] https://www.securonix.com/blog/securonix-threat-research-ube... [3] https://news.ycombinator.com/item?id=28770590 [4] https://github.com/praetorian-inc/noseyparker
But yes, there are multiple good explanations for why they would lock down the API.
I suspect the change has more to do with security concerns (and having an audit trail) than performance.
It's still not great though, since that only pins version numbers, and not hashes.
You probably don't want to manually generate requirements.txt. Instead, list your project's immediate dependencies in the setup.cfg/setup.py file, install that in a venv, and then 'pip freeze' to get a requirements.txt file. To recreate this in a new system, create a venv there, and then 'pip install -c requirements.txt YOUR_PACKAGE'.
The whole thing is pretty finicky.
It's a packaging/release engineering limitation more than a language incompatibility thing.
On the one hand, Nosey Parker is effectively a special-purpose `grep` with a bunch of security-relevant patterns built-in, including one for PEM-encoded keys: <https://github.com/praetorian-inc/noseyparker/blob/main/data...>
On the other hand, to naively run the check you describe, you would need access to a copy of all of GitHub, which isn't feasible.
What you can do with Nosey Parker is use its GitHub enumeration features to specify your GitHub organization and a list of GitHub usernames you are interested in, and scan against just those. This will implicitly list all the relevant public repositories, clone them, and scan their entire history.
For your use case, another thing you could do is use the new GitHub code search (<https://cs.github.com>) to regex search for particular keys or tokens. That new search seems to cover lots of the public content available on GitHub.
Also, to put some color on this use case: in offensive security engagements (aka "red team" engagements) at Praetorian, we frequently find leaked credentials or tokens on GitHub or elsewhere, which allow us deeper access into the client's systems. It's a significant problem.
The development so far has focused on security-related uses, especially finding hardcoded credentials in source code and log files. For the serial number and license spelunking you describe, Nosey Parker would need a couple more pieces to work well:
(1) Regex rules that would match serial numbers. I think it would be tough to write a regex for this that would match precisely and with good recall, i.e., producing few false positives and few false negatives. Not a showstopper, but not ideal.
(2) Built-in unarchiving support. Nosey Parker will handle textual files on disk just fine, but won't do useful things for many binary file types such as zip files. Your old external drives probably have tons of non-textual files like that. I would like to eventually add built-in unarchiving support to Nosey Parker; this would allow a way for textual content to be extracted from binary files and then scanned like the rest.
OTOH, things like parser combinators are much more ergonomic in Haskell than Rust.
There was no cluster to set up or administer using DuckDB.
I have successfully used DuckDB like above for preparing an ML dataset from about 100GB of input.
DuckDB is undergoing rapid development these days. There have been format-breaking changes and bugs that could lose data. I would not yet trust DuckDB for long-term storage or archival purposes. Parquet is a better choice for that.
> Strangely, these bugs were found by the CI of ClickHouse, and not by any of the hundreds of other products using these libraries.
A lot of software projects that have "OSS Fuzz integration" have very naive or superficial fuzz harnesses that couldn't ever get thorough coverage.
If the code is never executed by the fuzzer, bugs won't be found by fuzzing.
In your case, fuzzing your application code, you likely end up executing code in those dependencies that their own fuzzers miss.
This is separate from the 30% App Store commission.