Show HN: Unblob – extraction suite for 30+ file formats
github.com
github.com
If you're interested in something similar that can put things back together after you've modified them, check out OFRAK:
https://github.com/redballoonsecurity/ofrak
It's designed with embedded systems in mind, but has support for all kinds of other stuff, too. It also has some very advanced binary patching capabilities.
I work on it as part of my day job.
It looks like unblob has the right behavior by default that I have to alias for `dtrx`:
alias dtrx='dtrx --one=inside'
But I'll probably want to create an alias for unblob to change default depth to 1.Not really in the same class of tools as unblob, but handy to have around regardless.
These blind (or "magic"; https://en.m.wikipedia.org/wiki/List_of_file_signatures) extractors can be handy is a very wide range of applications.
From firmware/device image hacking to reverse engineering to data recovery.
It's in Python and is able to deserialize Unity archives, treating them as a serialization format rather than a simple archive format. Feel free to email me if you want to integrate something like this or you have questions :)
Possible alternative to Universal Extractor. https://github.com/Bioruebe/UniExtract2
As someone who uses `binwalk` extensively in a professional setting, with tooling built around `binwalk`, it would be useful to see (a) how `unblob` would integrate and (b) if it could be a replacement or supplemental.
The problems with this:
- very noisy, finding a lot of false positives (license code, format inside another format, etc).
- very slow, trying to extract irrelevant things
- imprecise, because it finds patterns in the middle of a file, where it's actually not relevant on the first level of extraction
unblob solves these problems by being smarter about the file formats, recognizing them by their specification, for example unpacks format header structs and carves out files based the information in the header (size, offset). See a simple example for NTFS [1].
We also went to great lengths preventing unnecessary work by skipping formats inside another [2]. We are using hyperscan [3] instead of grepping byte sequences with Python, which is orders of magnitudes faster. It can also handle 4Gb+ files because of this which binwalk cannot.
It's used for a year now in production and it's way more precise and faster than binwalk. We are getting less false-positives too, and even if unblob fails to extract everything, we still get meaningful information out of firmwares, where binwalk just failed with no output previously.
[1]: https://github.com/onekey-sec/unblob/blob/main/unblob/handle...
[2]: https://github.com/onekey-sec/unblob/blob/main/unblob/proces...
Does anyone know of something similar for text file formats? In particular something that makes it easy to work with legacy fixed width record file formats?
EDITED to correct autocorrect.
I'm going to pretend I didn't see the XML example but I think I could produce CSV files with this.
Thanks!
tar -xf example.tar -C temp-dir && mv temp-dir/* $(echo temp-dir/* | sed -r 's,.*/([^/]+)-[0-9.]+,./\1,')[1]: https://unblob.org/formats/