Reverse engineering my router's firmware with binwalk
embeddedbits.org
embeddedbits.org
Another cool tool I learned about recently is signsrch. It's more for reverse engineering binaries of software that implements encryption of some type. It'll find signatures in the binaries of these encryption methods, giving you a place to look when, for example, reverse engineering a file format that you suspect is encrypted in some way.
https://www.oreilly.com/library/view/learning-malware-analys...
So the tool was called Golem. It had tables for defining opcode to assembler pattern matching, that could be written for any machine (instead of just the one I was cracking).
It worked iteratively. You ran it over the binary once, it produced arbitrary labels from jump-points. You could annotate that output by changing the labels to something human-readable (e.g. Loop-back, Main, TimerISR etc) and add comments.
The next iteration would read that back in to build a symbol table, rescan the binary and re-output. But this time it would understand that the symbols were always on opcode boundaries, distinguish data table from code entry points (because you marked them) etc. So it would do a better job of staying in sync with the code.
Once I was done with that project (and had re-compilable source for the radio module) I put it away and never thought of it again.
Ghidra is old too, although only recently public. It couldn't be older than Java, which is from 1996.
Are they iterative? Can you add human clues/cues so they do a better job the next time?
IDA Pro supports dozens of processor architectures. I count about 70, not including model variations and not including community support. https://www.hex-rays.com/products/ida/processors/
Ghidra supports "X86 16/32/64, ARM/AARCH64, PowerPC 32/64/VLE, MIPS 16/32/64/micro, 68xxx, Java / DEX bytecode, PA-RISC, PIC 12/16/17/18/24, Sparc 32/64, CR16C, Z80, 6502, 8051, MSP430, AVR8, AVR32, and variants of these processors."
Binary Ninja officially supports x86, x64, ARMv7, Thumb2, ARMv8, PowerPC, MIPS, 6502. Community support adds AVR, MSP430, and VMNDH-2k12.
Hopper Disassembler supports "x86{16,32,64}, Dalvik, avr, ARM, java, PowerPC, Sparc, MIPS"
it is interactive (so by definition iterative)
Then IDA went on Windows and today it's multiplatform.
In the first example he uses the "--signature" and "--term" flags, these are unnecessary. Running binwalk with no flags will produce the same output.
To extract part of the file, he also uses dd with the "skip" and "count" options painfully calculated. You can just use:
binwalk --dd='.*' img.bin
and it will extract everything that matches the pattern - the pattern above will extract all found files.
I have to work with some old structural analysis software. The material and element definitions come in an obscure file format ".PF3CMP". I know it contains text like the material names, and numbers/letters for the material properties.
Ultimately its my goal to be able to write these files from matlab or python, instead of using the horribly clunky user interface. But first I need to know the structure of the file, and I'm not even sure how to begin figuring that out.
[0] is what it looks like when opened in a hex editor
It uses some techniques that might be relevant, like monitoring different parts of a file as you make different changes (like accelerating or decelerating). In your case it might be possible to compare between different material definitions for example.
[0]: https://kaitai.io/
A good test would be if you can name/tag/comment items in the file, you can search for these strings.
I'll also echo the other comment about reverse engineering the reading functions. Some formats only include certain structures if necessary so even if you have a lot of files you might be missing some example data to complete the picture.
It might be easiest to just start writing a utility that parses it, first making guesses and then refining as you generate and test more files like you mentioned in another reply. You already know what the magic bytes are at the start of the file - PF3CMP.
You can get it with WSL on Windows, or even just install git and you’ll get git-bash for another easy option.
https://techdocs.broadcom.com/content/broadcom/techdocs/us/e...
I've contacted the developer but they will not release the format of the files to me.
Quite a good, short intro into the subject as well!
Using binwalk for CTF challenges is actually a new insight for me :)
This is what happens whey you pay peanuts for embedded devs and outsource development to the cheapest sweatshop you can find so your products can meet a competitive price point.
Sadly this will not change until there's regulation in place to hold manufacturers accountable for their massively obvious vulnerabilities since nobody cares that they're flooding the market with potential botnet hosts when they're overworked, paid miserably and have a manager constantly breathing down their neck.
So it's about $400.00 for a router that has updated firmware(pfSense). Or you can cheap out and spend only $100.00. This is what you get by doing that.
You should also purchase a device that includes enough storage space and RAM to support more than the bare minimum; that will help keep things future proof.
23296 0x5B00 LZMA compressed data, properties: 0x5D, dictionary size:
8388608 bytes, uncompressed size: 97476 bytes
64968 0xFDC8 XML document, version: "1.0"
So it looks like the size of the bootloader should be 64968 - 23296 = 41672. But he extracts 41162: $ dd if=archer-c7.bin of=u-boot.bin.lzma bs=1 skip=23296 count=41162
Curious if anybody knows why 41162; is this a block-size alignment requirement?At the step where they remove the header with
dd if=uImage of=Image.lzma bs=1 skip=72
It results in a file that if I try and un compress it with `unlzma Image.lzma` it complains with "Compressed data is corrupt"I don't know where the magic number "72" comes from. Is it likely that could be different on my machine (a mac)?
[edit: I think there's something else wrong - if I use `mkImage` to examine the uImage file I only get:
mkimage -l uImage
GP Header: Size 27051956 LoadAddr 78a267ff
Instead of image information]So you'll need to check what binimage says about your image, the uImage header isn't necessarily fixed in size. Also see the comment above about the --dd switch, though mind the reply to that pointing out you might want to check what it finds before just letting it write a pile of files.
Relevant quotes: "By using the Products or Services in any way, you agree to the Terms. " "Also, modifying, translating, adapting, or otherwise creating derivative works and improvements, decompiling, decoding, reverse engineering, disassembling, or otherwise reducing the code used in any software in connection with the Services into a readable form in order to examine the source code or construction of such software and/or to copy or create other products based (in whole or in part) on such software, is prohibited."
i actually prefer to run Tomato, but archer c7 is not broadcom :(
can anyone offer advice about dd-wrt vs openwrt (considering trying openwrt).
What reasons do you have to stay on dd-wrt?
mostly that i've used it before. can i gui-flash to openwrt from dd-wrt? i've done tftp flashes before but they're pretty fiddly with getting the stupid 30-30-30 or whatever timing right. also i think these routers try to "pull" from a tftp server rather than having you push to one that they bootstrap - i've never been able to get the "pull" variant to work.
would be hell of a lot easier if the router could be booted into something like android's (arm's?) fastboot or flashmode mode so i can just push an image.
Openwrt also has a handy failsafe built into a lot of models. It boots a stripped down http server where you can upload recovery firmware.
Used to swear by dd-wrt, now I prefer openwrt.
That’s how I flashed from stock to OpenWRT on 3+ Archer units anyway. Make sure not to keep settings.
Not being Broadcom is a very good thing.
Atheros and Intel I believe both have good open-source support.
I put openwrt on my c7 V5 and could barely get any bars.
Flashed back to the stock and was back in business.
Another thing I've read is the third party firmwares don't get hardware access to NAT resulting in speed hits.
Cheers
> Another thing I've read is the third party firmwares don't get hardware access to NAT
i read that too :(
Here is the exact `factory-to-ddwrt` image I used (this will depend on which version you have): ftp://ftp.dd-wrt.com/betas/2019/10-15-2019-r41328/tplink_archer-c7-v2/
Does it mean I downloaded corruption zip file from TP-Link site? How I can extract kernel image? Binwalk says about Image.lzma: 0 0x0 LZMA compressed data, properties: 0x6D, dictionary size: 8388608 bytes, uncompressed size: 3164228 bytes
That would be pretty hilarious if it was true.
For the vendors with access to closed-source drivers and chipset info they can likely support devices not supported on the open source packages.
Edit: Per Wikipedia, "Qualcomm's QCA Software Development Kit (QSDK) which is being used as a development basis by many OEMs is an OpenWrt derivative"
It also notes Ubiquiti's wireless router firmware as being derived from OpenWRT, but I thought I remembered discussion of Ubiquiti being derived from a different open source distribution - unless perhaps the routers and wireless devices don't share a code base.
Looking into the equivalent firmware[1] for my Archer C7 v2, I didn't find any OpenWRT bits though. I was honestly a little bit disappointed.
I guess the difference between hardware revisions might be more fundamental than I assumed.
DECIMAL HEXADECIMAL DESCRIPTION
--------------------------------------------------------------------------------------------------------
0 0x0 TP-Link firmware header, firmware version: 1.-15188.3, image version: "",
product ID: 0x0, product version: -956301310, kernel load address: 0x0,
kernel entry point: 0x80002000, kernel offset: 16384512, kernel length:
512, rootfs offset: 855873, rootfs length: 1048576, bootloader offset:
15204352, bootloader length: 0
71520 0x11760 Certificate in DER format (x509 v3), header length: 4, sequence length: 64
98560 0x18100 U-Boot version string, "U-Boot 1.1.4 (Mar 5 2018 - 13:57:29)"
98736 0x181B0 CRC32 polynomial table, big endian
131584 0x20200 TP-Link firmware header, firmware version: 0.0.3, image version: "",
product ID: 0x0, product version: -956301310, kernel load address: 0x0,
kernel entry point: 0x80002000, kernel offset: 16252928, kernel length:
512, rootfs offset: 855873, rootfs length: 1048576, bootloader offset:
15204352, bootloader length: 0
132096 0x20400 LZMA compressed data, properties: 0x5D, dictionary size: 33554432 bytes,
uncompressed size: 2451644 bytes
1180160 0x120200 Squashfs filesystem, little endian, version 4.0, compression:lzma, size:
9878520 bytes, 789 inodes, blocksize: 131072 bytes, created: 2018-03-05
06:16:10
[1] https://static.tp-link.com/2018/201806/20180611/Archer%20C7(...https://openwrt.org/toh/tp-link/archer-c7-1750 (Scroll down to the Info Links table and the Wikidevi Info column)
v1 to v2 upgrades the Flash (8MB to 16MB) and uses a slightly different AN+AC wifi chip. v2 and v3 seem pretty similar at a glance. v4 is rated at 12v 2a rather than 2.5a; using a completely different BGN(2.6ghz) chip and also different ethernet chip/switch. v5 is lower power still at 1.5a, but it's less obvious where that change happened due to lack of pictures. A guess based on the simpler antenna list is that it uses less antenna.
It's Open source too for anyone that wants to run it.
image name: "MIPS OpenWrt Linux-3.3.8"
I would say you are true.
Embedded device will be hard coded to look at a fixed point and start booting from there, there’s no UEFI. How will you ensure boot-loaders get unpacked precisely where they need to be?
And that doesn’t even touch the idea of having a router understand a file system before any firmware code is loaded.
Routers really are quite different from PCs.
AR7 platform, for example, the MIPS core runs a small ROM that initializes RAM, then reads some blocks from flash. Not sure how much code you'd need to unpack a tar.gz but completely possible.
In the past, each of those would be a separate MTD partition with a seperate device file. You just dd them over those files.
One of these days I'm going to log in to the admin interface and find candy crush installed.
1. For M4s your storage is typically some kind of SPI flash which doesn't act like the traditional desktop flash you're dealing with. You have to manually specify the address you're reading/writing & you have to do it on block boundaries (multiple KB). You're generally looking at 8-64MB. 2. For M0 your storage is typically flash built-in with potentially even more restrictions. 3. These devices have very little RAM. Decompression means you have to have a way of enforcing constraints on the amount of space you'll need. Aside from the space needed regularly for decompression you may need to buffer the decompressed content in-memory to align with block boundaries. All of this means development time, increased costs & risk for something you may not be able to pull of.
If your vendor actually internally compresses their image then great but generally they don't for all the same reasons (+ sometimes this is touching ROM code in the chip).