When a linker is invoked to generate an executable from a bunch of object files, it will lay out their sections in memory, compute the addresses of the symbols in the virtual address space and apply the relocations based on the final addresses of the symbols onto the section bytes.
The trick to delinking is figuring out where those relocations were applied in order to undo them and get back relocatable bytes. Then, you create relocation tables based on what you've just undone as well as a symbol table, package it all and you'll get an object file.
The really tricky part is the analysis for spotting the relocation spots. I'm leveraging Ghidra to do the bulk of the work, but it still requires some work to convert references into relocation spots (fairly easy on x86, nightmarishly difficult on MIPS) as well as collecting all the required data and serializing the object file itself, hence this extension to automate all of that.
The big thing is relocations. Object files have granular relocations that executable files don't; at least on Windows, executable images just have minimal relocations that point to addresses of code and data that will need to be fixed up if the executable is relocated, but the executable image itself is only able to be relocated with regards to its image base address as a whole unit. In contrast, object files contain symbol-level relocations. To be able to accurately reconstruct this information, you need to annotate the disassembly with somewhat accurate information about symbols.
The other big difference with an object file is well, it is not linked. None of the symbols are "resolved". This is particularly easy to fix actually: during delinking if the symbol is outside of the current scope then it just needs to be replaced with an unresolved symbol, pretty much. Then when relinking, another object file or library needs to provide the symbol so that it can be linked back up.
(There are a few smaller differences too, like the lack of an entrypoint, but it really isn't a whole lot of significance.)
This means that the boundaries for which you carve object files out of an executable image or shared object is actually completely arbitrary. It obviously was segmented into translation units when it was originally compiled, but nothing really cares about those boundaries at the linking stage. (Of course, you probably want to try to figure it out if possible, since it will probably be very hard to do a matching decompilation if your object boundaries are incorrect.)
Take this with a grain of salt as despite having literally worked on this problem I feel like I might be messing up some of the details a bit. I've been meaning to write a blog post about object files, though there are actually a couple of good ones already floating around.
Can you expound on this? Why are the boundaries not important at the linking stage - aren't you linking the wrong code then? Or did I not understand the point here?
How you decide to slice up the original program is up to you. It doesn't have to follow the boundaries from the original object files.
On 32-bit and 64-bit Windows (both NT and 9x), EXE, DLL and OBJ files are all COFF. EXE/DLL files stick an MS-DOS stub and the "PE\0\0" header before the COFF header, OBJ files start with the COFF header directly.
This is different from DOS, 16-bit Windows and OS/2, where the standard object file format (Intel OMF) was completely unrelated to the executable formats (MZ for DOS; NE for 16-bit Windows and OS/2 1.x)
That said, I do not consider PE to be similar enough to COFF to warrant calling them the same format. Aside from having completely different headers, PE has many additional structures that don't exist in COFF, and even the structures that are shared are fairly different between PE and COFF. PE barely uses anything from COFF, and it doesn't even use it the way that COFF does. Of course ELF objects also have some differences between object file and shared object/executable, but it's at least very clearly the same format the whole time, just utilized a bit differently.