Tangentially related: I am currently scoping out an idea for how language models could be used to augment decompilers like Ghidra.
At a surface level, this was partially an intellectually interesting project because it is similar to a language translation project, however instead of parallel sentence pairs, I will probably probably be creating a parallel corpus of "decompiled" C code which will have to be aligned to the original source C code that produced the binary/object file.
Then I realized, the only way I could reasonably build this corpus would be by having some sort automated flow for building arbitrary open source C projects...
Perhaps I will attempt this project with a Go corpus instead.