An automated tool to check all depency changes for suspicious code and flag that to a human could be valuable. Whether or not that could be done in a way where wading through the false positives is worth it, I'm not sure, but it's a reasonable idea.
They have a lot of expertise in this, it feels like they could branch out in finding malware in source code, instead of in binaries.
It would be better to run the code in a sandbox and log all the syscalls it makes, similar to how Cuckoo's malware sandbox works.
In real-life it is not of much relevance. I haven't seen a practical case in my life where I would conclude that it is undecidable to say what a function does (happy to see some practical examples if someone has some). For a theoretical program that "uploads credentials iff a sub-program halts" you would probably block it and live with a possible false-positive.
Any and all functions that involve templating/plugins/metaprogramming/etc (common in various web frameworks) that can effectively run arbitrary code and thus are undecidable. Also any function that does deserialization and thus will instantiate novel objects with their code, e.g. Python code that uses pickle will often be undecidable as its execution will depend on what exactly is unpickled.