"Your first assumption might be that geod has hard-coded specific algorithms for each supported game, training the emulator in how to detect specific objects and create good-looking 3D models for them."
I don't know why anyone would assume that, given how poorly it works. My current assumption would be that game-specific algorithms would be the only way to make this work given how ambiguous NES graphics are. I suppose a top-tier neural network expert could make a game-agnostic algorithm... but it would probably handle a frame per minute on a multi-thousand-dollar machine.
I wish someone would come up with better filters for emulating CRT glow so the game would look like a series of photos of the game being played on a top-of-the-line CRT (PVM / BVM). But that would probably be way too slow, too. Actually I have an idea of how to do it...
If I could get an array of pixels per frame into Unity or some other game engine and make each pixel a light... or even just make each pixel a partially transparent shaded sphere bitmap and multiply them together where they over lap... it would probably run too slow but I once made something like that and looked pretty good and ran plenty fast albeit at a sub-NES resolution.