This is an interesting thought! A couple of other scattered thoughts I had about this:
- Engine evaluation of a leaf of the tree will always be different and more sophisticated than human heuristics. So there's a problem where a human can't be expected to follow down some lines. Of course, this is always changing, as humans seek to understand engine heuristics better. Carlsen's "blunder" at move 33 was a good example of this, from my memory.
- Maybe there's a difficulty metric like "sharpness", some function of the number of moves which do not incur a significant centipawn loss. Toward the end of game 6, Carlsen faced a relatively low sharpness on his moves, whereas Nepomniachtchi faced a high sharpness, and despite the theoretical draw, this difference will prove to be decisive between humans. This seems like it could interact in interesting ways with your difficulty metric - for example, what does it mean if sharpness is only revealed at high depth?
- It would be interesting to take the tree generated by stockfish, and weight the tree at each node by the probability that a human player would evaluate the position as winning. Then you could give a probability of ending up at each terminal position of the tree. Maybe some sort of deep learning model trained on players previous games? Time controls add such a confounding factor to this, but it would be so interesting to see "wild engine lines" highlighted in real-time.