Elo is appropriate for chess, where there is no initial game-state variance, and no built-in advantage for either competitor except who goes first; that can be addressed by averaging, by using the results of tournaments where the competitors swap colors, or simply by maintaining a separate Elo as white and as black.
Similarly for Starcraft you can track Elo separately for Terran/Zerg/Protoss. (Technically you would also need to do the same by map, but anyway...)
With MTG, you have a huge effect from the quality of the deck. Unless you have each player play with each deck, there's no way to de-convolute the quality of the deck vs the quality of the player. And if you did have that data, Elo couldn't leverage it -- you'd need a more sophisticated model to account for that statistical effect.
Then there's the game-state variance you allude to... Regardless of how good you are at MTG, and even how good your deck is, you're going to lose a lot of games due to mana flood / mana screw / etc. When that happens to either player, the outcome of the game does not contain useful information about skill. Of course if you sample enough games, you can still figure out what is skill and what is chance, but using Elo with low-count datasets is bound to be misleading because it is designed for games of pure skill, where game outcomes contain information about relative skill levels 100% of the time. Maybe you could establish some rules about what games are appropriate to use as indicators of relative skill, and which ones must be discarded?
Anyway it's an interesting idea. Here's related reading for the MMR score used in Magic Arena:
https://hareeb.com/2021/05/23/inside-the-mtg-arena-rating-sy...