Elo scoring two years of Magic: The Gathering games
dylanlott.com
dylanlott.com
Elo is appropriate for chess, where there is no initial game-state variance, and no built-in advantage for either competitor except who goes first; that can be addressed by averaging, by using the results of tournaments where the competitors swap colors, or simply by maintaining a separate Elo as white and as black.
Similarly for Starcraft you can track Elo separately for Terran/Zerg/Protoss. (Technically you would also need to do the same by map, but anyway...)
With MTG, you have a huge effect from the quality of the deck. Unless you have each player play with each deck, there's no way to de-convolute the quality of the deck vs the quality of the player. And if you did have that data, Elo couldn't leverage it -- you'd need a more sophisticated model to account for that statistical effect.
Then there's the game-state variance you allude to... Regardless of how good you are at MTG, and even how good your deck is, you're going to lose a lot of games due to mana flood / mana screw / etc. When that happens to either player, the outcome of the game does not contain useful information about skill. Of course if you sample enough games, you can still figure out what is skill and what is chance, but using Elo with low-count datasets is bound to be misleading because it is designed for games of pure skill, where game outcomes contain information about relative skill levels 100% of the time. Maybe you could establish some rules about what games are appropriate to use as indicators of relative skill, and which ones must be discarded?
Anyway it's an interesting idea. Here's related reading for the MMR score used in Magic Arena:
https://hareeb.com/2021/05/23/inside-the-mtg-arena-rating-sy...
How well a player chooses their deck is one of the factors that determines how good a player is. You can say the same thing about the other games: I'd probably have a better rating in chess if I didn't only play somewhat unsound gambits, and I'd definitely have a better rating in Starcraft if I didn't only do 2port wraith in TvZ.
A deck is something you have. A build order, or a chess opening, is something you _know_ and therefore more or less what I'd be comfortable calling skill.
[0] Let's be clear, Jeff's deck is bad and he's going to lose a lot, even if he's a time-traveling supercomputer with the diplomatic finesse of Otto von Bismarck.
I think the "piloting" ability is mostly (but not entirely) independent of the deck. You can see this most plainly in draft, where everyone is basically playing with a new deck. There are "soft" skills that are contextual and format-dependent, like knowing the cards that you need to play around (white just foretold turn 2, is that a Doomskar? Maybe I shouldn't play a creature this turn, etc). There are "hard" skills that are almost always valid (generally wait until after combat to spend mana, cast instants during your opponent's end step, etc).
But certainly deckbuilding talent is not necessary, because anybody can grab a decklist and head to TCGPlayer.
On that note I'd guess that draft (or sealed) tournaments are the best scenario to measure pure skill using Elo alone, since going into a tournament, everyone has equal chance to open good cards.
Top tier commander decks can easily cost over $10,000 I suspect the vast majority of players do not have access to them.
I think perhaps we could agree on an intermediate claim that Elo could work amongst the pool of players that can build any deck they want, which is a big pool for MTGA, and definitely a smaller pool for physical Standard, and probably very small for Commander.
I wonder if you could decompose the score by playing some games with a fixed deck too? Eg Arena challenges. How much of your overall Elo is just picking the right deck for the local meta, vs raw quality of plays?
You split your deck into two stacks. One with land and one with everything else. For your starting hand, you take 3 land and 4 of everything else.
Each draw phase, you pick which stack you draw from.
That’s it. Everything else stays the same but mana floods / screws completely stop.
One knock on effect I’d predict is higher mana value cards would be substantially more playable. I expect a deck of walls + counterspells + removal + big finishers like the Eldrazi or even just Baneslayer Angel to be much more effective than it is now.
On the other end of the spectrum super low to the ground aggro strategies also get a huge bonus by simply never having to draw a land again.
Probably Storm (play a bunch of cheap spells, typically with a discount or with effects that give you mana when you cast a spell) gets a huge boost as well as they can ensure they never fizzle out. Once the engine is going they’ll always win unless they get countered.
What lose out here are all the decks in the middle. The midrange, “fair” decks that are just trying to curve out with the best play each turn.
And all that’s not counting the rules headache with cards like Oracle of Mul-Daya, Fact or Fiction, Treasure Hunt, or Dark Confidant. Which pile does my Maze of Ith go in? Cultivate? Sol Ring? Faceless Haven?
With that said it’s also my personal opinion that variance just makes the game more enjoyable and widens the group of players you can compete against, as long as you have the emotional capacity to not take losses personally.
I personally suspect that aggro would completely dominate a separate land deck meta if pushed hard enough. But I'm all on board for an alternative game design that invents new interesting questions to ponder, addresses a pain point of mtg design, and most of all makes it more fun for a kid.
Agreed, I probably should have listed it first.
Even in the short term, look at when Arena ran their Treasure events where each upkeep that player would get a treasure token. By the end of the first day the event was dominated by “mono red” decks with 13 land and free splashes.
It's more that it keeps the game fun as he's just getting into it. You're guaranteed that both players are going to have playable draws.
For any setting with more advanced players there would definitely be side effects and a more polished set of rules in place for those special cards and circumstances.
I touched on low to the ground aggro decks in the very next paragraph, and I agree that's probably the biggest issue.
However aggro decks becoming more powerful does not mean control decks cannot also become more powerful. Decks aren't a one dimensional plot of aggro to control with midrange in the middle. (though to be fair I did say "on the other end of the spectrum" in my initial post)
When I say higher mana value cards are more playable, I mean in the sense that they can be cast "on curve" much more often, because if you want to cast a 7 drop on turn 7 without ramp you can choose do so, while with a normal deck you might expect to hit 7 mana between turns 8 and 11. (Not a hard calc, just a gut check)
If you have to answer your opponents' threats 1-1 early you can refill with answers, and if you get ahead with 2-1s or 3-1s you can keep drawing land to play your expensive threat on curve.
You don't often see it happen during high level play.
I used to be rather careless in how I planned the mana of my decks and rarely took a mulligan. I faced mana issues all the time. After putting more planning into my mana base and deciding on a careful strategy for when to take a mulligan I now rarely experience those issues. When I do it is mainly because I break my own rules out of greed when refusing to admit a hand with great cards is too low on mana.
That being said, the inherent randomness of MTG maybe means that in an ill-defined, abstract sense, it takes "more skill" to improve 100 Elo points in MTG than in Chess, because X% of your games have no meaningful decisions so you have fewer places to take advantage of your superior decision-making and, further, this probably has real implications for reasonable choices of K if you're running, say, MTG Arena, but the article is pretty clear that they're not doing anything especially rigorous when picking K in the first place, and honestly (IMO) it probably doesn't matter a whole lot if you're running a Friday night beer league with some friends or whatever.
[0] I agree with the sibbling comment that deck selection and deckbuilding is a large part of what magic players mean when they discuss skill, and it seems very reasonable to allow those things to be included in our model.
What makes this different from blind build order choices in Starcraft? The greed > safe > rush > greed interactions often set one player ahead pretty arbitrarily in the very early game.
What you are describing with the build order exists also. Many Magic decks have a single game plan ("rush", etc) and can only minimally adapt between games in a match (by swapping cards with a 15-card sideboard). The degree to how uneven a matchup is can vary a lot, and some decks are hybridized so it doesn't just devolve in to rock/paper/scissors
Back then they were trying to identify the best magic players in the world. I think it started at 1600.
In the mid 90s it was hard to get locally sanctioned games so I manually tracked games in my college town and used it as my first blog ever to post the scores of local players for a year or so. I wish the site was archived but I never kept a copy when I went to work and abandoned the site. I remember going through the formula to calculate the score and I couldn’t get it to work in excel or javascript so I manually calculated it out on paper for a few dozen games a week. I think back about how much time I could have saved if I just had a little more programming skills.
Even if you have well-defined, sequential 2-player matches, where a widely-used model[0] exists in the literature, there are a wide variety of ways to estimate player ratings from game results, which all have their own assumptions and various tuning parameters.[1] If your domain also includes team-based or multiplayer matches (or some other weird feature that you want to account for), you then get to decide whether you want to try and hack together something using Elo because it'll be "close enough"[5], or whether you want to try and use (or build!) a more sophisticated system which captures those nuances, such as Microsoft's TrueSkill(TM) system[6].
[0] The so-called "Bradley-Terry model" https://en.wikipedia.org/wiki/Bradley%E2%80%93Terry_model
[1] Beyond Elo, which is described in the article, you have things like Glicko/Glicko-2[2], which still does incremental updates to participants after each match but try to track rating uncertainty/player growth in a more sophisticated way, to systems like Edo[3] and Whole-History rating[4], which attempt to find the maximum-likelihood rating for all players at all points in time simultaneously.
[2] https://en.wikipedia.org/wiki/Glicko_rating_system
[4] https://www.remi-coulom.fr/WHR/WHR.pdf
[5] This is (obviously) the approach taken in the article, and IMO is probably the right answer unless you're a huge nerd who's interested in wasting a ton of time for not-much practical benefit.
⸻
1. A tie is not a strictly neutral event in many ELO scoring systems: usually it means that the higher-ranked player loses some ELO while the lower-ranked player gains some, just not as much as in a straight victory.
For team-based play (like with Spades), on Board Game Arena, they treat the partners as having tied which is, I think, incorrect. A better approach is probably to treat it as a match between two players where each team's ELO is the mean of their individual member's averages. The tie approach means that a strong player is penalized for having a weaker partner.
I have a number of table zaps recorded in subjective data, and I've considered a similar approach; treating a game recorded as a table zap as A winning a game against B, C, and D, however you're right, that doesn't handle the case of D being out, and then B and C being zapped at the same time. I think that's a subtle but important distinction.
Re: two player matches - That's a much better approach to my naive interpretation, definitely going to implement that sometime.
It's more accurate to say that the entry cost into the Vintage format is 10k (or whatever the cost of the deck you want to play). You can't just throw money at the deck and increase its winrate unless your deck starts out suboptimal.
1000 or 2000 in equipment is absolutely normal for a lot of competitive endeavors.
This would be a really out-of-the-box way to compensate for the deck quality bias I allude to in my other post -- normalize the effect of the deck on game outcomes by using a static "deck quality" score.
I suspect that coming up with halfway decent "deck quality scores" is an extremely difficult problem, though. It's not much of a leap from there to imagine using a computer to solve for the best possible deck in the format, the implications of which are terrifying for competitive Magic (and would be priceless to card speculators)
For this reason, the "assume a sphere with no friction" joke here is that deck selection, lock-in / mulligan processes, information asymmetry, and turn order, all being assumed to be equal and at that player's local maximum.
And I imagined this all as a service people would pay for, lol.
Actually, it was a bit more abstract: it scored a game state where it considered this players permanents, the opponents permanents, this players life total, opponents life total(s).
It uses the public pairings and results that were published each round for events all the way back to the 90s. Unfortunately, there are less competitive MTG events these days, so most peoples' ratings stop in early 2020, but that's another topic altogether.
my understanding is it handles new players and teams better than straight elo.
My use for it was to keep track of winners during a mario kart tournament and see if it could predict the winners.
it did ok.
I absolutely do! Thanks for the link, I’ll definitely check it out.