I would love to know why the third match turned sour. I suspect (as an amateur with no ML background) that that matchup was under-trained.
Like I could imagine OpenAI getting stuck in a subset of the draft pool for which it trained against, like maybe the top 10 of 18 champs. And then picking outside of that meta causes it to fall back on much less robust training/strategy.