Fine-tuning Mistral 7B on Magic the Gathering Draft
generallyintelligent.substack.com
generallyintelligent.substack.com
One thing it did make me think about was that these models are suitable for things that don't have a natural definitive answer. That is, picking the perfect card given a set of picks is probably combinatorially impossible to solve. But picking a good card given a set is possible and LLMs can approach human level performance.
I think this leads to a set of problems that current LLMs may be fine-tuned to solve.
I'm guessing LLM decision is rather average, but that the LLM has no easy way of spending the extra time to gather information around said high stakes decisions like a human would.
Just like like LLMs, some humans are better tuned than others for specific tasks, as well as in general.
With high-stakes decisions, you're surrendering the decision-making power to the AI because you don't understand the output well enough to verify it.
Basically, and AI can give ideas but not advise.
Where are the dates in whole foods? (A: with nuts, not fruits and veggies)
How can I steam bao without a steamer basket? (A: saucepan, 1" water, balled up aluminum foil, plate, baos, lid)
Any guess as to when this photo was taken? It looks like anywhere from the 70s to the 90s. (A: the photo paper has a logo that postdates a 2003 company merger)
I had to buy some spare tools where I cared more about price than quality and it helped me choose some suitable brands
As mentioned, you can tell it a bit about a person (and feed in their wishlist if they have one) and it'll help you pick something they'll probably like
Finding something to do to spend an afternoon in a city while traveling
In general, anything where there is no objective best answer (meaning I can ask it to generate multiple possibilities and filter out the bad ideas) and where I value speed over correctness.
I've been going ham with this. Pay for ChatGPT Plus for work, so GPT-4's been helping me design encounters, plan sessions, and brainstorm ideas for one-shots. It gives me a sort of bland, vanilla idea for something, I suggest a twist on it, it goes, "oh great idea, here that is with your changes:" and I iterate with it from there.
Likewise I love theming characters, plotlines, and settings after songs, bands and albums, so I'll dump in a bunch of lyrics and ask ChatGPT to help me intertwine subtle references into descriptions, names, and plot points.
Room to grow.
*observations from the recent karpathy llm talk.
Surely, but we can't gloss over the fact that this was accomplished by a single person.
Consider a PM involved in this project, feeding in requirements from a business. Instead of the "just get it done at any cost" mentality of a single person you would have KPIs and business objectives that would muddy the water.
I just mean to say that there is a gulf between what can be done by a single hacker in his basement when they have no constraints other than their imagination compared to what can be accomplished by a business. Sometimes the single-hacker achievement doesn't scale.
So, it is impressive that this is possible for a single person at all. But, from a business/operation perspective, I don't actually think that is as relevant as it may seem.
Do you mean that you are looking at the draft picks from https://www.17lands.com/leaderboard and then sorting by Win Rate? Didn't you mean to choose Match Wins or Trophies? Otherwise, you're not measuring the best players on the service. You're training on draft choices where most choices were very good - i.e., win rate sort will show you the luckiest players, not the best ones. That will naturally show up in any validation or testing you do too.
Shouldn't this be compared not to an LLM baseline, but to a baseline where an "Elo" style score is computed for each card compared to others from the 17lands data; then, until you have two colors, suggest the best scoring card, or when you do have color(s), suggest the best scoring card within that color or a land?
I think it is possible for the LLM to have some semblance of rules knowledge, but it is more likely that it is picking up on card rarity, costs and "Big" more than anything else for unseen cards.
Your "accuracy" on the draft seems poor. I'm not sure it means what you think it means. Are you saying that when looking at the high win rate choices, where all the choices were mostly good, you happened to pick the choice that isn't the same as the player who originated the data? It actually seems harder to make a choice among all good choices.
Anyway, there is quite a bit going on here.
Ahh no just unclear in the post, I'm filtering to players in 17lands with a > 62% match win rate who are drafting at a high ranking (>=diamond rank). I look at all of those players' drafts though, even the ones where they do poorly.
> Your "accuracy" on the draft seems poor. I'm not sure it means what you think it means. Are you saying that when looking at the high win rate choices, where all the choices were mostly good, you happened to pick the choice that isn't the same as the player who originated the data? It actually seems harder to make a choice among all good choices.
Accuracy here is making the same choice from a given pack as one of the good players. Obviously subjective so not a perfect metric, but a decent check on ability to emulate a high-quality drafter.
I mean 62% might feel like a good number, but it's arbitrary, you'd have to justify how you chose it, and just eyeballing it, it is filtering out a lot of very good players with many, many more match wins.
Perhaps you can sort by Latest Rank, and filter out people with 2 or fewer trophies. Or you will have to validate with known bad draft choices in the prompt, to see what it does. Suffice it to say, I still don't think the 17Lands data represents what you think it does.
Like without a direct discussion about measuring and accounting for luck in the draft... for all I know the data is seriously flawed. It probably isn't, but it's maybe one of many, many issues to address when dealing with strategy card game AI problems.
Definitely not perfect data though, and agree that defining good in this context is hard -- a lot of the variance of "good" depends on how you play the cards either way. All good points!
Hmm, but there are a lot of players with greater than a 62% lifetime win rate with very few drafts, but there may be many of those players... do you see? The win rate isn't a good filter. You chose it, you are trying to justify it, and I'm not convinced, not without the hard numbers.
I'm not confused about what filter you chose. I just think it's a bad filter, and you haven't thought very deeply about how it affects the data, which includes presumably your test and validation data - however you're choosing to test and validate, apparently by hand, by some eyeballed examples.
Anyway I think you have to compare with a non-LLM, non-random baseline to have any sense if this stuff is working at all. I could be dead wrong. I would maybe compare with a community draft picker.
- Emulation
- Advisor
In case of MTG player emulation for example, I think it makes sense to group data by some rankable criteria like winrate to train rank-specific models that can mimic players of each rank.
I would select from all games played on sufficiently high level.
Edit - I did the math. From the data on the MTG Elo Project, top Magic players have about a 70-75% game win percentage over an average tournament player. They have the top player at ~2300 Elo with the average being around 1500 (in matches), and have scaled the Elo system so that a 200 point gap is a 60% chance to win a best-of-three match (this is NOT the same as Chess Elo scoring).
This is really smart, I didn't think about this! Will add it to my list of things to try, great idea!
> Domain adaptation over subreddits/forums before finetuning may help as well.
I was thinking about this too (along with transcribing draft youtube videos), I'd definitely be curious how much this helps.
Also - why qlora rather than a full finetune? Using LambdaLabs, it'd cost roughly the same as your quote. Cheaper I think if you're willing to gamble with fp8: https://github.com/mosaicml/llm-foundry/tree/main/scripts/tr.... And fewer hyperparameters to tune as well
If so, and if he's correct in his assumption that these sets are out of the bot's training cutoff window, then surely it's purely coincidence if it ends up being a good drafter? The bot would have literally no way to know what cards work well with its previous picks, what signals have been sent and received in the draft so far, etc. Not even the best human player could take (for example, from the sample prompt) "Gadwick's First Duel -- {1}{U} (uncommon)" and figure out what works well with that (if they've never seen the card before).
It would just end up picking generically good draft cards that share a color with its previous picks. Which is already what pick-order-based heuristics have always done.
Not quite -- there's a few ways the model learns the full card text:
* The models are trained on card trivia completions as well, where they're asked to complete the full text of the card as well as information about it (type, CMC, etc.)
* The models do still have to learn next token completion on the cards in packs, meaning they learn to predict the full text of the cards while making draft picks as well.
Net net, the bots learn the text of the new cards pretty comprehensively.
The best performing draft AI's I've seen leverage representation learning in some form.
Still some fun things about LLM representations -- you can do fun things like give the bots preferences / personality in a system prompt which is entertaining!
The resulting model was so much worse than just formatting everything plaintext. This was with MPT-30B, 15 special tokens, 300M training tokens, and a full finetune.
I may have made a mistake, but I haven't seen any open source finetunes successfully add a large number of tokens yet either.
Adding new tokens needs a ton of data to train what the token means. Reusing existing tokens, will allow you to easily teach that a sequence of tokens now has a new meaning after fine tuning.
> Adding new tokens needs a ton of data to train what the token means.
But how much? 300M tokens is fine for a simple version of ChatML with ~4 tokens. Not for 15, at least in my case. How's this relationship scale?
Just trying to offer one datapoint for what doesn't work, with the hedge that I might have just had a bug
a simple input might be <cards you hold> 1 14 56</end><cards to pick> 5 64 2</end> -> predicted token is the draft pick.
Then train a transformer based network from scratch.
I'd be curious about the difference in success w/ drafts on a new 2/2 bear with a different name, and cards with a new keyword 'fizzbangitude 7' as well.
The base checkpoint takes a lot of compute to make, but that's what holds most of it's "knowledge" so to speak.
Making a NN from scratch means you'll have to somehow map the cards into inputs. I have limited knowledge of how MTG works, but most TGG have text descriptions and complex effects. Mapping text to logic is what LLMs are really good at, otherwise you're starting from scratch and will also need a relatively large amount of compute before it starts displaying any type of decent behaviour.
It's also easy for most software devs to do this - finetuning mostly consists of collecting text and feeding it into a finetuning script. You don't need to know linear algebra, what a "convolution" is, etc. to do finetuning.
I was actually just looking into fine-tuning an LLM for Magic: The Gathering this week -- I've been building a small card-similarity browser using semantic embeddings of cards to find functionally or flavorfully similar cards.
I've just been using InstructorXL, but either Instructor doesn't have enough innate knowledge of the game, or else I need to work on better prompts, but so far I've tried 9 different prompts, and none of them seem to perform very well for generating embeddings:
https://github.com/HanClinto/MtgMatrix/blob/main/data/create...
So my next step was to try and download a dataset of similar cards (I have some ideas on this), and I was trying to see if I could use this to do triplet-loss training of a large embedding model or something.
Aaaaand, that's as far as I've gotten. I haven't actually figured out _how_ to hook all of that up, but your post is extremely inspirational for me. Thank you for posting this!!
A question, when GPT-4 contradicts in explanation, how much of them were in fact correct?
It was mostly when a card is good in a vacuum but not as good in a specific set. WOE (which this was trained on) skewed pretty aggressive, so GPT-4 was tended to overvalue strong expensive cards (compared to what good players thought at least).
Sorry if I missed this, but how much did it cost total to do the fine-tune? Is that the 40 hour number (~$27)?
Also, very cool writeup. Thanks for sharing!
I think if you add up all of the learning and testing I did, probably closer to ~$50 total
I was under the assumption that finetuneing LLMs was useful only when you need to change the model's tone (speak like a pirate, voldemort etc).
Are there other examples where LLMs were trained to reason a particular way?
The only issue there is that sometimes the RLHF seeps through, which can be solved by system prompting even harder.
A lot of why I tried this out was to test the limits of this belief, you see a lot of talk like this out there and it sounded like nonsense to me.
Finetuning is fundamentally not much different than continued pretraining; if you feed the model high-quality and high-volume data I think it's reasonable to expect it to acquire new skills
I've seen this effort previously, pretty exciting stuff:
Why would he need to write a game simulator from scratch?
Surely there are OSS versions which are "good enough" ?
(disclosure: not a MtG expert)
Is it managing mana curves, creature count, mana fixing priority etc? I'd also love to know what the common characteristics of "missing" are. Does it undervalue a splash, or removal, or fast/slow.
MTGA and MO development team do a lot of work to put card effects into if-then rule, but unfortunately their work is not visible to the players :)
For learning to draft in one environment and then applying it to a new set of cards, like done here, it would become far more difficult to do without an LLM. The variations and nuances of cards rules text would have to be encoded as features, which would be extremely cumbersome. The LLM gets you some level of understanding of that for free.
2. The model is effectively trained to predict the next token based on the previous tokens in each of these examples, which has the side effect here of teaching it to make a draft pick based on the contents of a pack.
Nothing too fancy, just next word prediction more or less