Interesting to see only a handful of complex scenarios. I've always suspected ML game agents need hundreds of tiny puzzles with hundreds of variations each to learn game mechanics properly. Like:
The factory is not powered, place the missing power pole(s)
The factory is missing items, place the missing belt(s)
Craft and place these 200 assembly machines
The assembly machine is not running for some reason, fix it
The factory production is too low, double it
Get to this other point in the factory as fast as possible
Fix the brownout
All of the above with and without bots
Programmatically generating a few thousand example scenarios like these should be relatively easy. Then use it like an IQ test question bank: draw a dozen scenarios from the bank and evaluate performance on each based on time & materials used.I hypothesize that ML agents learn faster when evaluated on a sample from a large bank of scenarios of smoothly increasing complexity where more complex scenarios are presented after it scores sufficiently high on lower complexity scenarios.