interesting results. why does reload/cross-tile have worse results? would be nice to see some examples of failed results (how close did it to solving?)
Reload are challenging because of how the agent-action loop works. But the models were pretty good at identifying when a tile contained an item.