MineDojo – Building Open-Ended Embodied Agents with Internet-Scale Knowledge
minedojo.org
minedojo.org
If we were to be extra generous, then the audio and visuals in minecraft could be a toy case of training an agent with a multimodal sensorium (can the agent avoid enemies sneaking up on them using sound cues, find mobs based on sounds, know lava is nearby by the bubbling etc.).
In this case though the agent is given just RGB pixels and some structured data about their inventory, a small number of nearby voxels, their health etc.
You'd have to be very generous to call structured data like that and some RGB pixels a multi-modal embodied agent though.
You're not the target audience, why should they appeal to you? This is a website accompanying a paper about reinforcement learning.