To use Colab under this arrangement, each person playing the game has to load the model into their personal VM, which is being billed by Google as 6GB of data transfer from GCP into Colab.
The obvious questions;
1) Was there a place they could have hosted the model closer to Colab so that they weren’t being charged for egress bandwidth — and also, ideally, so that 6GB of data wasn’t actually being moved very far?
2) The underlying model is 6GB, but I’m curious how much memory is required for an individual user’s world state and how hard it would be to have a single GPU handling multiple user sessions?
Presumably it would be possible to multiplex multiple sessions with a single GPU? You would have to serialize the game state, receive the next input, load the prior state, feed the new input through the model, return the resulting text output, and re-serialize the state until the next input comes through.
What I don’t know if that’s at all practical based on the amount of data that would have to be serialized? Is the 6GB model data separate and static throughout the game, with an isolated block of data for the current world-state? Or does playing the game fundamentally alter the state of the model, meaning you would have to reload the whole thing just to process the next command?