Do you? What's the technical detail here? Why can't you get the model's prediction, even for that first token?
In the matmul, it'd just zero out all parameters. In older models, you'd still have bias vectors but I think recent models don't use those anymore. So the output would be zero probability for each token, if I'm not mistaken.