Is it my understanding that the reward model is also similar to an LLM (with the difference being it predicts a score instead of the next token)?
Everything else in the model before that final layer is exactly identical, architecture-wise.
In the case of a reward model, are you streaming in the list of tokens; if so, what is the output after each token? Or are you feeding in all of the tokens in one shot, with the predicted reward as the output?
You can check the examples from the TRL library for more information.
What library is that? Thanks!