Batched reward model inference and Best-of-N samplingraw.sh34 points·rawsh··0 commentsOpen articleSaveView on HN