Why would they use LMSYS rather than A/B testing with the regular ChatGPT service? Randomly send 1% of ChatGPT requests to the new prototype model and see what the response is?
Maybe they did both? Maybe they have been secretly A/B testing for a while, and only now started to use LMSYS as a secondary data source.
They do both. Source: just got served an A/B test talking to GPT-4 using the web chat interface.
How do you measure the response? Also, it might be underaligned, so it is safer (from the legal point of view) to test it without formally associating it with OpenAI.
GPT regularly gives me A/B responses and asks me which one is better.
I often get this with code, but when I try to select the text to copy the code to test it that is immediately treated as a vote.
This frustrates me, just because I'm copying code doesn't mean it's the better choice. I actuallly want to try both and then vote but instead I accidentially vote each time.
Perhaps to give a metric to include in the announcement when (if) such a new model is released.