For some reason I had assumed testing this would be more sophisticated than just checking the thumbs up/down stats and user "vibes"
For some reason I had assumed testing this would be more sophisticated than just checking the thumbs up/down stats and user "vibes"
Retest on benchmarks whether it accomplishes tasks with the same success rates. Prose is only one thing.
Messing with the randomness may make the problem solving capabilities weaker. Probably it doesn't but this is the answer to what else I would want them to do.
Yes - maybe saying "vibes" was minimizing the effort but what I am trying to say is that even the controlled testing is just asking users whether quality is impacted or not. Which is subjective and thats what I meant by when I said "vibes"
Don't get me wrong - I have no idea how one would go about testing this with other methods; I was just stating my assumption.
Since they rolled this out to all users I had assumed there would be other testing involved.
If we did, we probably wouldn’t need LLMs in the first place, i.e. we could just generate text using explicitly programmed algorithms.