Benchmarked output quality versus actual output quality are very different things. Some usecases are at the very fringe of model intelligence and depth of intelligence and logic suffers.
Also interested in how this watermarking push makes sense when considering RSI.