At one point we considered adding artifical delay to responses because irrational users dont trust something that finishes fast, even if its the same quality.
How empirical are your comparisons of new and old outputs?
How empirical are your comparisons of new and old outputs?