> For instance, in customer service, an RL agent might discover that sometimes asking a clarifying question early in the conversation, even when seemingly obvious, leads to much better resolution rates.
Why does this read to me as the bot finding a path of “Annoy the customer until they hang up and mark the case as solved.” ?