What was the training data?
PS - please don't let this me used at a way to prevent human interaction. Chatbots are a disaster and literally the worst possible application of ML, as a shitty interface to a menu system. I hope this will be used in a way that is not consumer-hostile and that the company actively resists ignorant business attempts to use it to avoid paying for customer support.
And we are not building a chatbot, we're building something collaborative that you can work with to accomplish the stuff you want to do!
And was the feedback data used to train the model with reinforcement learning? Or did you request users to "correct" the action and get a supervised signal?