We used a combination of human demonstrations and feedback data!
You need custom software both to record the demonstrations and to represent the state of the Tool in a model-consumable way.
And was the feedback data used to train the model with reinforcement learning? Or did you request users to "correct" the action and get a supervised signal?