1. Use some sort of instruction tuning to get the thing "good enough" that it gives decent results 75% of the time and the other 25% a human has to take over. 2. Use the actual usage data as training input. Punish bad behaviors and show the model what the human did to solve the problem. 3. Use this training loop to progressively have the model take over a larger % of the time.
…and I think if you can't get (1) good enough to be worth using it's going to be really hard to get the loop going.