This sounds quite cool, and I am very picky but this seems quite clean.
It wasn't clear if this was for fine-tuning or not, the general explanation could perhaps be clearer, the first few sentences of the README still don't make it very clear (I can sort of guess, having seen DSPy, AutoPrompt etc).
It would be awesome if this DID also explore if it needed to do: prompt tuning, fine-tuning, or soft-prompt tuning etc, I am on the lookout for a tool that does this. Obviously a general open source Q* like solution would be amazing but I get that might be a bit of a different beast! Part of my issue is that there are so many things that can be tweaked and I often don't know the most time and cost efficient thing to optimise. I get that prompt tuning is often going to be the best thing to do, especially first. But for efficient inference, shorter prompts may well be needed. Though maybe clever model key-value caching is starting to make this less of an issue, but it's still faster to have as short as prompts as possible, still, fine tuning or even resuming pretraining may be the best thing to do sometimes.
BTW I would strip `gpt-3.5-turbo` from all examples, as it's more expensive than the better 4o-mini.
I hope to check this out more later.
Nice work!