This is a good idea, I used gradio and streamlit to list outputs from different models to check the output manually. But using CLI makes more sense for running multiple use cases and evaluate.
You have lots of steps to run, I would suggest:
1. Create a config file (yaml or json) to define prompts, variables, models, and output file.
2. Create an init command which will create empty files with the required structure. For example:
`promptfoo init`
output will be:
config.yaml var.json prompts.json
Good luck!