So, the question becomes "how do you guarantee that people will do the task you ask them to do?". Well, simply put: you can't.
Although I see what you mean about unconstrained inputs. If you can give it any number as an input then it's going to be difficult to validate, except at the most superficial level, eg "input 10, generates 10 items matching the expected constraint."
I am skeptical that any real use case wouldn't effectively mean writing the validation code is as much work as just writing a single implementation (and testing it) in the first place.
- At this time Marvin is only tested with GPT-3.5 and GPT-4 to reduce the surface area of LLM differences. You're welcome to try others, but the prompts are optimized for performance of those two models
- I expect that prompts optimized for one family of models will not automatically work with others, so as time allows we may end up with branching prompts based on your model choice. Marvin does a lot of work to manipulate the user prompt before sending it to the LLM so I think there's an excellent chance that we could deliver similar results for the same user input.
- Even with GPT-3.5 and GPT-4 we see a remarkable difference between them, with 4 needing far less complex prompts than 3.5 to get the same result. However, given both the cost and availability of GPT-4, we decided to make 3.5 our default to make sure everyone can use the library. Therefore we do our best to make sure prompts work with 3.5 and expect that that is a good bet of compatibility with 4.