I do hope we converge on a standardized API and schema for this. Testing and integrating multiple LLMs is tiresome with all the silly little variations in API and prompt formatting.
But it's too bleeding edge, you are asking a lot.
Just do the work and don't be spoiled senseless