And removing the chat interface as much as possible. Many benchmarks are better with text completion models, but they keep insisting on this horrible interface for their models.
Fine tuning is there to ensure you get the output format you want without the extra garbage. I swear they have tuned their models to waste tokens.