We've been seeing success using it for LLM-as-a-judge tasks.
We set up an evaluation criteria and used o1 to evaluate the quality of the prod model, where the outputs are subjective, like creative writing or explaining code.
It's also useful for developing really good few-shot examples. We'll get o1 to generate multiple examples in different styles, then we'll have humans go through and pick the ones they like best, which we use as few-shot examples for the cheaper, faster prod model.
Finally, for some study I'm doing, I'll use it to grade my assignments before I hand them in. If I get a 7/10 from o1, I'll ask it to suggest the minimal changes I could make to take it to 10/10. Then, I'll make the changes and get it to regrade the paper.