OpenAI reasoning models: Advice on prompting
simonwillison.net
simonwillison.net
This databricks blog/paper seemingly contradicts this:
> OpenAI o1 models show a consistent improvement over Anthropic and Google models on our long context RAG Benchmark up to 128k tokens.
https://www.databricks.com/blog/long-context-rag-capabilitie...