I think it's often the other way around. The differences in perceived coding productivity that many people attribute to claude vs codex is more the harness than the model it's running (if running comparable classes of models).
I've been assuming a harness is basically a set of tools and a TUI for passing text to the model and the model coming back with tool calls and user responses.
Are the tools really that complex and different between harnesses?