This surprises me a bit. I've found much better token usage with a proper/efficient MCP as even with a deep /skill defining usage, parsing MCP results is generally just better/more efficient than parsing CLI results. I say this having written a CLI tool explicitly for harness usage, and leveraging MCPs for the same.
I'm sure there are bad MCPs and great CLI tools that parse poorly/well via harness, but I'd be curious on an better research study.