Full threadrtuin·Any benchmarks other than computer use/agentic coding published yet? Curious to compare more broadly with other modelsView on HN