Full threadanish_m·What are the SOTA benchmarks for LLMs now? Love the progress on opensource models, but would like to see an uncontaminated and objective framework to evaluate them.View on HN