ParentFull threadwongarsu·Tbf, most of the "real benchmarks" have issues that are just as bad. Assessing LLM performance is just hardView on HN