Measuring What Matters: Construct Validity in Large Language Model Benchmarksarxiv.org1 point·Cynddl··0 commentsOpen articleSaveView on HN