Measuring What Matters: Construct Validity in Large Language Model Benchmarksoxrml.com3 points·Cynddl··2 commentsOpen articleSaveView on HN