MLCommons provides MLPerf which is a series of benchmarks anyone can run and compare results. For devices like the BeagleY-AI the "edge" category provides a series of various benchmarks to run covering things like object detection, speech recognition, and now smaller LLM's.