Ayumi's LLM RP Ranking, using llama.cpp grammar for automated quality metrics
ayumi.m8geil.de
ayumi.m8geil.de
More info on the metric itself and older tests is here, though the page is not up-to-date yet with the new testing:
https://rentry.co/ayumi_erp_rating#ayumi-llm-character-iq-al...
This is super cool! I've grown distrustful of standard LLM ranking metrics, even when I'm pretty sure the testers aren't outright cheating, but this seems like a really novel test. What's more, no one is trying to game this RP/ERP testing, and the current ranking aligns with some of my own testing for LLM "smartness," even in models with no focus on roleplaying at all.