Show HN: Reliably Incorrect – explore LLM reliability with data visualizations
adamsohn.com
adamsohn.com
Ran "what's the best X for Y?" across 6 LLMs (ChatGPT, Perplexity, Gemini, Claude, DeepSeek, Mistral) for ~200 B2B SaaS tools across 34 categories. In 60%+ of categories, models converge on the same "default three" and everything else is effectively invisible. Not wrong — just erased. Single-turn, so it never shows up in p_step^N.
A verification layer catches "false." But there's no layer catching "the space of correct answers was silently pruned." Curious if your framework could be extended from correctness per step to coverage per response.