Braid: Bounded reasoning for LLMs using symbolic Mermaid graphs
arxiv.org
arxiv.org
We evaluate BRAID across GSM-Hard, SCALE MultiChallenge, and AdvancedIF.
Key findings:
- Structured symbolic reasoning improves accuracy on complex tasks
- Smaller models often match or outperform larger models using classic prompting
- Significant cost reductions (up to 74× performance-per-dollar)
- Even SOTA models see accuracy gains when pure performance is the goal
All benchmarks and detailed logs are public: https://benchmark.openserv.ai
Happy to discuss methodology, evaluation choices, limitations, or failure cases.