The goal is to move beyond simple binary classification ("is this hallucinated?") and give users a detailed, interpretable view of where and how an LLM's response may go off the rails.
Key Features:
Token-level confidence visualization (color-coded)
Flagged issues: hallucinations, overconfidence, uncertainty
Terminal rendering (with color)
HTML report generator (interactive)
Markdown output
Stats: confidence range, flagged spans, density
Demo data with realistic hallucination cases
Modular: add custom renderers, integrate into pipelines
Run it via CLI, or integrate as a library into your own tools or dashboards.
GitHub: https://github.com/Mattbusel/LLM-Hallucination-Detection-Scr...
Would love feedback, ideas for integration (e.g. Jupyter, browser extension, VSCode plugin), and thoughts on how to best measure and surface LLM epistemic uncertainty.