3,064 karma · joined February 23, 2008
Cerebras has used optillm for optimising inference with techniques like CePO and LongCePO.
Unlike most hallucination detection approaches that require separate LLM calls (which add cost and latency), this is a lightweight classifier built on HuggingFace transformers. It's adaptive, meaning it continuously improves as it processes more examples.
Technical approach:
- Uses a prototype memory system that maintains class examples for quick adaptation
- Combines transformer embeddings with an adaptive neural layer
- Trained on the RAGTruth benchmark dataset across QA, summarization, and data-to-text tasks
- Achieves 80.7% recall overall (51.5% F1), with strongest performance on data-to-text generation
Example usage:
from adaptive_classifier import AdaptiveClassifier
# Load pre-trained detector
detector = AdaptiveClassifier.from_pretrained("adaptive-classifier/llm-hallucination-detector")
# Format input with context, query and response
input_text = f"Context: {your_context}\nQuestion: {your_question}\nAnswer: {llm_response}"
# Get prediction
prediction = detector.predict(input_text)
# Returns: [('HALLUCINATED', 0.72), ('NOT_HALLUCINATED', 0.28)]
Current limitations:
- Performance varies by task type (stronger on data-to-text, weaker on summarization precision)
- Initial version focuses on binary classification; token-level detection is planned
- The model is relatively small, so it won't catch subtle nuanced hallucinations that require deep domain knowledge
The library's wider goal is to enable adaptive classification for use cases where models need to continuously learn from new examples. We've also built LLM routers and configuration optimizers with it.
Would love feedback from anyone working on RAG systems or LLM evaluation. What metrics or capabilities would be most useful to you in a hallucination detector?
Project: https://github.com/codelion/adaptive-classifier
Docs: https://github.com/codelion/adaptive-classifier#hallucinatio...
i've also been experimenting with different chunking strategies to see if that helps maintain coherence over larger contexts. it's a tricky problem.
it's interesting that journals are starting to require raw data... that shift could definitely drive adoption of these kinds of systems. btw, i've seen some projects focusing on verifiable data provenance, might be relevant here.
Have you tried explicitly framing the prompt to reward identifying risks and downsides? For example, instead of asking "Is this a good investment?", try "What are the top 3 reasons this company is likely to fail?". You might get more critical output by shifting the focus.
Another thought - maybe try adjusting the temperature or top_p sampling parameters. Lowering these values might make the model more decisive and less likely to generate optimistic scenarios.
The other benefit of a DSL like Semgrep's is that LLMs have become very good at generating it. See https://github.com/lambdasec/autogrep on how to automatically generate Semgrep rules from existing CVEs.
On the other hand, shipping a whole browser feels so wasteful, especially for smaller apps. I've been experimenting with Tauri recently, which uses the system's webview. It's a nice middle ground, and the performance is noticeably better than Electron from what I've seen.