136 karma · joined November 25, 2025
My current project is Steer, an open-source library for deterministic AI agent reliability.
GitHub: https://github.com/imtt-dev/steer
I write about the 'Confident Idiot' problem and the shift from passive monitoring to active teaching.
Blog: https://steer-labs.com
Green Line (Capability): MMLU State-of-the-Art (48% → 94%).
Red Line (Control): Organizations with effective mitigation for inaccuracy/hallucination (lagging at ~52%).
Sources:
MMLU: Official Repository (hendrycks/test) and Model Technical Reports.
Risk: McKinsey “State of AI” Annual Reports (2023–2025).
The Gap: Intelligence is surging. Control is lagging. We are currently in the gap (High capability, low trust).
How are you solving the gap?
Potential Solution: Move verification outside the model using deterministic “Reality Locks” (Regex, SQL AST, Entropy)
Prompting ("Be concise") is brittle. I switched to Shannon Entropy.
The Hypothesis: Code/Data is mathematically "messy" (High Entropy). Slop is smooth and predictable (Low Entropy).
I wrote a filter that blocks responses if entropy dips below ~3.5.
The Payoff: It captures the blocked slop as a dataset for DPO. I use the math to gather data today so I can fine-tune a natively quiet model tomorrow.
I spent the break reading the recent Stanford/Harvard paper on agentic adaptation [1]. Their research provides mathematical proof for what I experienced in Q4: supervising only final outputs is a dead end. Agents learn to "ignore tools and improve likelihood," meaning they learn to lie more convincingly to pass evaluations while the underlying logic rots.
I call this the Agent Lobotomy.
The agent I have in production today is significantly dumber than the one I demoed in December. I was forced to strip autonomy, remove context, and add human checkpoints because I could not trust the probabilistic output. We are stuck in an Autonomy Retreat, creating an Authority Bottleneck [2] where agents are relegated to assistive tasks because the tail risk of autonomous action is too high.
I built Steer (open source) to stop the bleed. In v0.4.0, I moved the architecture to an Agent Service Mesh pattern. Instead of decorating every function, you patch the framework (e.g. PydanticAI) at the entry point. It auto-discovers tools and enforces a reliability policy globally via deterministic Reality Locks.
The real unlock is the data. By capturing the delta between a Blocked Response and a Taught Fix, Steer acts as a synthetic data factory for DPO. It moves reliability from a runtime tax to a training asset, allowing you to eventually refactor your prompt monolith into fine-tuned model weights.
I've put together three cookbooks showing how this stops the lobotomy in SQL and RAG workflows: 1/ Framework Patching: https://github.com/imtt-dev/steer/blob/main/steer/cookbook/p... 2/ SQL Security Lock: https://github.com/imtt-dev/steer/blob/main/steer/cookbook/s... 3/ RAG Grounding Guard: https://github.com/imtt-dev/steer/blob/main/steer/cookbook/r...
References: [1] https://arxiv.org/abs/2512.16301 [2] https://cloudedjudgement.substack.com/p/clouded-judgement-12...
I’ve realized that a 3,000-token system prompt isn't "logic", it's legacy code that no one wants to touch. It’s brittle, hard to test, and expensive to run. It is Technical Debt.
My thesis is that we need to stop treating prompts as the "program" and start treating them as temporary specs that eventually get compiled into the model weights via fine-tuning.
I built Steer (open source) to automate this "refactoring" process. It helps you climb the "Deliberation Ladder":
1. The Floor (Validity): Use Steer's deterministic verifiers (regex, AST, JSON Schema) to block objective failures in real-time. Don't ask an LLM if JSON is valid; check it with code.
2. The Ceiling (Quality): Use `steer export` to turn those captured failures into a fine-tuning dataset, training the model to handle nuance and "vibes" without a massive prompt.
Curious if others are seeing this "Prompt Bloat" in production?
That thread [1] blew up, so I’m sharing the open-source implementation (v0.2) that solves it.
Steer is an active reliability layer for Python agents. It sits between your LLM and the user to enforce hard constraints.
Unlike passive observability tools that just log errors, Steer creates a feedback loop:
1. Catch: It uses deterministic verifiers (like Regex, AST parsing, JSON Schema) to block hallucinations in real-time.
2. Teach: You fix the behavior in a local dashboard (`steer ui`).
3. Train: v0.2 adds a "Data Engine" that exports these runtime failures into an OpenAI-ready fine-tuning dataset.
The goal isn't just to block errors; it's to use those errors to bootstrap a model that stops making them.
It is Python-native, local-first, and framework agnostic.
My thesis is that even in those "fuzzy" workflows, the agent's process is full of small, deterministic sub-tasks that can and should be verified.
For example, before the AI even attempts to analyze the X-ray for cancer, it must: 1/ Verify it has the correct patient file (PatientIDVerifier). 2/ Verify the image is a chest X-ray and not a brain MRI (ModalityVerifier). 3/ Verify the date of the scan is within the relevant timeframe (DateVerifier).
These are "boring," deterministic checks. But a failure on any one of them makes the final "judgment" output completely useless.
steer isn't designed to automate the final, high-stakes judgment. It's designed to automate the pre-flight checklist, ensuring the agent has the correct, factually grounded information before it even begins the complex reasoning task. It's about reducing the "unforced errors" so the human expert can focus only on the truly hard part.
My thesis isn't that we can stop the hallucinating (non-determinism), but that we can bound it.
If we wrap the generation in hard assertions (e.g., assert response.price > 0), we turn 'probability' into 'manageable software engineering.' The generation remains probabilistic, but the acceptance criteria becomes binary and deterministic.
I built an open-source library to enforce these logic/safety rules outside the model loop: https://github.com/imtt-dev/steer