HNHacker News
TopNewBestAskShowJobs

luulinh90s

31 karma · joined February 24, 2026

submissionscomments
luulinh90s··on Steering interpretable language models with concept algebra
We haven’t benchmarked our steering for scaffolding function-calling in an agent loop yet (and the model we are using is just a base model), so I can’t give a quantitative claim. But concept-based steering should be a good fit for keeping the agent on task and enforcing behavioral guardrails around tool use.

In practice, you can treat concepts as soft/hard constraints to bias the agent toward: (1) calling tools only when needed, (2) selecting the right tool/function, or (3) using the correct argument schema.

luulinh90s··on Steering interpretable language models with concept algebra
Hi! Thanks for checking.

We haven’t published the concept dictionary yet.

We plan to release it in soon with other important artifacts.

luulinh90s··on Show HN: Steerling-8B, a language model that can explain any token it generates
in the "Performance" section of the post: https://www.guidelabs.ai/post/steerling-8b-base-model-releas..., the authors show the model lags behind llama 8b but worth noting that llama 8b trained on > 2x more computes (see the FLOPs axis)