HNHacker News
TopNewBestAskShowJobs

giang_at_glai

5 karma · joined February 26, 2026

submissionscomments
giang_at_glai··on Show HN: Steerling-8B, a language model that can explain any token it generates
Actually, the model is forcing the response to be generated inside the attribution modules.
giang_at_glai··on Steering interpretable language models with concept algebra
We will share a technical write-up soon that addresses both of your questions: (1) steering vs. prompt engineering, and (2) how effectively our steering suppresses undesired generations.

If you have joined our waitlist, we will notify you as soon as it is available.

giang_at_glai··on Steering interpretable language models with concept algebra
Author here.

This post shows “concept algebra” on language model: inject, suppress, and compose human-understandable concepts at inference time (no retraining, no prompt engineering).

There’s an interactive demo on the post.

Would love feedback on: (1) what steering tasks you’d benchmark, (2) failure cases you’d want to see, (3) whether this kind of compositional control is useful in real products.

Related: https://news.ycombinator.com/item?id=47131225