HNHacker News
TopNewBestAskShowJobs

rohanucla

113 karma · joined May 21, 2026

submissionscomments
rohanucla··on Cleaning up after AI rockstar developers
Nice Blog!
rohanucla··on Sem: New primitive for code understanding – not LSPs, but entities on top of Git
I mean we are trying to be faster than LSPs, LSPs are a little slow for enterprise grade codebases
rohanucla··on Sem: New primitive for code understanding – not LSPs, but entities on top of Git
There is a skill.md for the agent to know about the cli, I can make update the same with more examples.
rohanucla··on Sem: New primitive for code understanding – not LSPs, but entities on top of Git
Thanks a lot Alex! for this reply, it keeps us pumped.
rohanucla··on Sem: New primitive for code understanding – not LSPs, but entities on top of Git
no setup just configures your git diff to use sem by defult, you will find the sem mcp directory on github repostiory, also there's skill.md file which will tell your agent on how to use sem.
rohanucla··on Sem: New primitive for code understanding – not LSPs, but entities on top of Git
sorry if you consider that as hijack, it was just a user's request to use this as default plugin on their git. But I will add it to let the users know thanks for the feedback
rohanucla··on Sem: New primitive for code understanding – not LSPs, but entities on top of Git
It doesn't override git diff at all, sem is its own standalone CLI. git diff continues to work exactly as before. You do sem setup only when you want to change your default git diff behavior, other wise after installing sem you can use it straight away using sem commands.
rohanucla··on Sem: New primitive for code understanding – not LSPs, but entities on top of Git
sem doesn't override git diff, it's a completely separate command (sem diff). Your regular git diff should work exactly as it always has after installing sem.

If you want to change your git diff default behavior then you can do sem setup.

rohanucla··on Sem: New primitive for code understanding – not LSPs, but entities on top of Git
This is actually the exact scenario we just spent the last few weeks optimizing for. On a 71K-file TypeScript monorepo, sem was previously choking entirely (DNF), and now completes in 6.5s with the topology cache warm. On a 100K-file generated fixture, sem impact went from 90s cold down to about 1s warm. The key was building a SQLite-backed cache that stores the dependency graph structure so repeat runs skip re-parsing unchanged files entirely.
rohanucla··on Sem: New primitive for code understanding – not LSPs, but entities on top of Git
Appreciate it!
rohanucla··on Sem: New primitive for code understanding – not LSPs, but entities on top of Git
That's a really compelling use case actually
rohanucla··on Sem: New primitive for code understanding – not LSPs, but entities on top of Git
Thanks! The data artifacts angle is really interesting. in some ways the problem is even harder there because data pipelines have less explicit structure than code, I guess.
rohanucla··on Sem: New primitive for code understanding – not LSPs, but entities on top of Git
haha definitely!
rohanucla··on Sem: New primitive for code understanding – not LSPs, but entities on top of Git
git is actually great, and there are not much of the issues as the world says about it, and the best is to build complimentary layers that makes it even stronger is the best bet I guess.
rohanucla··on Sem: New primitive for code understanding – not LSPs, but entities on top of Git
I am sorry, should have put up a warning there, but You can do sem unsetup, if you go to the github, you will understand more about the way to reverse it.
rohanucla··on Sem: New primitive for code understanding – not LSPs, but entities on top of Git
Lemme give you an example. when you're working in a 100K-file TypeScript monorepo and you change a utility function that parses API responses. git diff tells you that you changed n lines in that function. What it doesn't tell you is which services, components, and tests actually depend on that function across the repo. You're left grepping for the function name, hoping nobody aliased the import or re-exported it through a barrel file. sem impact gives you that full downstream dependency list in seconds, so you know exactly what to review and test before you ship.
rohanucla··on Sem: New primitive for code understanding – not LSPs, but entities on top of Git
Ha, the regex approach is honestly how a lot of people start with this problem and you can get surprisingly far with it until you hit the edge cases around aliased imports, re-exports, and nested scopes where things start falling apart. That's basically why we went with tree-sitter under the hood it gives you the actual parse tree so you don't have to keep patching regex patterns for every new language construct.
rohanucla··on Sem: New primitive for code understanding – not LSPs, but entities on top of Git
This is a really interesting direction, you're essentially talking about data flow or taint analysis, where you track how a value propagates through copies and transformations rather than just following call edges. Honestly pure static analysis gets you partway there but it hits real limits once you run into dynamic dispatch, runtime branching, or serialization boundaries where data gets written somewhere and read back in a completely different part of the codebase.

We're on the structural side right now with call graphs and dependency edges, but a hybrid approach that combines the static graph with runtime instrumentation to fill in the gaps is definitely something I'd love to explore. Thanks for the feedback.

rohanucla··on Sem: New primitive for code understanding – not LSPs, but entities on top of Git
Thanks for pointing it out. I agree with you here, my testing process was quite specific to sem's output but also would love any suggestion from you of how you would design the whole testing process for this kind of tool?

I can also give my thought process, because I was more interested in figuring out the model's inherent search results and understanding without sem.

rohanucla··on Sem: New primitive for code understanding – not LSPs, but entities on top of Git
Really glad you've been using it, and yeah that's exactly the direction I've been thinking about. The line diff as the default view in code forges has always felt like an accident of history it definitely was easy to compute, but not what's actually useful for understanding what changed.
rohanucla··on Sem: New primitive for code understanding – not LSPs, but entities on top of Git
What I've been more interested in lately is structural intelligence as a field in whole.

Things with LLMs break because our infra was always designed for analyzing lines(tools like grep fuzzy matching) and working on quite small sections of code. LLMs struggle with this in cases when they have to analyze different parts of a codebase they either get too much context where you're throwing whole files at them, or too little where they only see the function in isolation, with no real understanding of how the pieces actually connect to each other.

That's really the gap sem is trying to fill. With sem impact you can give an agent the precise blast radius of a change instead of guessing which files matter, and sem diff --patch lets you enforce that a change only touches specific functions and reject anything that bleeds outside that boundary something that's really hard to do with line-level diffs.

Your testing idea is actually closer than you might think. sem already extracts entity signatures, dependencies, and call graphs, so you could build a harness that gives the test-writing agent only the function signature with its dependency graph and behavioral contract, while withholding the implementation entirely. That would force the agent toward behavioral tests because it literally can't see the internals to mock them. I haven't built this harness myself yet but sem graph and sem inspect expose everything you'd need.

The general principle is that sem gives you a structural map of the codebase to both constrain and validate what the model produces, rather than treating code as flat text and hoping the model figures out the relationships on its own.

Another usecase can be about figuring out dead code present in the codebase.

Edit: Also one last thing because I started working on this while solving the fundamental issue of why merge conflicts were occuring with git, so you might also like the merge drive I open sourced on the same Github org - Weave

rohanucla··on Sem: New primitive for code understanding – not LSPs, but entities on top of Git
It can do that, but that's a small slice of what it does. sem parses your codebase into entities (functions, classes, methods) and builds a dependency graph across files.

So instead of line level analysis the whole granularity of seeing changes and tracking thing shifts to entities. It helps in attention mapping of your agent and lets you track the changes faster.

LSPs have been doing it for quite long but using treesitters is faster even tho type awareness is not great with this approach but overall working across multiple languages with a single tool can be quite helpful.

rohanucla··on Show HN: Anyone interested in a tool helps to explore C++ ASTs
Cool project. The bidirectional source-to-AST navigation is the killer combination, clang -Xclang -ast-dump gives you a wall of text that's hard to correlate back to code, especially with C++'s implicit conversions and template instantiations. Being able to click and see exactly what Clang produces is invaluable for anyone writing AST matchers or clang-tidy checks.

I work on a related project called sem (Ataraxy-Labs/sem) that takes ASTs in a different direction — extracting semantic entities from tree-sitter ASTs to build cross-file dependency graphs with git history tracking. Would be curious if you've thought about leveraging Clang's richer AST to do cross-file entity analysis for C++.

rohanucla··on Lazydiff: Terminal PR Review with AST-Aware Semantic Diff Rendering
Do lemme know about feedback and i am open to any kind of constructive criticism, it really helps me improve.