HNHacker News
TopNewBestAskShowJobs

curious_nile

1 karma · joined July 15, 2025

submissionscomments
curious_nile··on Show HN: We beat Google, Cognition, Claude Code at codebase docs generation
thanks. we map dependencies across the whole codebase first, then generate docs from that. Most tools just read files top to bottom and summarize.

The 7-10x coverage gap isn't really a prompting or model thing, we just don't stop after the important-looking files. CI, build system, test harness, internal utilities, all of it. References (2500+ vs claude code's 50) come from the same place. Once you've mapped the dependencies, every doc can point to every other thing that touches it.

curious_nile··on Show HN: We beat Google, Cognition, Claude Code at codebase docs generation
I'm Nilesh. My brother Abhishek and I built ProdE. Carnegie Mellon and IIT Delhi.

We benchmarked four AI code documentation tools: ProdE, DeepWiki, Claude Code, and Google Code Wiki. ProdE scored highest on usefulness for coding agents. 15% ahead of DeepWiki, 38% ahead of Google, 40% ahead of Claude Code.

I know this might feel like self praise, but we couldn't find an existing benchmark to use, so created one ourselves and open sourced it.

The biggest gap is coverage. Coding agents can only answer questions about parts of the codebase that are documented. If your docs cover routing but skip middleware, every middleware question becomes a hallucination. ProdE documents 114-140 files per project. Claude Code covers 13-17. So agents using Claude Code's docs are blind to roughly 90% of the codebase.

Zero hallucinations across all 9 evaluations. Every file path, function reference, and claim we checked pointed to real code. So it's not just that we cover more, what we cover is also accurate.

DeepWiki did really well here -- 5x more diagrams per project than us, best visual docs by far. Claude Code had the strongest writing quality of the four.

Honestly, if I saw this post I'd also assume the vendor rigged it. So here's everything we did to make it not that. Claude Opus judges all four tools using a published rubric. Claude Code's output was renamed to doc_x/ so the judge couldn't tell it was Claude Code. ProdE launched after Claude's training cutoff, so the judge had no prior knowledge of our tool. We don't use Claude anywhere in our pipeline. 9 evaluation passes across 3 open-source repos (FastAPI, Pydantic, Mermaid), all pinned to exact commits to tackle the non deterministic outputs.

We scored usefulness for coding agents and readability for humans as separate things, because what makes docs good for agents is different from what makes them good for humans. Agents need lots of references to specific files and functions. Humans need clear writing and good diagrams. The tool with the best writing scored lowest on usefulness for agents. Ofcourse the usefulness for Humans is better judged by humans.

Blog (full analysis): https://prode.ai/blogs/we-benchmarked-ai-code-documentation-... Repo (everything, run it yourself): https://github.com/abhishek-curiousboxai/code-documentation-... You can fork it and re-run. Everything is MIT licensed.

curious_nile··on Show HN: ProdE – code change impact and root cause analysis for large codebases
This is awesome, thanks for sharing and building it.

Yeah, that is a challenge we also faced and gave it a good amount of thought. Something like exposing LSP find references capabilities to the agent. Didn't follow this approach though ourselves.

On the documentation front, with ProdE we actually map out the features mentioned in docs to the actual code files, and diff against new commits to see if any file was edited. Edit means docs need to be updated. Working well this way.

Nia is great, but i believe they are indexing the open source world. Are they indexing private repos as well?

curious_nile··on Show HN: ProdE – code change impact and root cause analysis for large codebases
Maker here (Abhishek). ProdE does code change impact + root cause analysis for large codebases.

Backstory: we tried a bunch of “AI-native” dev tooling and kept seeing the same thing, each tool builds its own partial code map inside its own UI. We wanted code intelligence that’s independent, so the whole toolchain can query the same ground truth.

What it does: - Change impact analysis: answer “which services break if I change the auth middleware?” before you ship. - RCA: when tests/prod fail, surface likely causes + related context.

Try: https://prode.ai Requires: signup + connecting your repos (read-only).

Fallback: if you can’t connect a repo but still want to evaluate it, email abhishek@prode.ai and I’ll add you to an OSS-indexed workspace (e.g., Supabase).

Question: should code intelligence live inside the IDE, or as a shared service Slack/Jira/CI/IDEs can all query?

curious_nile··on Show HN: ProdE – Give AI coding tools context for multi-repo codebases
Hi HN, Abhishek here, co-founder of ProdE.ai

Every team has that one engineer who knows the entire codebase, but they're always buried in requests and meetings.

Today, our AI coding tools (Cursor, Windsurf, Copilot) are the ones needing that engineer's time the most. They're powerful, but when faced with complex multi-repo, microservice environments, they fail completely, hallucinate, and generate garbage code because they lack context.

So, we built ProdE. We think of it as the essential knowledge layer for AI-powered development. It connects to your repos and acts as a 'senior engineer brain' for your tools and your team.

How it works is simple:

1. Connect: You link your GitHub/Bitbucket/GitLab repos. 2. Learn: ProdE analyzes the code, structure, and history to build a persistent knowledge layer. 3. Integrate: ProdE seamlessly plugs into your existing AI tools via MCP. There's no new UI to learn; your current tools just start working better.

We do not train any AI models on your code ever, it is completely proprietary and hosted on GCP with servers exclusively located in USA.

To get your feedback, we're making it free for codebases up to 100k lines of code. For larger projects, the first month is on us. No credit card required.

Right now, we are focused on getting the experience right for Python, TypeScript and C#, with support for other languages on the roadmap based on demand.

We just launched and would love this community's honest feedback. What do you like the most? What did we get wrong? What are we missing?

You can check it out here: https://prode.ai/