HNHacker News
TopNewBestAskShowJobs

Jet_Xu

54 karma · joined June 4, 2024

Systemizing Vibe Coding through AI Harnessing Engineering. Creator of DocMason an open-source repo-native agent app for analyst-grade answers over complex private files. The repo is the app. Codex is the runtime.

https://github.com/JetXu-LLM/DocMason

submissionscomments
Jet_Xu··on Ask HN: LeetCode, anyone still doing it?
It depends on whether nowadays FLAG will still need engineer write code during interview.
Jet_Xu··on AI still can't figure out PowerPoint
You are right but it is not because the weakness of AI, it is because all AI generation app are built by those coder only know .md files. They do not know what a consultant level slides looks like.

See this demo in the middle I show some top level strategy business slides (they are all generated by AI) https://youtu.be/Sq3a5qxsLwM

I am also working on such tool :)

Jet_Xu··on AI Can't Read an Investor Deck
Try DocMason it is born for agentic analyze messy office files.

https://github.com/JetXu-LLM/DocMason

Demo video: https://youtu.be/Sq3a5qxsLwM

Jet_Xu··on Is RAG Dead? Long Context, Grep, and the End of the Mandatory Vector DB
If you want to get knowledge from KB within 10 seconds, then traditional RAG (embedding+vector based) is still the most efficient way -- maybe more in consumer facing search / chat area.

If you want to get most precise and useful information from KB in minutes, then Agentic RAG is the key.

Jet_Xu··on LLM Wiki – example of an "idea file"
I believe Multimodal KB+Agentic RAG is a suitable solution for personal KB. Imagine you have tons of office docs and want to dig some complex topics within it. You could try https://github.com/JetXu-LLM/DocMason

Fully retrieve all diagram or charts info from ppt and excels, and then leverage Native AI agents(e.g. Codex) to conduct Agentic Rad

Jet_Xu··on Show HN: DocMason – AI Agent Knowledge Base for local complex office files
I pressed Hide, but still not work. And I do not have a "delete" button under this post. It is really awkward...
Jet_Xu··on [dead]
I am so sorry that I post two duplicate post in Show HN for same project

please just refer to this post with main content in the body (pure hand write content no AI generation ^-^): https://news.ycombinator.com/item?id=47640770

And could anyone help or tell me how to delete this post?

Jet_Xu··on Ask HN: Do you code remotely from your phone?
No, too harmful for my eyes...
Jet_Xu··on Why is Hacker News such an oldschool page?
What I want is markdown format support in Hacker news. But other features really not important :)
Jet_Xu··on DeepSeek v3 beats Claude sonnet 3.5 and way cheaper
Please refer to my recent AI Code review performance test include DeepSeek V3: https://news.ycombinator.com/item?id=42547196
Jet_Xu··on What do you check first in PR review? Help shape our AI Code review tool
A few weeks ago, I shared LlamaPReview (an AI PR reviewer) here on HN and received great feedback [1]. Now I'm trying to understand how experienced developers prioritize different aspects of code review to make the tool more effective.

When you open a PR, what's the first thing you check? Is it:

- Overview & Architecture Changes - Detailed Technical Analysis - Critical Findings & Issues - Security Concerns - Testing Coverage - Documentation - Deployment Impact

I've set up a quick poll here: https://github.com/JetXu-LLM/LlamaPReview-site/discussions/9

Current results show an interesting split between "Detailed Technical Analysis" and "Critical Findings", but I'd love to hear HN's perspective:

1. What makes you trust/distrust a PR at first glance? 2. How do you balance between architectural concerns and implementation details? 3. What information do you wish was always prominently displayed?

Your insights will directly influence how we structure AI Code Review to match real developers' thought processes.

[1] Previous discussion: https://news.ycombinator.com/item?id=41996859

Jet_Xu··on Nvidia Steps Up Hiring in China to Focus on AI-Driven Cars
Interesting timing, considering China just launched an antitrust probe into Nvidia [1] and there were earlier reports about H20 order suspension [2]. This hiring push seems to be a delicate balancing act:

- Expanding autonomous driving R&D while navigating export controls - Maintaining Chinese market presence despite regulatory pressures - Building local expertise when H100/A100 sales are restricted

The automotive sector might be strategically chosen as it's less impacted by current chip restrictions. Worth noting that China remains Nvidia's largest market in Asia, accounting for ~20% of their revenue.

[1] https://www.reuters.com/technology/china-investigates-nvidia... [2] https://www.tomshardware.com/tech-industry/artificial-intell...

Jet_Xu··on Show HN: GitBook Documentation Downloader for LLMs
Interesting approach! While converting docs to markdown works, I've found that technical documentation (especially for complex repos) needs more than just text conversion for effective LLM consumption.

The key challenge is preserving repository context - like code dependencies, architectural decisions, and evolution patterns. Have others experimented with knowledge graph approaches for maintaining these relationships when processing repos for LLMs?

Jet_Xu··on Show HN: Replace "hub" by "ingest" in GitHub URLs for a prompt-friendly extract
Interesting approach! While URL-based extraction is convenient, I've been working on a more comprehensive solution for repository knowledge retrieval (llama-github). The key challenge isn't just extracting code, but understanding the semantic relationships and evolution patterns within repositories.

A few observations from building large-scale repo analysis systems:

1. Simple text extraction often misses critical context about code dependencies and architectural decisions 2. Repository structure varies significantly across languages and frameworks - what works for Python might fail for complex C++ projects 3. Caching strategies become crucial when dealing with enterprise-scale monorepos

The real challenge is building a universal knowledge graph that captures both explicit (code, dependencies) and implicit (architectural patterns, evolution history) relationships. We've found that combining static analysis with selective LLM augmentation provides better context than pure extraction approaches.

Curious about others' experiences with handling cross-repository knowledge transfer, especially in polyrepo environments?

Jet_Xu··on Show HN: Open-Source Colab Notebooks to Implement Advanced RAG Techniques
Interesting discussion! While RAG is powerful for document retrieval, applying it to code repositories presents unique challenges that go beyond traditional RAG implementations. I've been working on a universal repository knowledge graph system, and found that the real complexity lies in handling cross-language semantic understanding and maintaining relationship context across different repo structures (mono/poly).

Has anyone successfully implemented a language-agnostic approach that can: 1. Capture implicit code relationships without heavy LLM dependency? 2. Scale efficiently for large monorepos while preserving fine-grained semantic links? 3. Handle cross-module dependencies and version evolution?

Current solutions like AST-based analysis + traditional embeddings seem to miss crucial semantic contexts. Curious about others' experiences with hybrid approaches combining static analysis and lightweight ML models.

Jet_Xu··on Lessons Learned: Migrating to Mistral-Large-2411 for Production Code Reviews
I'd like to share our technical journey migrating our code review system from Mistral-Large-2407 to 2411, and the key challenges we overcame. Here are the most interesting findings:

1. Prompt Pattern Evolution

  - Initial challenge: Direct model upgrade led to significant quality degradation
  - Root cause: Changes in 2411's prompt processing architecture
    # Previous prompt format for Mistral-Large-2407
    <s>[INST] user message[/INST] assistant message</s>[INST] system prompt + "\n\n" + user message[/INST]
    # New optimized prompt format for Mistral-Large-2411
    <s>[SYSTEM_PROMPT] system prompt[/SYSTEM PROMPT][INST] user message[/INST] assistant message</s>[INST] user message[/INST]
  - Solution: Implemented enhanced prompt patterns through LangChain
2. API Integration Insights

  - Built custom HTTP client interceptor for debugging
  - Discovered crucial differences in message formatting
  - Leveraged LangChain's abstraction layer effectively
3. Key Technical Improvements

  - Enhanced review focus through optimized prompts
  - Improved output reliability and format compliance
  - Eliminated response truncation issues

This is implemented in our AI Code Review Github APP LlamaPReview [https://jetxu-llm.github.io/LlamaPReview-site/]. Happy to discuss specific implementation details or share more technical insights about working with Mistral-Large-2411 in production.
Jet_Xu··on Show HN: Email Inbox for Bots
Is that possible this bot could understand and remember all my history emails, and then conduct Agentic RAG before helping on reply?
Jet_Xu··on AI Agents: A $4.6T opportunity to transform the software industry
I've been working on an AI-driven development vision, building on the "System of Agents" concept discussed in the article.

The key insight is that to achieve truly automated development, we need AI agents across the entire development lifecycle - not just in isolated tasks.

Our system implements a "Retrieval Augmented - Complex Coding Task accomplishment loop" with: - Multiple LLM-based agents handling different aspects of development - A Task and State Management System for agent coordination - Dual knowledge retrieval combining self-repository learning and GitHub public knowledge - Integrated development tools library (editing, testing, deployment) orchestrated through multi-agent collaboration

The key differentiator is the holistic integration - every step from initial development to deployment is agent-aware and interconnected. This creates a true "Service-as-Software" development environment where AI doesn't just assist, but actively drives the development process.

We're seeing promising results in automated code generation, self-healing test suites, and intelligent CI/CD pipelines that learn from deployment patterns.

Would love to hear thoughts from the community, especially from those working on similar full-lifecycle automation approaches.

Jet_Xu··on Show HN: LlamaPReview – AI GitHub PR reviewer that learns your codebase
Thank you for such detailed and thoughtful feedback! Really appreciate the time you took to analyze our claims and point out the areas needing more clarity.

You're absolutely right about the marketing copy - we should be more precise and transparent about what we actually do vs. what's aspirational.

Regarding "understanding code relationships and dependencies": We're building a knowledge graph of the entire repository that captures code relationships, function calls, and module dependencies. This graph is then used with GraphRAG to fetch relevant context for each PR, allowing the LLM to understand the broader impact of changes.

Important to note: We take privacy very seriously. All code analysis happens in-memory during PR reviews - we don't permanently store any source code or build persistent knowledge bases from customer code. The knowledge graph is generated and used on-the-fly for each review session.

This approach helps us work around context window limitations while providing meaningful insights. However, I should note that this feature is still under active development - we're continuously improving the graph construction and relevancy matching.

Would love to hear your thoughts on this approach. We're committed to building something genuinely useful for developers rather than just another LLM wrapper.

Jet_Xu··on Show HN: LlamaPReview – AI GitHub PR reviewer that learns your codebase
Thanks for these important questions!

The service runs on secure cloud infrastructure and processes code in-memory during PR reviews - we don't permanently store any source code. We use enterprise-grade LLMs (can't disclose specific models due to licensing) and implement context-aware analysis without fine-tuning on customer code.

When we say "learning", we mean analyzing the codebase context during PR reviews to understand patterns and relationships, not training or building persistent knowledge bases. This ensures both privacy and effectiveness.

We're working on open-sourcing parts of the implementation - will share more soon!

Jet_Xu··on Show HN: LlamaPReview – AI GitHub PR reviewer that learns your codebase
The core code of this LlamaPReview is from my open source project llama-github. But there are some other code currently still have not open source.
Jet_Xu··on Show HN: LlamaPReview – AI GitHub PR reviewer that learns your codebase
Maybe you could refer to my open source project llama-github
Jet_Xu··on Show HN: LlamaPReview – AI GitHub PR reviewer that learns your codebase
Good questions! Code review and generation are quite different tasks. Review is about pattern recognition and consistency checking, while generation requires understanding business logic and system design.

LlamaPReview works best at: - Spotting potential issues (like off-by-one errors) - Identifying patterns across the codebase - Maintaining coding standards

For complex architectural decisions, it serves as an assistant rather than a replacement - helping senior developers save their time to focus their attention where it matters most.

Jet_Xu··on Show HN: LlamaPReview – AI GitHub PR reviewer that learns your codebase
Thanks for raising this important question. We will not store any code in our database. But we will leverage SaaS LLM API (e.g. GPT/Claude/Mistral) to help on the PR review - during this step, for sure we need to send code to these SaaS LLM for analyze. This is the main reason why we mentioned "collecting users code" in our privacy.
Jet_Xu··on Show HN: LlamaPReview – AI GitHub PR reviewer that learns your codebase
Interesting perspective on the timing of feedback. We chose PR reviews because they're a natural integration point where developers already expect feedback, and it's when context is most complete. However, we're exploring ways to provide earlier feedback without being intrusive.

The key is finding the right balance between immediate assistance and allowing developers to maintain their flow. Would love to hear more about your experiences with different feedback timing approaches.

Jet_Xu··on Show HN: LlamaPReview – AI GitHub PR reviewer that learns your codebase
Thanks for mentioning PR Agent. While there are several tools in this space, LlamaPReview focuses on deep codebase understanding and context-aware reviews(advanced functions still under evolution). We'd love to hear about your experiences and what specific features you find most valuable in code review tools.
Jet_Xu··on Show HN: LlamaPReview – AI GitHub PR reviewer that learns your codebase
Thanks for raising this question. Currently, we're offering a free tier to gather community feedback and improve the service. We use enterprise-grade LLMs to ensure high-quality reviews while maintaining reasonable operational costs. Our focus is on building a valuable tool for developers first, and we'll be transparent about any future pricing changes.
Jet_Xu··on Show HN: LlamaPReview – AI GitHub PR reviewer that learns your codebase
The "learning" process involves analyzing your codebase's context during PR reviews - we don't train on your data (we even will not save them but only calculate in memory). Instead, we use advanced context retrieval to understand:

- Project structure and architecture - Coding patterns and conventions - Dependencies and relationships between components

This allows us to provide more relevant and context-aware reviews while maintaining data privacy (some advanced features still is under developing)

Jet_Xu··on Show HN: LlamaPReview – AI GitHub PR reviewer that learns your codebase
Great points about code review reliability. LlamaPReview is designed to be a complementary tool for senior developers, not a replacement for human review. Here's our approach:

1. It helps save senior developers' time by handling routine checks and providing initial insights 2. It analyzes the entire codebase context to provide more meaningful reviews 3. It's particularly useful for identifying patterns and relationships across the codebase

The goal is to make human reviewers more efficient, allowing them to focus on complex architectural decisions and critical business logic. We've seen positive results from both open-source and commercial projects using this approach.

Jet_Xu··on Show HN: LlamaPReview – AI GitHub PR reviewer that learns your codebase
Thanks for raising these important questions about data privacy and security. Let me clarify:

1. Code Processing: All code analysis happens in-memory during the PR review process. We don't permanently store any of your source code.

2. Data Retention: We only store the PR comments we generate, not the underlying code. This helps maintain a history of our suggestions while protecting your IP.

3. Privacy Focus: We take data privacy seriously and have successfully worked with both open-source and closed-source projects. We're always open to suggestions on how to further enhance our privacy measures.

If you have specific privacy requirements or suggestions, I'd be happy to discuss them.

Page 1 of 2Next →