54 karma · joined June 4, 2024
https://github.com/JetXu-LLM/DocMason
See this demo in the middle I show some top level strategy business slides (they are all generated by AI) https://youtu.be/Sq3a5qxsLwM
I am also working on such tool :)
https://github.com/JetXu-LLM/DocMason
Demo video: https://youtu.be/Sq3a5qxsLwM
If you want to get most precise and useful information from KB in minutes, then Agentic RAG is the key.
Fully retrieve all diagram or charts info from ppt and excels, and then leverage Native AI agents(e.g. Codex) to conduct Agentic Rad
please just refer to this post with main content in the body (pure hand write content no AI generation ^-^): https://news.ycombinator.com/item?id=47640770
And could anyone help or tell me how to delete this post?
When you open a PR, what's the first thing you check? Is it:
- Overview & Architecture Changes - Detailed Technical Analysis - Critical Findings & Issues - Security Concerns - Testing Coverage - Documentation - Deployment Impact
I've set up a quick poll here: https://github.com/JetXu-LLM/LlamaPReview-site/discussions/9
Current results show an interesting split between "Detailed Technical Analysis" and "Critical Findings", but I'd love to hear HN's perspective:
1. What makes you trust/distrust a PR at first glance? 2. How do you balance between architectural concerns and implementation details? 3. What information do you wish was always prominently displayed?
Your insights will directly influence how we structure AI Code Review to match real developers' thought processes.
[1] Previous discussion: https://news.ycombinator.com/item?id=41996859
- Expanding autonomous driving R&D while navigating export controls - Maintaining Chinese market presence despite regulatory pressures - Building local expertise when H100/A100 sales are restricted
The automotive sector might be strategically chosen as it's less impacted by current chip restrictions. Worth noting that China remains Nvidia's largest market in Asia, accounting for ~20% of their revenue.
[1] https://www.reuters.com/technology/china-investigates-nvidia... [2] https://www.tomshardware.com/tech-industry/artificial-intell...
The key challenge is preserving repository context - like code dependencies, architectural decisions, and evolution patterns. Have others experimented with knowledge graph approaches for maintaining these relationships when processing repos for LLMs?
A few observations from building large-scale repo analysis systems:
1. Simple text extraction often misses critical context about code dependencies and architectural decisions 2. Repository structure varies significantly across languages and frameworks - what works for Python might fail for complex C++ projects 3. Caching strategies become crucial when dealing with enterprise-scale monorepos
The real challenge is building a universal knowledge graph that captures both explicit (code, dependencies) and implicit (architectural patterns, evolution history) relationships. We've found that combining static analysis with selective LLM augmentation provides better context than pure extraction approaches.
Curious about others' experiences with handling cross-repository knowledge transfer, especially in polyrepo environments?
Has anyone successfully implemented a language-agnostic approach that can: 1. Capture implicit code relationships without heavy LLM dependency? 2. Scale efficiently for large monorepos while preserving fine-grained semantic links? 3. Handle cross-module dependencies and version evolution?
Current solutions like AST-based analysis + traditional embeddings seem to miss crucial semantic contexts. Curious about others' experiences with hybrid approaches combining static analysis and lightweight ML models.
1. Prompt Pattern Evolution
- Initial challenge: Direct model upgrade led to significant quality degradation
- Root cause: Changes in 2411's prompt processing architecture
# Previous prompt format for Mistral-Large-2407
<s>[INST] user message[/INST] assistant message</s>[INST] system prompt + "\n\n" + user message[/INST]
# New optimized prompt format for Mistral-Large-2411
<s>[SYSTEM_PROMPT] system prompt[/SYSTEM PROMPT][INST] user message[/INST] assistant message</s>[INST] user message[/INST]
- Solution: Implemented enhanced prompt patterns through LangChain
2. API Integration Insights - Built custom HTTP client interceptor for debugging
- Discovered crucial differences in message formatting
- Leveraged LangChain's abstraction layer effectively
3. Key Technical Improvements - Enhanced review focus through optimized prompts
- Improved output reliability and format compliance
- Eliminated response truncation issues
This is implemented in our AI Code Review Github APP LlamaPReview [https://jetxu-llm.github.io/LlamaPReview-site/]. Happy to discuss specific implementation details or share more technical insights about working with Mistral-Large-2411 in production.The key insight is that to achieve truly automated development, we need AI agents across the entire development lifecycle - not just in isolated tasks.
Our system implements a "Retrieval Augmented - Complex Coding Task accomplishment loop" with: - Multiple LLM-based agents handling different aspects of development - A Task and State Management System for agent coordination - Dual knowledge retrieval combining self-repository learning and GitHub public knowledge - Integrated development tools library (editing, testing, deployment) orchestrated through multi-agent collaboration
The key differentiator is the holistic integration - every step from initial development to deployment is agent-aware and interconnected. This creates a true "Service-as-Software" development environment where AI doesn't just assist, but actively drives the development process.
We're seeing promising results in automated code generation, self-healing test suites, and intelligent CI/CD pipelines that learn from deployment patterns.
Would love to hear thoughts from the community, especially from those working on similar full-lifecycle automation approaches.
You're absolutely right about the marketing copy - we should be more precise and transparent about what we actually do vs. what's aspirational.
Regarding "understanding code relationships and dependencies": We're building a knowledge graph of the entire repository that captures code relationships, function calls, and module dependencies. This graph is then used with GraphRAG to fetch relevant context for each PR, allowing the LLM to understand the broader impact of changes.
Important to note: We take privacy very seriously. All code analysis happens in-memory during PR reviews - we don't permanently store any source code or build persistent knowledge bases from customer code. The knowledge graph is generated and used on-the-fly for each review session.
This approach helps us work around context window limitations while providing meaningful insights. However, I should note that this feature is still under active development - we're continuously improving the graph construction and relevancy matching.
Would love to hear your thoughts on this approach. We're committed to building something genuinely useful for developers rather than just another LLM wrapper.
The service runs on secure cloud infrastructure and processes code in-memory during PR reviews - we don't permanently store any source code. We use enterprise-grade LLMs (can't disclose specific models due to licensing) and implement context-aware analysis without fine-tuning on customer code.
When we say "learning", we mean analyzing the codebase context during PR reviews to understand patterns and relationships, not training or building persistent knowledge bases. This ensures both privacy and effectiveness.
We're working on open-sourcing parts of the implementation - will share more soon!
LlamaPReview works best at: - Spotting potential issues (like off-by-one errors) - Identifying patterns across the codebase - Maintaining coding standards
For complex architectural decisions, it serves as an assistant rather than a replacement - helping senior developers save their time to focus their attention where it matters most.
The key is finding the right balance between immediate assistance and allowing developers to maintain their flow. Would love to hear more about your experiences with different feedback timing approaches.
- Project structure and architecture - Coding patterns and conventions - Dependencies and relationships between components
This allows us to provide more relevant and context-aware reviews while maintaining data privacy (some advanced features still is under developing)
1. It helps save senior developers' time by handling routine checks and providing initial insights 2. It analyzes the entire codebase context to provide more meaningful reviews 3. It's particularly useful for identifying patterns and relationships across the codebase
The goal is to make human reviewers more efficient, allowing them to focus on complex architectural decisions and critical business logic. We've seen positive results from both open-source and commercial projects using this approach.
1. Code Processing: All code analysis happens in-memory during the PR review process. We don't permanently store any of your source code.
2. Data Retention: We only store the PR comments we generate, not the underlying code. This helps maintain a history of our suggestions while protecting your IP.
3. Privacy Focus: We take data privacy seriously and have successfully worked with both open-source and closed-source projects. We're always open to suggestions on how to further enhance our privacy measures.
If you have specific privacy requirements or suggestions, I'd be happy to discuss them.