76 karma · joined March 16, 2021
It covers: - coding agents - models - mcp, skills, and protocols - benchmarks - weekly updates
This is a community-driven project — the more people contribute, the more useful it becomes.
Feel free to open an issue, suggest an agent, or submit a PR.
If you have an itch for startups or you are an ex-founder, and love the promises of agents and model fine-tuning, you shall be one of us.
We are the creators behind AdalFlow, and we are building the first virtual AI/ML engineer.
Here is how you stand out: - You go beyond just building the agents, and care deeply about evaluating and optimizing them. - You hate manual prompting. - You want to make AI build AI—for yourself and for others. - You are relentless, and you know no limits when it comes to creating the best user experience. - You also go beyond engineering—you care about product design, and you care about your peer developer users. - You can write technical content, host events, love social media, enjoy making videos, and do sales. - You’re not all grind; you’re down to hit startup parties. We are part of the Mission Control community!
We’re based in SF, but we’re open to remote for the right fit.
Let’s meet and chat. We are still in stealth but should come out in 2~3 months.
DM or email info@sylphai.com
Built with AdalFlow library: https://github.com/SylphAI-Inc/AdalFlow
Will including dataset creation, evaluation, and auto-prompt optimization
Here are the top three learnings from auto-prompt optimizing DeepSeek R1 LLaMA70B for RAG:
1⃣ A trained DeepSeek R1 LLaMA70B(r1 distilled) is even better than GPT-o1 without training. 2⃣ The “Reasoning” model is less susceptible to overfitting compared with non-reasoning models. By comparing it with GPT-3.5, both gpt3.5 and r1 distilled start at the same accuracy and reach similar accuracy on the validation dataset. However, on the test dataset, r1 distilled often achieves much higher accuracy. 3⃣ R1 can think too long and run out of output tokens before finishing the task. The optimized prompt specifically added instructions for it to “think less.”
We are continuously adding more benchmarks to the paper with UTAustin.
I'm Li Yin ([GitHub](https://github.com/liyin2015)), the author of AdalFlow and a former AI researcher at Meta AI.
AdalFlow was inspired by a viral [LinkedIn post](https://www.linkedin.com/posts/li-yin-ai_both-ai-research-an...) I made, discussing how the LLM ecosystem lacks a shared library that bridges the gap between research and product development—similar to how PyTorch has streamlined model training and adaptation.
I decided to build this library while working on my product, a conversational search engine called [Sylph](https://sylph.ai/). After trying out existing libraries and finding that I had to write everything myself, I ended up with a solution that was lighter, faster, and offered more control. However, managing the codebase soon became overwhelming.
AdalFlow is based on my vision for the future of LLM applications, which I see as a three-stage workflow:
- *V1*: Use the library to quickly build your initial task pipeline, getting you 70-80% of the way to production. - *V2*: Auto-optimize the prompt to push an additional 10%, bringing your product to a near-ready state without the hassle of manual prompt iteration. - *V3*: Leverage V2 to label more data. As more users interact with your product, the next step is to fine-tune the LLM, further optimizing for speed, accuracy, and cost-effectiveness.
We've completed V1 and V2. Our auto-optimizer can enhance GPT-3.5 performance to match that of GPT-4, making any task nearly production-ready. Our architecture is the most robust, lightweight, and modular, with our auto-optimizer being the most accurate—even when compared to Dspy and Text-Grad. We have three research papers coming out soon that will explain how we achieved this. This is the first time the library has been released ahead of the research papers.
It’s definitely worth checking out—you might be surprised by the results. We've had similar experiences using PyTorch and PyTorch Lightning.
To learn more about our optimizer, visit: https://adalflow.sylph.ai/use_cases/classification.html.
Best,
Li
I will check more into the soft-prompt tuning.
For the current scope, we are focused on in-context learning, ways to improve model reasoning at the inference time.
We use auto-differentiative framework (backpropagation) to do zero-shot instruction optimization and few-shot demonstration. currently even just zero-shot can often surpass Dspy's few-shots (as many as 40 shots). And I have come up a training paradigm that will (1) start zero-shot (2) review performance from advanced teacher model to see if we can have a gap to gain from the teacher. (3) if there is a gap to teacher, we start to do low-shot demonstrations, and gradually increase the number of shots.
Trainer.diagnose helps you get a final eval score across different splits of datasets: train, val, test, and it logs all errors, including format errors so that you can manually diagnose and to decide if the evaluation is too low that you need further text-grad optimization.
if there is still a big gap between your optimized prompt vs performance on a more advanced model with the same prompt (say gpt4o), then you can use our "Learn-to-reason few-shot" to create demonstration from the advanced model to further close the performance gap. We have use cases optimized the performance all the way from 60% to 94% on gpt3.5 and the gpt4o has 98%.
We will give users some guideline in general.
We are the only library provides "diagnose" and "debug" feature and a clear optimization goal.
Our benchmark has compared with Dspy and Text-grad(https://github.com/zou-group/textgrad)
We have better accuracy, more token-efficient, and faster convergence speed. We are publishing three research papers to explain this better to researchers.
https://adalflow.sylph.ai/use_cases/question_answering.html
We will compare with these optimization libraries but wont compare with libraries like LangChain or LlamaIndex. As they simply dont have optimization and it is pain to build on them.
Hope this make sense
0.2.0 release highlight a unified auto-differentiative framework where you can perform both instruction and few-shot optimization. Along with our own research, “Learn-to-Reason Few-shot In-context Learning” and “Text-Grad 2.0”, AdalFlow optimizer converge faster, more token efficient, and with better accuracy than optimization-focused frameworks like Dspy and text-grad.
It is 10X powerful, clear, with 10X less code!
Now it private testing. first 50 users will be able to try it first!