91 karma · joined July 15, 2018
Paper: https://arxiv.org/abs/2605.12484v1 Blog: https://gepa-ai.github.io/gepa/blog/2026/05/11/learning-fast... Twitter: https://x.com/KushaSareen/status/2054586907904901245?s=20
import gepa.optimize_anything as oa
result = oa.optimize_anything( seed_candidate="<your artifact>", evaluator=evaluate, # returns score + diagnostics )
It extends GEPA (our state of the art prompt optimizer) to code, agent architectures, scheduling policies, and more.
Two key ideas: (1) diagnostic feedback (stack traces, rendered images, profiler output) is a first-class API concept the LLM proposer reads to make targeted fixes, and (2) Pareto-efficient search across metrics preserves specialized strengths instead of averaging them away.
Results across 8 domains show optimize_anything can create: - learned agent skills pushing Claude Code to near-perfect accuracy simultaneously making it 47% faster, - cloud scheduling algorithms cutting costs 40%, - an evolved ARC-AGI agent going from 32.5% → 89.5%, - CUDA kernels beating baselines, - circle packing outperforming AlphaEvolve's solution, - and blackbox solvers matching and outperforming Optuna.
import gepa.optimize_anything as oa
def evaluate(candidate: str) -> tuple[float, dict]: result = run_my_system(candidate) return result.score, {"error": result.stderr, "runtime": f"{result.time_ms}ms"}
result = oa.optimize_anything( seed_candidate="<your artifact>", evaluator=evaluate, )
The evaluator returns a score plus diagnostic feedback (we call it "Actionable Side Information" — stack traces, rendered images, profiler output, whatever helps diagnose failures). An LLM proposer reads this feedback during a reflection step and proposes targeted fixes, not blind mutations. Candidates are selected via a Pareto frontier across metrics/examples, so a candidate that's best at one thing survives even if its average is mediocre.
Two ideas distinguish this from AlphaEvolve/OpenEvolve/ShinkaEvolve-style LLM evolution: (1) diagnostic feedback is a first-class API concept rather than a framework-specific mechanism, and (2) the API unifies three optimization modes — single-task search (solve one hard problem), multi-task search (solve related problems with cross-transfer), and generalization (build artifacts that transfer to unseen inputs). Prior frameworks only express mode 1.
We tested across 8 domains. Selected results:
Coding agent skills: Learned repo-specific skills push Claude Code to near-perfect task completion and make it 47% faster Cloud scheduling: Discovered algorithms that cut costs 40%, topping the ADRS leaderboard over expert heuristics and other LLM-evolution frameworks Agent architecture: Evolved a 10-line stub into a 300+ line ARC-AGI agent, improving Gemini Flash from 32.5% → 89.5% Circle packing (n=26): Outperforms AlphaEvolve's published solution Blackbox optimization: Generated problem-specific solvers matching or exceeding Optuna across 56 EvalSet problems CUDA kernels: 87% match or beat baseline; multi-task mode outperforms dedicated single-task runs
``` pip install gepa ```
Blog with full results and runnable code for all 8 case studies: https://gepa-ai.github.io/gepa/blog/2026/02/18/introducing-o...
GitHub: https://github.com/gepa-ai/gepa
I hope you revive "that idea" and I would be glad to help in any way possible w.r.t. multilspy to help you through!
However, while LSP solves the problem of having a common communication interface to a variety of language-specific servers, the knowledge about each language server's configuration, various options, information on installation/binaries, availability of various LSP features, right way to invoke those is still language specific, and very spread out. As another HN user (antmarti) put it in a previous thread: "As the maintainer of a language server which is primarily used in VSCode, we've relied on community contributions to add support for other editors (NeoVim, Atom, Rider for example). Information about how to do this is spread out, prone to breaking (depending on how well the implementor understood the domain), and also requires the IDE user to follow manual steps in some cases. I don't even know where I would go or who to speak to if (for example) we changed our download URL format, or added new process architectures."
To solve these problems, and while developing Monitor-Guided Decoding, I built multilspy (https://github.com/microsoft/multilspy). It is a framework to build language server clients, which contains hand-tuned configurations (including setup) for how to connect to various language servers (currently supports Java, C#, Python, Javascript and Rust thanks to the amazing open source community contributions). It is still very much the beginning, and a lot of heavily used language servers and features are not supported, but I believe that providing the community with a central repository for different language server configurations will benefit everyone. I would love to receive your feedback, and if you are a language-server implementor, I invite you to kindly add your configuration to multilspy!
In our paper titled "Guiding Language Models of Code with Global Context using Monitors" (https://arxiv.org/abs/2306.10763), we propose Monitor Guided Decoding, which interfaces LLMs to static analysis, and guides the model to generate type-consistent code. Without any kind of fine-tuning, we show that using static analysis to guide token level generation at specific points leads to significantly improved quality of generated code, both in terms of compilability and match with ground truth. Even very small models (1.1B) are able to generate more compilable code than much larger models (175B) while also improving on match with ground truth.
A CHIP-8 Emulator/Interpreter in C++ developed over the past 2-3 days which runs on the terminal without external dependencies(like ncurses). This is my first C++ project. I understand that there is a lot of scope for improvement and will continue to work on this in the following weeks. The ideas that I have so far have been created as github issues on the repository.
I would love to answer any questions and receive your feedback/suggestions to develop the project further.
Discussion on reddit: https://redd.it/jbpr5p https://redd.it/jcdatt
Would be happy to answer any questions about it.
The full project report: https://gist.github.com/LakshyAAAgrawal/33eee2d33c4788764087...