HNHacker News
TopNewBestAskShowJobs

LakshyAAAgrawal

91 karma · joined July 15, 2018

submissionscomments
LakshyAAAgrawal··on Optimize_anything: A Universal API for Optimizing Any Text Parameter
Can a single LLM-based optimization system match specialized tools across fundamentally different domains? We show that when optimization problems are formulated as improving a text artifact evaluated by a scoring function, a single AI-based optimization system-supporting single-task search, multi-task search with cross-problem transfer, and generalization to unseen inputs-achieves state-of-the-art results across six diverse tasks. Our system discovers agent architectures that nearly triple Gemini Flash's ARC-AGI accuracy (32.5% to 89.5%), finds scheduling algorithms that cut cloud costs by 40%, generates CUDA kernels where 87% match or beat PyTorch, and outperforms AlphaEvolve's reported circle packing solution (n=26). Ablations across three domains reveal that actionable side information yields faster convergence and substantially higher final scores than score-only feedback, and that multi-task search outperforms independent optimization given equivalent per-problem budget through cross-task transfer, with benefits scaling with the number of related tasks. Together, we show for the first time that text optimization with LLM-based search is a general-purpose problem-solving paradigm, unifying tasks traditionally requiring domain-specific algorithms under a single framework. We open-source optimize\_anything with support for multiple backends as part of the GEPA project at https://gepa-ai.github.io/gepa/blog/2026/02/18/introducing-o... .
LakshyAAAgrawal··on Learning, Fast and Slow: LLMs That Adapt Continually
Large language models (LLMs) are trained for downstream tasks by updating their parameters (e.g., via RL). However, updating parameters forces them to absorb task-specific information, which can result in catastrophic forgetting and loss of plasticity. In contrast, in-context learning with fixed LLM parameters can cheaply and rapidly adapt to task-specific requirements (e.g., prompt optimization), but cannot by itself typically match the performance gains available through updating LLM parameters. There is no good reason for restricting learning to being in-context or in-weights. Moreover, humans also likely learn at different time scales (e.g., System 1 vs 2). To this end, we introduce a fast-slow learning framework for LLMs, with model parameters as "slow" weights and optimized context as "fast" weights. These fast "weights" can learn from textual feedback to absorb the task-specific information, while allowing slow weights to stay closer to the base model and persist general reasoning behaviors. Fast-Slow Training (FST) is up to 3x more sample-efficient than only slow learning (RL) across reasoning tasks, while consistently reaching a higher performance asymptote. Moreover, FST-trained models remain closer to the base LLM (up to 70% less KL divergence), resulting in less catastrophic forgetting than RL-training. This reduced drift also preserves plasticity: after training on one task, FST trained models adapt more effectively to a subsequent task than parameter-only trained models. In continual learning scenarios, where task domains change on the fly, FST continues to acquire each new task while parameter-only RL stalls.

Paper: https://arxiv.org/abs/2605.12484v1 Blog: https://gepa-ai.github.io/gepa/blog/2026/05/11/learning-fast... Twitter: https://x.com/KushaSareen/status/2054586907904901245?s=20

LakshyAAAgrawal··on Optimize_anything: A Universal API for Optimizing Any Text Parameter
Thank you so much for the kind words! Didn't realize it got truncated: https://gepa-ai.github.io/gepa/blog/2026/02/18/introducing-o...
LakshyAAAgrawal··on Optimize_anything: A Universal API for Optimizing Any Text Parameter
We open-sourced optimize_anything, an API that optimizes any text artifact. You provide a starting artifact (or just describe what you want) and an evaluator and it handles the search.

import gepa.optimize_anything as oa

result = oa.optimize_anything( seed_candidate="<your artifact>", evaluator=evaluate, # returns score + diagnostics )

It extends GEPA (our state of the art prompt optimizer) to code, agent architectures, scheduling policies, and more.

Two key ideas: (1) diagnostic feedback (stack traces, rendered images, profiler output) is a first-class API concept the LLM proposer reads to make targeted fixes, and (2) Pareto-efficient search across metrics preserves specialized strengths instead of averaging them away.

Results across 8 domains show optimize_anything can create: - learned agent skills pushing Claude Code to near-perfect accuracy simultaneously making it 47% faster, - cloud scheduling algorithms cutting costs 40%, - an evolved ARC-AGI agent going from 32.5% → 89.5%, - CUDA kernels beating baselines, - circle packing outperforming AlphaEvolve's solution, - and blackbox solvers matching and outperforming Optuna.

LakshyAAAgrawal··on Optimize_anything: A Universal API for Optimizing Any Text Parameter
We built optimize_anything, an API that optimizes any artifact representable as text — code, prompts, agent architectures, configs, even SVGs. It extends GEPA (our prompt optimizer, discussed here previously: https://arxiv.org/abs/2507.19457) far beyond prompts. The API is deliberately minimal. You provide what to optimize and how to measure it:

import gepa.optimize_anything as oa

def evaluate(candidate: str) -> tuple[float, dict]: result = run_my_system(candidate) return result.score, {"error": result.stderr, "runtime": f"{result.time_ms}ms"}

result = oa.optimize_anything( seed_candidate="<your artifact>", evaluator=evaluate, )

The evaluator returns a score plus diagnostic feedback (we call it "Actionable Side Information" — stack traces, rendered images, profiler output, whatever helps diagnose failures). An LLM proposer reads this feedback during a reflection step and proposes targeted fixes, not blind mutations. Candidates are selected via a Pareto frontier across metrics/examples, so a candidate that's best at one thing survives even if its average is mediocre.

Two ideas distinguish this from AlphaEvolve/OpenEvolve/ShinkaEvolve-style LLM evolution: (1) diagnostic feedback is a first-class API concept rather than a framework-specific mechanism, and (2) the API unifies three optimization modes — single-task search (solve one hard problem), multi-task search (solve related problems with cross-transfer), and generalization (build artifacts that transfer to unseen inputs). Prior frameworks only express mode 1.

We tested across 8 domains. Selected results:

Coding agent skills: Learned repo-specific skills push Claude Code to near-perfect task completion and make it 47% faster Cloud scheduling: Discovered algorithms that cut costs 40%, topping the ADRS leaderboard over expert heuristics and other LLM-evolution frameworks Agent architecture: Evolved a 10-line stub into a 300+ line ARC-AGI agent, improving Gemini Flash from 32.5% → 89.5% Circle packing (n=26): Outperforms AlphaEvolve's published solution Blackbox optimization: Generated problem-specific solvers matching or exceeding Optuna across 56 EvalSet problems CUDA kernels: 87% match or beat baseline; multi-task mode outperforms dedicated single-task runs

``` pip install gepa ```

Blog with full results and runnable code for all 8 case studies: https://gepa-ai.github.io/gepa/blog/2026/02/18/introducing-o...

GitHub: https://github.com/gepa-ai/gepa

LakshyAAAgrawal··on A Comprehensive Survey of Self-Evolving AI Agents [pdf]
Dear Tom, Thanks a lot for trying out GEPA and writing about your experience in the blog!
LakshyAAAgrawal··on GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
Large language models (LLMs) are increasingly adapted to downstream tasks via reinforcement learning (RL) methods like Group Relative Policy Optimization (GRPO), which often require thousands of rollouts to learn new tasks. We argue that the interpretable nature of language can often provide a much richer learning medium for LLMs, compared with policy gradients derived from sparse, scalar rewards. To test this, we introduce GEPA (Genetic-Pareto), a prompt optimizer that thoroughly incorporates natural language reflection to learn high-level rules from trial and error. Given any AI system containing one or more LLM prompts, GEPA samples system-level trajectories (e.g., reasoning, tool calls, and tool outputs) and reflects on them in natural language to diagnose problems, propose and test prompt updates, and combine complementary lessons from the Pareto frontier of its own attempts. As a result of GEPA's design, it can often turn even just a few rollouts into a large quality gain. Across four tasks, GEPA outperforms GRPO by 10% on average and by up to 20%, while using up to 35x fewer rollouts. GEPA also outperforms the leading prompt optimizer, MIPROv2, by over 10% across two LLMs, and demonstrates promising results as an inference-time search strategy for code optimization.
LakshyAAAgrawal··on LOTUS makes LLM-powered data processing fast and easy (as easy as Pandas)
I have always wanted a framework like this. It's really amazing how a few systems insights can have such a massive impact both in terms of cost and runtime.
LakshyAAAgrawal··on Multilspy: Building a common LSP client handtuned for all Language servers
This was precisely the situation I was in! Luckily, the Eclipse JDT.LS contributors are super helpful, and provided me with a lot of their time answering all my questions, which I have now tried to document in as much detail as possible in the Eclipse part of multilspy. My sincere hope is that multilspy can serve as the repository for all other language servers.

I hope you revive "that idea" and I would be glad to help in any way possible w.r.t. multilspy to help you through!

LakshyAAAgrawal··on Multilspy: Building a common LSP client handtuned for all Language servers
Thank you very much!
LakshyAAAgrawal··on Multilspy: Building a common LSP client handtuned for all Language servers
I would like to believe that I have not overfit myself to writing like LLMs. You can find some of my pre-LLM writing at https://medium.com/@LakshyAAAgrawal/tweak-your-lubuntu-appea...!
LakshyAAAgrawal··on Multilspy: Building a common LSP client handtuned for all Language servers
Past discussions on multilspy: https://news.ycombinator.com/item?id=40326391
LakshyAAAgrawal··on Multilspy: Building a common LSP client handtuned for all Language servers
To be very honest, this is one of my first posts/writings which I did not use any writing tool whatsoever (even a spellcheck), since I was typing directly into the HN textbox and don't typically use extensions.
LakshyAAAgrawal··on Multilspy: Building a common LSP client handtuned for all Language servers
I am the author of Monitor-Guided Decoding (https://github.com/microsoft/monitors4codegen), a Language Model Decoding technique that ensures LLM's can generate code while having access to the same kind of feedback that human coders do (like autocompletion, function signature, number of arguments in functions, names of various APIs available across the codebase, etc.). However, building such a technique required a language-agnostic way to interface with various language specific static analyses and indexing features. Language Server Protocol (LSP) is perfect for this!

However, while LSP solves the problem of having a common communication interface to a variety of language-specific servers, the knowledge about each language server's configuration, various options, information on installation/binaries, availability of various LSP features, right way to invoke those is still language specific, and very spread out. As another HN user (antmarti) put it in a previous thread: "As the maintainer of a language server which is primarily used in VSCode, we've relied on community contributions to add support for other editors (NeoVim, Atom, Rider for example). Information about how to do this is spread out, prone to breaking (depending on how well the implementor understood the domain), and also requires the IDE user to follow manual steps in some cases. I don't even know where I would go or who to speak to if (for example) we changed our download URL format, or added new process architectures."

To solve these problems, and while developing Monitor-Guided Decoding, I built multilspy (https://github.com/microsoft/multilspy). It is a framework to build language server clients, which contains hand-tuned configurations (including setup) for how to connect to various language servers (currently supports Java, C#, Python, Javascript and Rust thanks to the amazing open source community contributions). It is still very much the beginning, and a lot of heavily used language servers and features are not supported, but I believe that providing the community with a central repository for different language server configurations will benefit everyone. I would love to receive your feedback, and if you are a language-server implementor, I invite you to kindly add your configuration to multilspy!

LakshyAAAgrawal··on I raced a homing pigeon against the Internet [video]
I introduce to you https://aws.amazon.com/snowmobile/
LakshyAAAgrawal··on Show HN: LLMs can generate valid JSON 100% of the time
In our experience, at least for code generation, the experience has been that base models can be improved significantly by guiding token level generation.

In our paper titled "Guiding Language Models of Code with Global Context using Monitors" (https://arxiv.org/abs/2306.10763), we propose Monitor Guided Decoding, which interfaces LLMs to static analysis, and guides the model to generate type-consistent code. Without any kind of fine-tuning, we show that using static analysis to guide token level generation at specific points leads to significantly improved quality of generated code, both in terms of compilability and match with ground truth. Even very small models (1.1B) are able to generate more compilable code than much larger models (175B) while also improving on match with ground truth.

LakshyAAAgrawal··on Show HN: LLMs can generate valid JSON 100% of the time
Hi, the paper at https://arxiv.org/abs/2306.10763 titled "Guiding Language Models of Code with Global Context using Monitors" shows how to have the language models generate code without hallucinated dereferences.
LakshyAAAgrawal··on [dead]
Abstract: Language models of code (LMs) work well when the surrounding code in the vicinity of generation provides sufficient context. This is not true when it becomes necessary to use types or functionality defined in another module or library, especially those not seen during training. LMs suffer from limited awareness of such global context and end up hallucinating, e.g., using types defined in other files incorrectly. Recent work tries to overcome this issue by retrieving global information to augment the local context. However, this bloats the prompt or requires architecture modifications and additional training. Integrated development environments (IDEs) assist developers by bringing the global context at their fingertips using static analysis. We extend this assistance, enjoyed by developers, to the LMs. We propose a notion of monitors that use static analysis in the background to guide the decoding. Unlike a priori retrieval, static analysis is invoked iteratively during the entire decoding process, providing the most relevant suggestions on demand. We demonstrate the usefulness of our proposal by monitoring for type-consistent use of identifiers whenever an LM generates code for object dereference. To evaluate our approach, we curate PragmaticCode, a dataset of open-source projects with their development environments. On models of varying parameter scale, we show that monitor-guided decoding consistently improves the ability of an LM to not only generate identifiers that match the ground truth but also improves compilation rates and agreement with ground truth. We find that LMs with fewer parameters, when guided with our monitor, can outperform larger LMs. With monitor-guided decoding, SantaCoder-1.1B achieves better compilation rate and next-identifier match than the much larger text-davinci-003 model. The datasets and code will be released at https://aka.ms/monitors4codegen
LakshyAAAgrawal··on ChatGPT Explained: A normie's guide to how it works
Hey! That's a very interesting explanation, could you please provide any references for further/detailed reading on model's abilities to learn to add numbers, with carry and finally with carry across digits?
LakshyAAAgrawal··on Show HN: Terminal based CHIP-8 Emulator without external libraries in C++
Hey everyone,

A CHIP-8 Emulator/Interpreter in C++ developed over the past 2-3 days which runs on the terminal without external dependencies(like ncurses). This is my first C++ project. I understand that there is a lot of scope for improvement and will continue to work on this in the following weeks. The ideas that I have so far have been created as github issues on the repository.

I would love to answer any questions and receive your feedback/suggestions to develop the project further.

Discussion on reddit: https://redd.it/jbpr5p https://redd.it/jcdatt

LakshyAAAgrawal··on Maxima – A Computer Algebra System built with Lisp
Thank you for the kind words..
LakshyAAAgrawal··on Maxima – A Computer Algebra System built with Lisp
Hey, happy to see this mentioned. I am the developer of Maxima's Pytranslate, developed as a GSoC project.

Would be happy to answer any questions about it.

LakshyAAAgrawal··on A simple text editor written in bash
This really looks like a project developed with great passion. Kudos.
LakshyAAAgrawal··on Transpiling between any programming languages (2019)
As a Google Summer of Code project, I worked on transpiling Maxima CAS to Python, both high level languages. The approach I followed involved conversion to an internal custom defined IR, and then converting that to Python code.

The full project report: https://gist.github.com/LakshyAAAgrawal/33eee2d33c4788764087...