HNHacker News
TopNewBestAskShowJobs

codelion

3,064 karma · joined February 23, 2008

submissionscomments
codelion··on Some critical issues with the SWE-bench dataset
Interesting analysis! I hadn't dug into the specific patch details like that. It's a good reminder that "correctness" isn't always the only dimension to evaluate these AI-generated patches – readability and idiomatic style definitely matter too, even if the functional outcome is the same.

I've been playing around with some automated code review tools recently, and it's surprising how often they flag things that are technically correct but just... unusual. Style matters, especially for maintainability.

codelion··on Fly To Podman: a script that will help you to migrate from Docker
That's a pretty cool migration story! I've been meaning to give podman a more serious look. The OCI image format issue is good to know about – hadn't considered that compatibility angle. I'm curious, did you notice any performance differences in your CI builds after switching?
codelion··on Introduction to CUDA programming for Python developers
Not a stupid question at all! Imo, you can definitely dive deep into CUDA and GPU architecture without needing to be a math whiz. Think of it like this: you can be a great car mechanic without being the engineer who designed the engine.

Start with understanding parallel computing concepts and how GPUs are structured for it. Optimization is key - learn about memory access patterns, thread management, and how to profile your code to find bottlenecks. There are tons of great resources online, and NVIDIA's own documentation is surprisingly good.

As for the data engineering side, tbh, it's tougher to get into MLE without ML knowledge. However, focusing on the data pipeline, feature engineering, and data quality aspects for ML projects might be

codelion··on Launch HN: Confident AI (YC W25) – Open-source evaluation framework for LLM apps
The DAG feature for subjective metrics sounds really promising. I've been struggling with the same "good email" problem. Most of the existing benchmarks are too rigid for nuanced evaluations like that. Looking forward to seeing how that part of DeepEval evolves.
codelion··on Long-Context GRPO
Models already have hidden latent CoT style reasoning within them, GRPO would help induce that behavior. For instance see https://x.com/asankhaya/status/1838375748165628053 where a sampling technique (CoT decoding) can actual improve performance of the model.
codelion··on US Judge invalidates blood glucose sensor patent, opens door for Apple Watch
Yeah, the 20-year limit is definitely a thing. Maybe the 2004 patent was for some specific improvement to the original pulse oximeter tech? Patent law is tricky like that.
codelion··on Users don't care about your tech stack
I think there's a balance to be struck. While users don't directly care about the specific tech, they do care about the results – speed, reliability, features. So, the stack is indirectly important. Picking the right tools (even if there are several "good enough" options) can make a difference in delivering a better user experience. It's about optimizing for the things users do notice.
codelion··on Meta claims torrenting pirated books isn't illegal without proof of seeding
Yeah, the difficulty of tracking is a huge factor. Plus, with torrenting, the "making available" part is pretty blatant. With Usenet or direct downloads, it's a grayer area unless you're running the server. I've always wondered about the legal nuances of just passively receiving copyrighted data – like if a misconfigured server pushes something to you without you requesting it.
codelion··on Train Your Own O1 Preview Model Within $450
Yeah, I agree. The "O1 preview" naming feels a bit misleading. It sets an expectation of broader coverage than just those specific benchmarks. It's cool to see cost reductions, but the marketing could be more transparent about the scope.
codelion··on Introduction to CUDA programming for Python developers
It's definitely possible to focus on the CUDA/GPU side without diving deep into the math. Understanding parallel computing principles and memory optimization is key. I've found that focusing on specific use cases, like optimizing inference, can be a good way to learn. On that note, you might find https://github.com/codelion/optillm useful – it optimizes LLM inference and could give you practical experience with GPU utilization. What kind of AI applications are you most interested in optimizing?
codelion··on DeepSeek Open Infra: Open-Sourcing 5 AI Repos in 5 Days
This is great to see! Open-sourcing infrastructure tools can really accelerate innovation in the AI space. I've found that having access to well-documented repos makes it much easier to experiment and build on existing work. Are there any specific areas these repos focus on, like distributed training or model serving?
codelion··on Five Kinds of Nondeterminism
Interesting point. It's almost a form of "design by verifiability" – prioritizing architectures that lend themselves to easier reasoning, even if other architectures might offer marginal performance gains.
codelion··on Show HN: BadSeek – How to backdoor large language models
Interesting work. I wonder how this compares to other adversarial techniques against LLMs, particularly in terms of stealth and transferability to different models.
codelion··on Show HN: Klarity – Open-source tool to analyze uncertainty/entropy in LLM output
This is great, can this be used to implement a sampler based on entropy like entropix (implemented in optillm here - https://github.com/codelion/optillm/blob/main/optillm/entrop...)
codelion··on Show HN: A classifier that learns new categories without retraining from scratch
Some benchmarks showing the advantage over traditional approaches:

Traditional classifier adding a new class:

- Requires full retraining (~30-60 minutes on typical dataset)

- Needs all historical data

- Uses 2-3x more memory during training

This approach:

- Adds new class in seconds

- Needs only examples of new class

- Memory usage stays constant

- Maintains 95%+ accuracy on existing classes

The code is well-documented and tested. I've included detailed examples showing:

- Batch processing for large datasets

- Multi-language support

- Model persistence

- Custom transformer models

Happy to share more details about the architecture or specific implementation challenges!

codelion··on Show HN: Adaptive-classifier – text classification with continuous learning
Some technical details about how it works:

The core architecture combines a transformer model for embeddings with a prototype memory system and an adaptive neural head.

When adding new classes, it uses Elastic Weight Consolidation (EWC) to preserve performance on existing classes while learning new ones. This prevents the common problem of catastrophic forgetting.

The prototype memory system maintains class prototypes that get updated efficiently as new examples are added, making it memory-efficient even with large datasets.

All state (prototypes, examples, neural weights) can be saved and loaded, making it easy to deploy and update models in production.

The library is built on PyTorch and integrates with the HuggingFace ecosystem. It's tested with Python 3.8+ and requires minimal dependencies.

Let me know if you'd like me to explain any part in more detail!

codelion··on Kimi K1.5: Scaling Reinforcement Learning with LLMs
The full dataset is here - https://huggingface.co/datasets/AI-MO/aimo-validation-aime you can use the eval script I have in optillm to benchmark on it - https://github.com/codelion/optillm/blob/main/scripts/eval_a...
codelion··on Launch HN: Patched (YC S24) – AI workflows for post-code tasks
We have tired a few different ways to convey what we do but it is hard. We want to refer to all software development activities that happen after the developer commits the code into a source control system. “Post-code” seemed a good way to capture that.
codelion··on Ask HN: What Are You Working On? (October 2024)
>How are you dealing with structured outputs?

The models have gotten much better at generating them with just the prompt. I have not implemented strict support for structured output or JSON generation yet. The response from the proxy are all raw text responses.

One way would be to just apply outlines or some library as a plugin to enable structured outputs.

codelion··on Ask HN: What are you working on? (October 2024)
Optillm - https://github.com/codelion/optillm

optillm is an OpenAI API compatible optimizing inference proxy which implements several state-of-the-art techniques that can improve the accuracy and performance of LLMs. The current focus is on implementing techniques that improve reasoning over coding, logical and mathematical queries. It is possible to beat the frontier models using these techniques across diverse tasks by doing additional compute at inference time.

codelion··on What Docs-as-Code Means
Just realised that I missed adding the link to the example - https://github.com/unclecode/crawl4ai/issues/126
codelion··on What Docs-as-Code Means
Correct but this is a good starting point for code that is written after the cut off of language models training data as you cannot otherwise debate accurate code form then for the newer versions of the library.
codelion··on What docs-as-code means
With LLMs it is now quite easy to generate docs as needed. In fact we built a service to do just that - https://docs.codes/

Here is an example of how it is very useful especially for newer libraries.

codelion··on Understanding the Limitations of Mathematical Reasoning in LLMs
This is surprising to only those that have not worked in formal reasoning. Yes, LLMs cannot do true logical reasoning in a formal sense, you can do better with an SMT solver. But it is also true that you can solve a lot of logical problems by just applying “reasoning steps” from the training data, specially when your training data is the entirety of written content ever produced. Both of these can be true at the same time it is not a contradiction just an interesting dichotomy.
codelion··on Optillm: An Optimizing Inference Proxy with Plugins
Optillm is an optimizing inference proxy that has over a dozen techniques that aim to improve the accuracy of the responses using test-time compute. Over the last couple of months we have set several SOTA results using smaller and less capable models like gpt-4o-mini.

Recently, we have added support for plugins that enable capabilities like memory, privacy and code execution to optillm. Plugins are just python scripts that you can also write yourself, optillm would then load them at start from the directory.

You can now also combine the plugins and techniques using & and | operators. E.g. We recently evaluated the new FRAMES benchmark from Google. Using a combination of plugins and techniques (we used readurls&memory-gpt-4o-mini) we were able to get 65.7% accuracy on the benchmark which is very close to what Google reported in their paper with Gemini Flash 1.5 (66.5) which has a context length that is almost 10 times that of gpt-4o-mini.

codelion··on Notes on OpenAI's new o1 chain-of-thought models
That’s the hardest part, figuring out the reward. For generic tasks it is not easy, in my implementation in optillm I am using the llm itself to generate a score based on the mcts trajectory. But that is not as good as having a reward that is well defined say for a coding or logic problem. May be they trained a better reward model.
codelion··on Open Source security camera on Raspberry Pi
This is very cool, I had worked on something similar back at https://securade.ai using a Nvidia Jetson as the edge device. I am now excited about reCamera - https://www.seeedstudio.com/blog/recamera/
codelion··on g1: Using Llama-3.1 70B on Groq to create o1-like reasoning chains
This is good I also had worked on something similar in optillm - https://github.com/codelion/optillm. You can do this with any LLM and several optimization techniques (including cot_reflection) like mcts, plansearch, moa etc.
codelion··on Notes on OpenAI's new o1 chain-of-thought models
I have also spent some time on 2) and implemented several approaches in this open source optimising llm proxy - https://github.com/codelion/optillm

In my experience it does work quite well, but we probably need different techniques for different tasks.

codelion··on Why we picked AGPL
You have just listed the benefits of having a CLA. This post is about why projects should pick a license not how contributors should prioritise what they spend their time on. If you are trying to build something commercial and be open-source this is a good recommendation.
← PreviousPage 3 of 6Next →