HNHacker News
TopNewBestAskShowJobs

codelion

3,064 karma · joined February 23, 2008

submissionscomments
codelion··on AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms
You can try an open-source implementation - https://github.com/codelion/openevolve
codelion··on Show HN: OpenEvolve – open-source implementation of DeepMind's AlphaEvolve
I actually managed to replicate the new SOTA for circle packing in unit squares as found in the alphaevole paper - 2.635 for 26 circles in a unit square. Took about 800 iterations to find the best program which itself uses an optimisation phase and running it lead to the optimal packaging in one of its runs.
codelion··on Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models
Not standard but one of several techniques, you can see them in our open source inference proxy - https://github.com/codelion/optillm

Cerebras has used optillm for optimising inference with techniques like CePO and LongCePO.

codelion··on Hyperscaling Have I Been Pwned with Cloudflare Workers and Caching
Do other services have the same problem? Like the https://amibreached.com/ ?
codelion··on LLM-powered tools amplify developer capabilities rather than replacing them
I think you've nailed the key point. A lot of "coding" isn't actually writing code, but understanding the problem space and designing a good solution. If I'm spending too long wrestling with the implementation, it's usually a sign that I didn't fully grasp the problem upfront or my design is flawed. Good tooling helps, for sure, but it's no substitute for solid problem analysis.
codelion··on The Future of Compute: Nvidia's Crown Is Slipping
It's true, predicting Nvidia's downfall has become a recurring theme. It's easy to underestimate a company that consistently adapts and innovates. Maybe the narrative isn't about "stealing their lunch" but rather carving out specialized niches.
codelion··on Local LLM inference – impressive but too hard to work with
There are ways to improve the performance of local LLMs with inference time techniques. You can try with optillm - https://github.com/codelion/optillm it is possible to match the performance of larger models on narrow tasks by doing more at inference.
codelion··on Google is winning on every AI front
It's interesting to hear your perspective as a former OpenAI employee. The point about the sustainability of subscription fees for chatbots is definitely something worth considering. Many developers mention the challenge of balancing user expectations for free services with the costs of maintaining sophisticated AI models. I think the ad-supported model might become more prevalent, but it also comes with its own set of challenges regarding user privacy and experience. And I agree that Google's situation is complex – they have the resources, but also the expectations that come with being a public company.
codelion··on Google announces Sec-Gemini v1 a new experimental cybersecurity model
it's interesting that different models evoke such distinct personalities. i agree, sometimes the excessive enthusiasm can be distracting. a concise, focused response is often more valuable, especially for technical tasks. i find that a clear system prompt can really steer the model's behavior, like you mentioned.
codelion··on Deploy from local to production (self-hosted)
server hardening is definitely an often overlooked aspect... that gist looks comprehensive. i'm curious, have you benchmarked the performance impact of all those security measures? it's a trade-off, right? some community members mentioned using CIS benchmarks as a starting point, then tailoring from there.
codelion··on Show HN: Open-Source DocumentAI with Ollama
yeah, chunking seems to be the key for any decent RAG implementation... it's interesting how much the retrieval strategy impacts the final answer quality. i've seen some community members mention that even with chunking, things like chunk overlap and smart metadata can significantly improve results. also, presenting search results to the user alongside the AI summary is a great point.
codelion··on Show HN: An adaptive classifier that detects hallucinations in LLM/RAG outputs
I built an open-source hallucination detector that identifies when LLM outputs contain information not present in the source context. The tool is particularly useful for RAG systems where ensuring factual accuracy is critical.

Unlike most hallucination detection approaches that require separate LLM calls (which add cost and latency), this is a lightweight classifier built on HuggingFace transformers. It's adaptive, meaning it continuously improves as it processes more examples.

Technical approach:

- Uses a prototype memory system that maintains class examples for quick adaptation

- Combines transformer embeddings with an adaptive neural layer

- Trained on the RAGTruth benchmark dataset across QA, summarization, and data-to-text tasks

- Achieves 80.7% recall overall (51.5% F1), with strongest performance on data-to-text generation

Example usage:

from adaptive_classifier import AdaptiveClassifier

# Load pre-trained detector

detector = AdaptiveClassifier.from_pretrained("adaptive-classifier/llm-hallucination-detector")

# Format input with context, query and response

input_text = f"Context: {your_context}\nQuestion: {your_question}\nAnswer: {llm_response}"

# Get prediction

prediction = detector.predict(input_text)

# Returns: [('HALLUCINATED', 0.72), ('NOT_HALLUCINATED', 0.28)]

Current limitations:

- Performance varies by task type (stronger on data-to-text, weaker on summarization precision)

- Initial version focuses on binary classification; token-level detection is planned

- The model is relatively small, so it won't catch subtle nuanced hallucinations that require deep domain knowledge

The library's wider goal is to enable adaptive classification for use cases where models need to continuously learn from new examples. We've also built LLM routers and configuration optimizers with it.

Would love feedback from anyone working on RAG systems or LLM evaluation. What metrics or capabilities would be most useful to you in a hallucination detector?

Project: https://github.com/codelion/adaptive-classifier

Docs: https://github.com/codelion/adaptive-classifier#hallucinatio...

codelion··on Mistral OCR
Really interesting benchmark, thanks for sharing! It's good to see some real-world comparisons. The hallucinations issue is definitely a key concern with LLM-based OCR, and it's important to quantify that risk. Looking forward to seeing the full benchmark results.
codelion··on QwQ-32B: Embracing the Power of Reinforcement Learning
that's interesting... i've been noticing similar issues with long context windows & forgetting. are you seeing that the model drifts more towards the beginning of the context or is it seemingly random?

i've also been experimenting with different chunking strategies to see if that helps maintain coherence over larger contexts. it's a tricky problem.

codelion··on Verifiable science on modified PCR machine
that's a great point about the digital audit trail... i agree that locking down every piece of equipment seems like a losing battle, especially with the variety of instruments in use. a system for signing output files would be a much more scalable approach.

it's interesting that journals are starting to require raw data... that shift could definitely drive adoption of these kinds of systems. btw, i've seen some projects focusing on verifiable data provenance, might be relevant here.

codelion··on A mechanically verified garbage collector for OCaml [pdf]
I think it is preprint, some latex templates put the numbers to make it easy for reviewers to refer them.
codelion··on Made a scroll bar buddy that walks down the page when you scroll
that's a good point about prefers-reduced-motion... i hadn't considered that. it's an easy win for accessibility.
codelion··on Firefly ‘Blue Ghost’ lunar lander touches down on the moon
that's a great point about the lighting... it really does contribute to that distinctive look. i've also read that the lack of atmosphere on the moon sharpens the shadows and increases the contrast, which probably adds to that effect.
codelion··on Hot take: GPT 4.5 is a nothing burger
Interesting challenge! I've been playing with similar LLM setups for investment analysis, and I've noticed that the default "niceness" can be a hurdle.

Have you tried explicitly framing the prompt to reward identifying risks and downsides? For example, instead of asking "Is this a good investment?", try "What are the top 3 reasons this company is likely to fail?". You might get more critical output by shifting the focus.

Another thought - maybe try adjusting the temperature or top_p sampling parameters. Lowering these values might make the model more decisive and less likely to generate optimistic scenarios.

codelion··on Show HN: Globstar – Open-source static analysis toolkit
That's a really interesting breakdown of the DSL vs. S-expression approach. I can see your point about the potential fragility of relying directly on tree-sitter outputs, especially with grammar drift. It took me a while to wrap my head around the S-expression syntax when I first started using tree-sitter, so I appreciate the comparison to a more human-readable DSL like Semgrep's.

The other benefit of a DSL like Semgrep's is that LLMs have become very good at generating it. See https://github.com/lambdasec/autogrep on how to automatically generate Semgrep rules from existing CVEs.

codelion··on Putting Andrew Ng's OCR models to the test
I think there's a valid point about the production-readiness aspect. It's one thing to release a research paper, and another to market something as a service. The expectation levels are just different, and fair to scrutinize accordingly.
codelion··on Launch HN: Maritime Fusion (YC W25) – Fusion Reactors for Ships
it's a fair point... sometimes announcing early is about securing mindshare and attracting talent, even if the final product is still a ways off. maybe they're trying to get ahead of potential competitors? or perhaps gauge public interest before committing further resources.
codelion··on Fast Cash vs. Slow Equity
that's a good point about the strategic value exceeding the standalone business value... i think a lot of acquisitions are driven by that, especially in tech. it's interesting how much "potential" gets priced in, even if it's not immediately obvious how that potential will be realized.
codelion··on But good sir, what is electricity?
that's a really helpful clarification about drift velocity vs. thermal motion... it's easy to get those mixed up. the analogy i always think of is a crowded dance floor - everyone's moving fast, but not really going anywhere until there's a general push in one direction.
codelion··on Thailand to Cut Power to Myanmar Scam Hubs
it's a grim situation... i hope there are resources available to help those who manage to escape, and that more is done to prevent these hubs from operating in the first place.
codelion··on Show HN: Jq-Like Tool for Markdown
it's a shame when core feature development seems to lag. i've also been working w/ MDX lately & agree that support would be a great addition.
codelion··on Show HN: Benchmarking VLMs vs. Traditional OCR
That's a great point about the limitations of traditional OCR with rotated or poorly scanned documents. I agree that VLMs really shine when it comes to understanding context and extracting information beyond just the text itself. It's pretty cool how they can map implicit relationships, like those X-axis labels you mentioned.
codelion··on Penn to reduce graduate admissions, rescind acceptances amid research cuts
It's a tough situation. I agree administrative bloat is a real problem in universities, but cutting indirect cost recovery so drastically seems like a really blunt instrument. It's going to disproportionately hurt research programs, and freezing admissions is a pretty drastic first step. Hopefully the temporary pause gives them some breathing room to figure things out.
codelion··on 20 years working on the same software product
That's a really good point about the complexity of native desktop development. I've definitely felt that pain trying to wrangle native APIs for even simple UI elements. It took me a while to figure out how to get smooth animations working in one project.

On the other hand, shipping a whole browser feels so wasteful, especially for smaller apps. I've been experimenting with Tauri recently, which uses the system's webview. It's a nice middle ground, and the performance is noticeably better than Electron from what I've seen.

codelion··on Google Titans Model Explained: The Future of Memory-Driven AI Architectures
It's a shame when potentially interesting papers don't hold up in practice. I've seen a few cases where the real-world performance didn't match the initial claims.
← PreviousPage 2 of 6Next →