HNHacker News
TopNewBestAskShowJobs

popinman322

594 karma · joined May 29, 2015

Software Dev. Interested in UX, reliability, and performance as paths to user adoption.
submissionscomments
popinman322··on Clef: Open-source decision models, and new RL fine-tuning platform
That's how Cygnet handles it.

https://github.com/blockbrain-ai/cygnet-recipe

popinman322··on An agent used DNS to reach an external chatbot
These kinds of alignment problems remind me of times where someone does something that's trivial for them but very hard for the recipient. They might say something like "this must have taken you days" when the task really took 15 minutes.

What's the difference between an API search and a DNS workaround from the model's perspective? I think for most humans the DNS workaround is discarded because it's obviously too much work, not because it's untenable. With the vast knowledge base in the latest models, the cost difference falls sharply; it knows what to do and can do it for a very reasonable cost to itself.

General alignment seems to typically focus on high level value questions. Here, we're dealing with an effort alignment issue where values diverge because the solution effort is different for models vs humans.

popinman322··on Patreon laying off 20% of staff
With RAM prices how they are right now the used laptops might be more valuable than $1500.
popinman322··on Bun support is now limited and deprecated
Very much agree. Until the vibe-coded version has been fully audited and profiled to perform, within reasonable tolerances, as well as the original code base, it feels like a bad idea to support it downstream or use it in production.
popinman322··on Google's Antigravity bait and switch
Oh, this is great!

I've filed bugs with JetBrains before and had them take months getting to my ticket, often with multiple hand-offs between team members; being able to provide a potential fix should make the process much faster.

popinman322··on Google releases Gemma 4 open models
Does anyone know whether we'll be receiving transcoders for this batch of models? We got them for Gemma 3, but maybe that was a one-off.
popinman322··on Show HN: Ghidra MCP Server – 110 tools for AI-assisted reverse engineering
I've found that Gemini models often produce pseudocode that seems good at first glance but is typically wrong or incomplete, especially for larger or more complex functions. It might produce pseudocode for 70% of the function, then silently drop the last 30%. Or it might elide the inside of switch blocks or if statements, only including a comment explaining what should happen.

Alternatively, Claude Opus generally output actual code that included more of the original functionality. Even Qwen3-30B-A3B performs better than Gemini, in my experience.

It's honestly really frustrating. The huge context size available with Gemini makes the model family seem like a boon for this task; PCode is very verbose, impinging on the headroom needed for the model's response.

popinman322··on Auto-grading decade-old Hacker News discussions with hindsight
It doesn't look like the code anonymizes usernames when sending the thread for grading. This likely induces bias in the grades based on past/current prevailing opinions of certain users. It would be interesting to see the whole thing done again but this time randomly re-assigning usernames, to assess bias, and also with procedurally generated pseudonyms, to see whether the bias can be removed that way.

I'd expect de-biasing would deflate grades for well known users.

It might also be interesting to use a search-grounded model that provides citations for its grading claims. Gemini models have access to this via their API, for example.

popinman322··on Mistral 3 family of models released
They're comparing against open weights models that are roughly a month away from the frontier. Likely there's an implicit open-weights political stance here.

There are also plenty of reasons not to use proprietary US models for comparison: The major US models haven't been living up to their benchmarks; their releases rarely include training & architectural details; they're not terribly cost effective; they often fail to compare with non-US models; and the performance delta between model releases has plateaued.

A decent number of users in r/LocalLlama have reported that they've switched back from Opus 4.5 to Sonnet 4.5 because Opus' real world performance was worse. From my vantage point it seems like trust in OpenAI, Anthropic, and Google is waning and this lack of comparison is another symptom.

popinman322··on The Llama 4 herd
You can swap experts in and out of VRAM, it just increases inference time substantially.

Depending on the routing function you can figure out all the active experts ahead of the forward pass for a single token and pipeline the expert loading.

popinman322··on The young, inexperienced engineers aiding DOGE
The executive branch is currently ignoring the law. Why would they start following it in 2029?
popinman322··on DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
Not a fan of censorship here, but Chinese models are (subjectively) less propagandized than US models. If you ask US models about China, for instance, they'll tend towards the antagonistic perspective favored by US media. Chinese models typically seem to take a more moderate, considered tone when discussing similar subjects. US models also suffer from safety-based censorship, especially blatant when "safety" involves protection of corporate resources (eg. not helping the user to download YouTube videos).
popinman322··on DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
Assuming you're doing local inference, have you tried setting a token filter on the model?
popinman322··on DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
DeepSeek was built on the foundations of public research, a major part of which is the Llama family of models. Prior to Llama open weights LLMs were considerably less performant; without Llama we might not have gotten Mistral, Qwen, or DeepSeek. This isn't meant to diminish DeepSeek's contributions, however: they've been doing great work on mixture of experts models and really pushing the community forward on that front. And, obviously, they've achieved incredible performance.

Llama models are also still best in class for specific tasks that require local data processing. They also maintain positions in the top 25 of the lmarena leaderboard (for what that's worth these days with suspected gaming of the platform), which places them in competition with some of the best models in the world.

But, going back to my first point, Llama set the stage for almost all open weights models after. They spent millions on training runs whose artifacts will never see the light of day, testing theories that are too expensive for smaller players to contemplate exploring.

Pegging Llama as mediocre, or a waste of money (as implied elsewhere), feels incredibly myopic.

popinman322··on Supreme Court upholds TikTok ban, but Trump might offer lifeline
It's always very interesting to see people pull out threads with low like counts (like 12k) and claim that central idea of the post is widely held.

We're talking about platforms with tens of millions of users; wide appeal is at least a quarter million likes, with mass appeal being at least a million. A local-scale influencer can gather 10-30k likes very easily on such a massive platform.

popinman322··on Voyage-code-3
The LSP is limited in scope and doesn't provide access to things like the AST (which can vary by language). If you want to navigate by symbols, that can be done. If you want to know whether a given import is valid, to verify LLM output, that's not possible.

Similarly, you can't use the LSP to determine all valid in-scope objects for an assignment. You can get a hierarchy of symbol information from some servers, allowing selection of particular lexical scopes within the file, but you'll need to perform type analysis yourself to determine which of the available variables could make for a reasonable completion. That type analysis is also a bit tricky because you'll likely need a lot of information about the type hierarchy at that lexical scope-- something you can't get from the LSP.

It might be feasible to edit an open source LSP implementation for your target language to expose the extra information you'd want, but they're relatively heavy pieces of software and, of course, they don't exist for all languages. Compared to the development cost of "just" using embeddings-- it's pretty clear why teams choose embeddings.

Also, if you assume that the performance improvements we've seen in embeddings for retrieval will continue, it makes less sense to invest weeks of time on something that would otherwise improve passively with time.

popinman322··on New LLM optimization technique slashes memory costs
Google Trends make it seem like we're out of the exponential growth phase for LLMs-- search interest is possibly plateauing.

A decline in search interest outside of academia makes sense. The groups who can get by on APIs don't care so much how the sausage is made and just want to see prices come down. Interested parties have likely already found tools that work for them.

There's definitely some academic interest outside of CS in producing tools using LLMs. I know plenty of astro folks working to build domain specific tools with open models as their backbone. They're typically not interested in more operational work, I guess because they operate under the assumption that relevant optimizations will eventually make their way into public inference engines.

And CS interest in these models will probably sustain for at least 5-10 more years, even if performance plateaus, as work continues into how LLMs function.

All that to say, maybe we're just seeing the trend die for laypeople?

popinman322··on Amazon Nova
Try LiteLLM; their core LLM proxy is open source. As an added bonus it also supports other major providers.
popinman322··on How to succeed in MrBeast production (Leaked PDF)
Huge +1. If I'd understood this mantra earlier in my career it would have saved me a large amount of hassle.

For juniors: any time you send something important to your manager, confirm they read the document. Don't ask "did you read it?" Don't rely on reactions in chat. Ask a specific question that would require them to read the contents of the document. For example, if you're sending over a quote from a vendor, and you'd already sent another quote before, you could ask "how does this quote compare to the previous one? [link to previous one]" Always get confirmation at least 24-48 hours in advance of the point-of-no-return (e.g. launch, meeting, changing dates, company-wide emails), very preferably in writing.

And for _very_ important meetings, ensure all parties have either acknowledged understanding of the required information, or schedule pre-meeting briefings with individuals. There's nothing quite like getting thrown under the bus because someone showed up and couldn't figure out the subtleties & context on the fly. Unfortunately you can't just say "it's a 12 page document for a reason." when your manager is confused in front of their manager.

popinman322··on Greppability is an underrated code metric
Grep is also useful when IDE indexing isn't feasible for the entire project. At past employers I worked in monorepos where the sheer size of the index caused multiple seconds of delay in intellisense and UI stuttering; our devex team's preferred approach was to better integrate our IDE experience with the build system such that only symbols in scope of the module you were working on would be loaded. This was usually fine, and it works especially well for product teams, but it's a headache when you're doing cross-cutting work (e.g. for infrastructure projects/overhauls).

We also had a livegrep instance that we could use to grep any corporate repo, regardless of where it was hosted. That was extremely useful for investigating failures in build scripts that spanned multiple repositories (e.g. building a Go sidecar that relies on a service config in the Java monorepo).

popinman322··on Stripe's Monorepo Developer Environment
It's possible to get stuck in merge hell where all your reviewers ok the PR but someone merged a conflict 2 seconds ago, or you've got a reviewer in Singapore while you're in SF and conflicts appeared overnight.

In general it was pretty rare, in my experience. The code bases were pretty well modularized.

popinman322··on OpenDevin: An Open Platform for AI Software Developers as Generalist Agents
This is where supporting machinery & RAG are very useful.

You can auto- lint and test code before you set eyes on it, then re-run the prompt with either more context or an altered prompt. With local models there are options like steering vectors, fine-tuning, and constrained decoding as well.

There's also evidence that multiple models of different lineages, when their outputs are rated and you take the best one at each input step, can surpass the performance of better models. So if one model knows something the others don't you can automatically fail over to the one that can actually handle the problem, and typically once the knowledge is in the chat the other models will pick it up.

Not saying we have the solution to your specific problem in any readily available software, but that there are approaches specific to your problem that go beyond current methods.

popinman322··on Ontario family doctor says new AI notetaking saved her job
Tangent here: really? I've found base Whisper has concerning error rates for non-US English accents; I imagine the same is true for other languages with a large regional mode to the source dataset.

Whisper + an LLM can recover some of the gaps by filling in contextually plausible bits, but then it's not a transcript and may contain hallucinations.

There are alternatives that share Whisper internal states with an LLM to improve ASR, as well as approaches that sample N-best hypotheses from Whisper and fine-tune an LLM to distill the hypotheses into a single output. Haven't looked too much into these yet given how expensive each component is to run independently.

popinman322··on Iterative reasoning preference optimization
Also, similar to Orca-Math but without a teacher model. They also followed an iterative DPO/KTO scheme, but with no length normalized NLL loss term.
popinman322··on Haystack DB – 10x faster than FAISS with binary embeddings by default
I remember stumbling upon an early discussion about this [0] a bit ago in the EleutherAI discord when searching for discussion about a paper; I'm glad to see it's turned into something public.

[0]: https://discord.com/channels/729741769192767510/730095596861...

popinman322··on Show HN: Finetune Llama-3 2x faster in a Colab notebook
Any news on when Unsloth's parallel full tuning will be available?
popinman322··on Using LLMs to Generate Fuzzers
You could likely also combine the LLM with a coverage tool to provide additional guidance when regenerating the fuzzer: "Your fuzzer missed lines XX-YY in the code. Explain why you think the fuzzer missed those lines, describe inputs that might reach those lines in the code, and then update the fuzzer code to match your observations."

This approach could likely also be combined with RL; the code coverage provides a decent reward signal.

popinman322··on Our next-generation model: Gemini 1.5
vs RAG: RAG is good for searching across >billions of tokens and providing up-to-date information to a static model. Even with huge context lengths it's a good idea to submit high quality inputs to prevent the model from going off on tangents, getting stuck on contradictory information, etc..

vs fine tuning: smaller, fine-tuned models can perform better than huge models in a decent number of tasks. Not strictly fine-tuning, but for throughput limited tasks it'll likely still be better to prune a 70B model down to 2B, keeping only the components you need for accurate inference.

I can see this model being good for taking huge inputs and compressing them down for smaller models to use.

popinman322··on Hi everyone yes, I left OpenAI yesterday
For the tasks my group is considering, even a 7B model is adequate.

Sufficiently accurate responses can be fed into other systems downstream and cleaned up. Even code responses can benefit from this by restricting output tokens using the grammar of the target language, or iterating until the code compiles successfully.

And for a decent number of LLM-enabled use cases the functionality unlocked by these models is novel. When you're going from 0 to 1 people will just be amazed that the product exists.

popinman322··on StreamingLLM: tiny tweak to KV LRU improves long conversations
Previous discussion, on a link to the implementation: https://news.ycombinator.com/item?id=37740932
Page 1 of 6Next →