HNHacker News
TopNewBestAskShowJobs

seamossfet

320 karma · joined January 21, 2023

Working on CAD for drug discovery.

Talk to me about computational chemistry, bioinformatics, and molecular biology!

submissionscomments
seamossfet··on Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Yeah, there's a tradeoff. pre-training with QAT needs more data and a ton of hardware, so it's easier to just take an existing open weight model and quantize it. This works, but it'll underperform a model that had quantization as a training target.

Matters more for 1-bit and 2-bit models.

seamossfet··on Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
>I'm a little skeptical of going below 4-bit quants due to the potential for significant degradation in quality

There's literature on this actually, if you want to get that low the model really should have quantization as a pretraining target otherwise larger models that are quantized after collapse.

In fact, to do this effectively, you need somewhere near 50x chinchilla to do QAT on super tiny targets like 1-bit or even 2-bit. These large models already need an astronomical amount of training data that doing proper QAT that small isn't really feasible unless there are breakthroughs in the architecture.

4-bit models and smaller _can_ perform well but not through just naively quantizing the existing weights of a large model.

seamossfet··on An agent used DNS to reach an external chatbot
How else would they get their marketing stories unless the agents can "break out" of containment?
seamossfet··on How to keep enjoying programming in a world of LLMs
>You literally ask the agent to do something and it does it

It's not that simple, you have to remember at the end of the day these things are just doing next token prediction. If you don't give it the proper tokens to attend to, then your outputs won't be satisfactory.

You can get stellar outputs from LLMs, but it really is a function of how well you manage your input tokens.

seamossfet··on Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
If you want to do a 1-bit model you have to QAT at pre-training with way more data than chinchilla to compensate for the cliffs (like 50x). Quantization on an existing pre-trained model will almost always collapse at 1-bit
seamossfet··on WebLLM: high-performance in-browser LLM inference engine
We use the ONNX runtime for small models in the browser https://github.com/microsoft/onnxruntime
seamossfet··on Emacs 31: An unofficial guide to Markdown-ts-mode
just switched to doom emacs this month, absolutely loving it tbh
seamossfet··on RAG Is Simpler Than You Think
yeah, but I mean even prose specific claude-isms aside; the information itself is a weird patchwork of concepts
seamossfet··on RAG Is Simpler Than You Think
I notice a lot of these AI written articles share this pattern where they'll present idea 1, then idea 2, and finally idea 3 which is some amalgamation of idea 1 and 2. Claude especially will present hybrid options and compromises to avoid having to make a choice then framing the hybrid option as the "best of both worlds" when they're borderline nonsensical.

"on the fly embedding" and "Sparse + dense reranking" don't really make sense how they're presented and smell like they came from a long claude-driven conversation after multiple cycles of these hybrid compromises across many turns.

seamossfet··on Embedded AI
Has anyone read this yet? I'm extremely skeptical of any book published after ~2024 just because of how much AI slop is in physical books these days.

Looks interesting, but I'd like to see some reviews on the content first before buying.

seamossfet··on A week of using Codex more than Claude
I think a part of this is that people tend to undervalue their own skills and expertise when talking about these anecdotes.

A lot of people in the comments do have a software engineering background. People at different skill levels in different backgrounds are going to be using these tools in different ways, and that's going to heavily impact their experiences with these models.

Sure, there are differences between Fable and Sol. But I've even seen people on here saying that they're getting better mileage out of Qwen models they're self hosting.

I think the driver is just as important than the car, when it comes to this sort of stuff.

seamossfet··on Terence Tao's ChatGPT conversation about the Jacobian Conjecture counterexample
I don't know if I agree with the premise that having access to AI results in dulling human intellect.

I feel like to get to Terry's level you need a combination of passion and aptitude for the subject. People that don't want to learn about a topic will always look for shortcuts, which I think represents the vast majority of people. Terry Tao is quite exceptional, and I think exceptional people will still exist even when the "easy" button is bigger than it's ever been.

seamossfet··on Terence Tao's ChatGPT conversation about the Jacobian Conjecture counterexample
The big take away for is the fact that the ONLY reason why chatgpt was able to get to this counterexample was because of the knowledge of the person driving the conversation.

I don't think chatgpt could have come to this on its own without the amount of steering he did, which just validates the idea that AI is not a replacement for human expertise but an amplifier.

seamossfet··on DNSGlobe – Rust TUI to watch DNS propagate around the world
Nothing wrong with vibecoding a toy project.
seamossfet··on Show HN: Tilde.run – Agent sandbox with a transactional, versioned filesystem
We're building a CAD for drug design, we often have to handle large and highly varied file formats. Protein structures, compounds, python scripts, lab notebook entries, instrumentation data, etc.

From a data structure and file ergonomics perspective, think of it as similar to Unity or UE4 for drug design. We have a huge variety of assets to manage alongside their relationships to each other, and the project files are local on the user's machine (with a collaboration / sync over the network between scientists working on the same project, hence where something like this would come in for us).

Many of those files are fine with a winning side strategy, but some of them might not be that clean. Take a protein structure defined by an `mmcif` file for example, if we clean the file by removing hydrogen atoms and another scientist repairs a side chain on that same file then we'd need a way to reconcile those differences.

On the agent side, our agents will generate small python scripts that manipulate the proteins, then cache and re-use those scripts as tools when possible. So preserving those scripts alongside the mutated asset and conversation history is something we've been working on.

seamossfet··on Show HN: Tilde.run – Agent Sandbox with a Transactional, Versioned Filesystem
Does this provide gitflow to handle conflicts from multiple agents touching the same file system or is it purely for single-branch sequential iterations on the filesystem?

I have a use case that could use this if it supports handling branching and merging file systems.

seamossfet··on Over-editing refers to a model modifying code beyond what is necessary
People often use "non-deterministic" to imply "unreliable" when that's not always the case. You can have an LLM that's clearly non-deterministic but also reliable for tasks.
seamossfet··on The M×N problem of tool calling and open-source models
Great article, but your site background had me trying to clean my laptop screen thinking I splashed coffee on it.
seamossfet··on The uncomfortable truth about vibe coding
The hardest part about using agents to code for me has always been working in teams. When you can cut through huge parts of the code with a chainsaw, how do you review multi-thousand line PRs?

It's really hard to do surgical changes with an AI agent, and it's even harder to review those changes. Even if I'm reviewing the specs and the code, the cognitive load on reviews feel like they've ballooned from what used to be a few hours to now taking me days to review these PRs.

seamossfet··on A compelling title that is cryptic enough to get you to take action on it
This is why I built [AI slop tool]. [Self promotion link to my vibe coded startup with no users]
seamossfet··on A compelling title that is cryptic enough to get you to take action on it
A false dichotomy that segments typical replies into one of two groups.

Group 1: A thinly veiled straw man that buckets everyone I disagree with, along with an attempt to appear as if I'm being unbiased

Group 2: The group I put myself in and provide better arguments for why this perspective is correct.

Vague motte and bailey statement that gives me plausible deniability when someone criticizes my analysis.

seamossfet··on Training mRNA Language Models Across 25 Species for $165
The problem with models like this is they're built on very little actual training data we can trace back to verifiable protein data. The protein data back, and other sources of training data for stuff like this, has a lot of broken structures in them and "creative liberties" taken to infer a structure from instrument data. It's a very complex process that leaves a lot for interpretation.

On top of that, we don't have a clear understanding on how certain positions (conformations) of a structure affect underlying biological mechanisms.

Yes, these models can predict surprisingly accurate structures and sequences. Do we know if these outputs are biologically useful? Not quite.

This technology is amazing, don't get me wrong, but to the average person they might see this and wonder why we can't go full futurism and solve every pathology with models like these.

We've come a long way, but there's still a very very long way to go.

seamossfet··on Scientists observe an immune signaling complex forming inside cells
This is awesome! The only limiter here is the resolution, I think this is fantastic for cellular level organelles but it doesn't quite get down to the same resolution something like x-ray diffraction does.

There's a huge trade off between resolution and scale that makes it hard to determine things like complex molecular dynamics and how those dynamics influence the broader functions of the cell.

That said, excited for more images like this! More data at that scale is always a good thing for researchers.

seamossfet··on Tell HN: Anthropic no longer allowing Claude Code subscriptions to use OpenClaw
Honestly, this is a good thing. OpenClaw as a concept was rather silly to run such a heavy model for. If you want something like OpenClaw to work you really need to figure out how to do it with an economical model.
seamossfet··on How to Write Unmaintainable Code (1999)
Honestly, it'd be really funny to try and make a CLAUDE.md file for slop maxxing.
seamossfet··on Cursor 3
I'm not convinced people who are doing real work on production applications with any sizable user base is writing code through only agents. There's no way to get acceptable code from these models without really knowing your code base well and basically doing all the systems thinking for the model.

Your workflow is probably closer to what most SWEs are actually doing.

seamossfet··on Cursor 3
Oh my god, this comment gave me flashbacks to when I was writing android apps in Eclipse + ADT
seamossfet··on Cursor 3
> We very much still believe this

That's good to hear, I might have jumped a little too quickly in my opinion. It's a bit of a Pavlovian response at this point seeing a product I very much love embrace a giant chat window as a UX redesign haha.

I would love to see more features on the roadmap that are more aligned with users like us that really embrace the Cursor 2 style with the code itself being the focal point. I'm sure there's a lot you can do there to help preserve code mental models when working with agents that don't hide the code behind a chat interface.

seamossfet··on Cursor 3
Yeah that's the disconnect though right? Even with the best frontier models, you need to do a lot of system design work, planning, and reviewing before you can let these models run.

These models are infinitely more effective when piloted by a seasoned software engineer and that will always be the case so long as these models require some level of prompting to function.

Better prompts come from more knowledgeable users, and I don't think we can just make a better model to change that.

The idea we're going to completely replace software engineers with agents has always been delusional, so anchoring their roadmap to that future just seems silly from a product design perspective.

It's just frustrating Cursor had a good attitude towards AI coding agents then is seemingly abandoning that for what's likely a play to appease investors who are drunk on AI psychosis.

Edit: This comment might have come off more callous than I intended. I just really love Cursor as a product and don't want to see it get eaten by the "AI is going to replace everything!" crowd.

seamossfet··on Cursor 3
Man, I wish they'd keep the old philosophy of letting the developer drive and the agent assist.

I feel like this design direction is leaning more towards a chat interface as a first class citizen and the code itself as a secondary concern.

I really don't like that.

Even when I'm using AI agents to write code, I still find myself spending most of my time reading and reasoning about code. Showing me little snippets of my repo in a chat window and changes made by the agent in a PR type visual does not help with this. If anything, it makes it more confusing to keep the context of the code in my head.

It's why I use Cursor over Claude Code, I still want to _code_ not just vibe my way through tickets.

Page 1 of 2Next →