Matters more for 1-bit and 2-bit models.
320 karma · joined January 21, 2023
Talk to me about computational chemistry, bioinformatics, and molecular biology!
Matters more for 1-bit and 2-bit models.
There's literature on this actually, if you want to get that low the model really should have quantization as a pretraining target otherwise larger models that are quantized after collapse.
In fact, to do this effectively, you need somewhere near 50x chinchilla to do QAT on super tiny targets like 1-bit or even 2-bit. These large models already need an astronomical amount of training data that doing proper QAT that small isn't really feasible unless there are breakthroughs in the architecture.
4-bit models and smaller _can_ perform well but not through just naively quantizing the existing weights of a large model.
It's not that simple, you have to remember at the end of the day these things are just doing next token prediction. If you don't give it the proper tokens to attend to, then your outputs won't be satisfactory.
You can get stellar outputs from LLMs, but it really is a function of how well you manage your input tokens.
"on the fly embedding" and "Sparse + dense reranking" don't really make sense how they're presented and smell like they came from a long claude-driven conversation after multiple cycles of these hybrid compromises across many turns.
Looks interesting, but I'd like to see some reviews on the content first before buying.
A lot of people in the comments do have a software engineering background. People at different skill levels in different backgrounds are going to be using these tools in different ways, and that's going to heavily impact their experiences with these models.
Sure, there are differences between Fable and Sol. But I've even seen people on here saying that they're getting better mileage out of Qwen models they're self hosting.
I think the driver is just as important than the car, when it comes to this sort of stuff.
I feel like to get to Terry's level you need a combination of passion and aptitude for the subject. People that don't want to learn about a topic will always look for shortcuts, which I think represents the vast majority of people. Terry Tao is quite exceptional, and I think exceptional people will still exist even when the "easy" button is bigger than it's ever been.
I don't think chatgpt could have come to this on its own without the amount of steering he did, which just validates the idea that AI is not a replacement for human expertise but an amplifier.
From a data structure and file ergonomics perspective, think of it as similar to Unity or UE4 for drug design. We have a huge variety of assets to manage alongside their relationships to each other, and the project files are local on the user's machine (with a collaboration / sync over the network between scientists working on the same project, hence where something like this would come in for us).
Many of those files are fine with a winning side strategy, but some of them might not be that clean. Take a protein structure defined by an `mmcif` file for example, if we clean the file by removing hydrogen atoms and another scientist repairs a side chain on that same file then we'd need a way to reconcile those differences.
On the agent side, our agents will generate small python scripts that manipulate the proteins, then cache and re-use those scripts as tools when possible. So preserving those scripts alongside the mutated asset and conversation history is something we've been working on.
I have a use case that could use this if it supports handling branching and merging file systems.
It's really hard to do surgical changes with an AI agent, and it's even harder to review those changes. Even if I'm reviewing the specs and the code, the cognitive load on reviews feel like they've ballooned from what used to be a few hours to now taking me days to review these PRs.
Group 1: A thinly veiled straw man that buckets everyone I disagree with, along with an attempt to appear as if I'm being unbiased
Group 2: The group I put myself in and provide better arguments for why this perspective is correct.
Vague motte and bailey statement that gives me plausible deniability when someone criticizes my analysis.
On top of that, we don't have a clear understanding on how certain positions (conformations) of a structure affect underlying biological mechanisms.
Yes, these models can predict surprisingly accurate structures and sequences. Do we know if these outputs are biologically useful? Not quite.
This technology is amazing, don't get me wrong, but to the average person they might see this and wonder why we can't go full futurism and solve every pathology with models like these.
We've come a long way, but there's still a very very long way to go.
There's a huge trade off between resolution and scale that makes it hard to determine things like complex molecular dynamics and how those dynamics influence the broader functions of the cell.
That said, excited for more images like this! More data at that scale is always a good thing for researchers.
Your workflow is probably closer to what most SWEs are actually doing.
That's good to hear, I might have jumped a little too quickly in my opinion. It's a bit of a Pavlovian response at this point seeing a product I very much love embrace a giant chat window as a UX redesign haha.
I would love to see more features on the roadmap that are more aligned with users like us that really embrace the Cursor 2 style with the code itself being the focal point. I'm sure there's a lot you can do there to help preserve code mental models when working with agents that don't hide the code behind a chat interface.
These models are infinitely more effective when piloted by a seasoned software engineer and that will always be the case so long as these models require some level of prompting to function.
Better prompts come from more knowledgeable users, and I don't think we can just make a better model to change that.
The idea we're going to completely replace software engineers with agents has always been delusional, so anchoring their roadmap to that future just seems silly from a product design perspective.
It's just frustrating Cursor had a good attitude towards AI coding agents then is seemingly abandoning that for what's likely a play to appease investors who are drunk on AI psychosis.
Edit: This comment might have come off more callous than I intended. I just really love Cursor as a product and don't want to see it get eaten by the "AI is going to replace everything!" crowd.
I feel like this design direction is leaning more towards a chat interface as a first class citizen and the code itself as a secondary concern.
I really don't like that.
Even when I'm using AI agents to write code, I still find myself spending most of my time reading and reasoning about code. Showing me little snippets of my repo in a chat window and changes made by the agent in a PR type visual does not help with this. If anything, it makes it more confusing to keep the context of the code in my head.
It's why I use Cursor over Claude Code, I still want to _code_ not just vibe my way through tickets.