HNHacker News
TopNewBestAskShowJobs

nbardy

819 karma · joined March 28, 2014

nbardy @ GitHub

Multimodal models

Language models for art

submissionscomments
nbardy··on GPT-6.1 Sol replaces GPT-6 Sol after just 7 days, with near-Astra intelligence
I think it's weirdly just a choice of deciding to cut releases.

We already know OpenAI has "bel" that is MUCH better than astra and is being used internally

nbardy··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Have you even tried the new models? Opus 5.5 is a clear leap. Go back a year ago and try them and tell me there is any sort of plateau.
nbardy··on Livenerf: Has Opus 5.5 been nerfed yet?
I’m pretty sure 99% of what people perceive is the old model training on the usage logs wherever they were stuck.

Step 1. Model can’t do something challenging Step 2. You try a bunch and fail Step 3. Anthropic trains on your usage data. Your current code base and current problem are now in domain Step 4. Model comes out and you’re shocked when it can tackle the thing you were stuck on Step 4. Codebase drifts significantly and you try new problems you thought were a similar level. Your code is less familiar and the problem doesn’t have a bunch of failure cases in the train set. Feels of it being worse on similar problems

nbardy··on Formalizing Fermat's Last Theorem
You can estimate the model size by looking at tokens per second and comparing to open source models
nbardy··on GPU World
One of the amazing things is that when every has one GPU, they will actually have 1k-10k agents at their disposal.

LLMs and KV caches have amazing performance characteristics with concurrent throughput. It scales very non linearly. So the token throughput within a batch scales WAY faster than the tokens per second of each user.

This is the reason the LLM providers have such crazy margins on their costs.

nbardy··on Nvidia agrees to acquire Hugging Face for $13B
They had incredible revenue growth the last few years and just broke 100M in revenue. I don't know what their internal spend was , but that was almost half of their recent round in ARR. Mostly likely they were profitable or on a clear trajectory to revenue growth. Huggingface hosts a lot of data and models, but mostly static cold storage is pretty cheap tbh.
nbardy··on I were 17, I'd learn how to build LLMs from scratch
This is a wildly incorrect and myopic view on the world.

Finetuning model is cheap and incredibly useful for deployment. You don't need to pre-train a frontier llm from scratch to make useful models.

There is tons of domains where you and fine-tune llms and deploy them for value in companies and for your own entrepreneurship ambitions. I have made this a big part of my career for the last few years and now I'm working on finetuning models for starting my own companies.

nbardy··on Grok 4.6
I think a lot of it is just time. The quality of a model is E * C

Where: E = Efficiency, and efficiency gains come from quality of data, quality of algorithms. C = Compute (Size of model, flops of train run)

So a better company can train a bigger and better model with less required compute which let's anthropic get there first. If another company does the same thing with a worse: model architecture, kernel, optimizer, etc... They will get there as well if they just run there train run with more flops for longer

Mythos was actually ready about 6 months ago. So if you have 6 months later or hardware setup and time to train you can get a lot done.

nbardy··on Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
Everyone keeps repeating this who doesn’t understand the underlying technology.

Small llms are still way more efficiently server on big GPUs.

Sharing server capacity takes advantage of the massive parallel throughput and sharing of memory bandwidth.

You are sharing the GPUs with thousands of concurrent users.

nbardy··on DeepMind's WeatherNext model achieves breakthrough forecasting cyclones
he got promoted to chief scientist and them being an expert at all modeling will really help compared to being llm only
nbardy··on “Code was never the hard part” is an insult to all programmers
yea, a lot of my prior work was in the "code is the easy part" I was a frontend engineer for years. And something like 90% of my job the code was not the hard part.

I loved writing GPU shaders or optimizing visualization performance, but most of the time it was wiring up netcode to UI elements that exist.

Ironically as I've moved into focusing on more GPU and kernel programming AI is now lapping me there anyway, however the impact of knowing what sort of algorithsm are state of the art in papers, what is causing memory bandwidth issues etc... does a lot to drive the machine.

nbardy··on Seedance 2.5
The faces look amazing. They can post train for identity preservation easy, that is just an adapter
nbardy··on Are AI Labs Pelicanmaxxing?
Then they’re legitimately getting better at svg which is a valuable tool.
nbardy··on America pays workers just 27% of what its wealth allows – the worst in the OECD
This is a weird number.

And reeks of the same sort of reallocation fallacy that makes people think the rich making too much makes them poor.

If we really just say 4xd everyone’s salary in America. Prices are gonna rapidly rise in everything.

Things don’t get materially better unless we build the material things we need. We need to come up with a way to build more houses. Lower the cost of healthcare. Not just increase everyone’s money supply.

nbardy··on GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]
These are true and it does make theses fields the first to fall, but also the hype comes form the fact that it can escape these conditions as well: - generalize to non verifiable domains (https://arxiv.org/abs/2507.17746)[This is on going work but has had steady progress in many angles of attack) - Visual Reasoning is alive and well in video diffusion and image models(see recent works around using diffusion model priors as world model features for physics reasoning) - Prior art exists for anything humans can do online?(Is this one even a hold back?)
nbardy··on GPT-5.6
We’re definitely going to need a lot of Gpu’s
nbardy··on Previewing GPT‑5.6 Sol: a next-generation model
Weird flex. There is cheaper better and faster models you could of moved to with an hour effort
nbardy··on An entire Herculaneum scroll has been read for the first time
The first bit was interesting and then you flipped right to generic cynicism.

They would be impressed with our technology even if it has downsides. Wisdom is knowing humans and technology and imperfect tools.

nbardy··on Petition against Meta's employee training data collection for ML models
Signing this sounds like a good way to get fired. Executive in corporations gets to make the decisions. Employment is at will, if you don’t like it you get to leave otherwise you’re not fulfilling your contract
nbardy··on SpaceX to buy Cursor for $60B
No, look a Composoer 2, it stands out starkly on its own in the pareto frontier on low cast and fast models.

Composer 2.5 was a huge leap with minimal compute from xAI.

They can compete with OpenAI and anthropic with xAI scale compute. They have a top notch model team and incredible training data and huge enterprise costumer contracts.

nbardy··on Claude Fable 5
It’s a bit misleading to say nothing special, as they are doing more than just increasing parameter count. Progress has been steady in all the sub components of training from data filtering and weighting to sparse attention, optimizers to up and down the stack various efficiency in training computing.

They’re using more compute, a bigger model and tons of training quality improvements to get more out of an equivalent model.

nbardy··on Anthropic, please ship an official Claude Desktop for Linux
This does feel like the perfect setup for Claude though.

Much easier to create a vm testing swarm of 100 disitributions with llms

nbardy··on Do transformers need three projections? Systematic study of QKV variants
This has been my thought for a long time. I think all that matters from attention is that there is crosswise comparison going on.

You need some amount of parallel compute and some amount of global comparison.

And the rest is basically a ways to parameters and scale.

(This is in theory, in practice you can get a lot of small % stability and efficiency improvements that really compound in algorithmic details of model architecture)

nbardy··on Anthropic's open-source framework for AI-powered vulnerability discovery
In general this is the way I see open source going.

We won't reuse open source libraries as libraries we import, but as design inspiration for the bespoke tools we make.

It's too cheap to make your own stuff and too expensive to be stuck with someone else primitives.

But grounding AI Coding in existing tools is incredibly powerful.

nbardy··on Claude Opus 4.8
Confidently yes. OpenAI for sure has been training larger models internally and distilling.

Pre-training scaling laws all support larger models being more cost effeceint to train then smaller models. And distillation is comparably cheap. So you can get the most juice by training the biggest model you can and distilling it.

nbardy··on Claude Opus 4.8
There is endless returns to frontier intelligence, just because most people can't make use of it doesn't mean someone can't make a ton of money off of it.

Most software engineers will just need cheap tokens.

But things like physics and drug discovery have no foreseeable upper bound.

nbardy··on Claude Opus 4.8
There is endless returns to frontier intelligence, just because most people can't make use of it doesn't mean someone can't make a ton of money off of it.

Most software engineers will just need cheap tokens.

But things like physics and drug discovery have no forseeable upper bound.

nbardy··on ClojureScript Gets Async/Await
Seems like wrapping async await functions with CSP was a better way to handle this . Clojure already had a nicer pattern for this
nbardy··on SpaceX says it has agreement to acquire Cursor for $60B
People keep saying this and they don't understand how businesses work.

Cursor has 1B in enterprise revenue. It doesn't matter if people can clone their product, those deals don't move slowly

nbardy··on Project Glasswing: Securing critical software for the AI era
There is step changes that actually merit this though. And a zero day machine IS one of those. It went from 4% zero day success rate to 85% on firefox.

Can you not see the significance of that?

Page 1 of 10Next →