HNHacker News
TopNewBestAskShowJobs

kadushka

454 karma · joined May 9, 2017

submissionscomments
kadushka··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
They use a meaningful portion of compute.
kadushka··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Just as $500 makes no sense now

Why? You can use in codex, right?

kadushka··on Claude Opus 5.5
We are being productive here!
kadushka··on Claude Opus 5.5
This makes perfect sense. There are no real improvements anymore (just benchmaxxing), and they explain it by "pacing the frontier".
kadushka··on Frontier AI on Your Own Hardware
Guess why I'm posting on HN right now instead of working?

because fable/astra is working for you? :)

kadushka··on Frontier AI on Your Own Hardware
thanks, fixed
kadushka··on Frontier AI on Your Own Hardware
CMU is #2 in the country for CS program. If those students don't think they can find a job, it's bleak.
kadushka··on Breaking the 1.58-bit Barrier for Ternary LLMs
the average information content of transformer LLMs at about 3-4 bits per parameter

The problem is that 4-bit block-wise quantization does not guarantee preserving 4 bits of useful information per parameter - not even on average. It simply assigns one of 16 quantization levels to each weight, with the whole block sharing the same scale/range.

How efficiently those 16 levels preserve the model’s information depends on the weight distribution, block size, range/clipping strategy, outliers, and which weights are actually important. Some weights may be represented almost exactly, while others lose much of their useful information.

A simple example is an outlier: if you choose the range to preserve a very large weight, much of the 16-level dynamic range is spent on that outlier, leaving coarse resolution for all the smaller weights in the block. So 4 bits of storage does not imply 4 bits of useful information preserved. Yes, QAT helps, but usually at the cost of learning efficiency. It takes longer to train a model to the same quality when using less precision, and sometimes we simply cannot get to the same quality level with not enough precision in the right places.

Another problem in quantization is that we don't really know which weights are sensitive - we can compute various sensitivity metrics, and some of these metrics will correlate with accuracy on some benchmarks, but not on others.

Another complementary option is, if the model is fast enough, we should be able to push up correctness by self-consistency voting at close to T=1. Smart/fast Zero-shot classifiers like the recent Jev could help with aggregation across answers too, extending applicability.

I'm not convinced by this argument - if such a method improves accuracy of a degraded quantized model, then it could in theory also help non-degraded full precision model. And if so, then we are back to square one, because this composite model will then get degraded due to quantization (baseline has improved!)

We do know one thing - increasing the size of the model usually makes it more robust to quantization. If going from 8 bits to 2 bits speeds things up by a factor of, say, 4x, then if we double the size of the model, we might still end up with an overall speedup. Finding this balance might become a hot area of research.

kadushka··on Breaking the 1.58-bit Barrier for Ternary LLMs
That's what I meant - we are currently use fp4 formats for training, and we cannot quite get away with that, despite dynamic quant and small block size - we still have to use quite a bit of higher precision (fp8 or even fp16) in various model components.
kadushka··on Breaking the 1.58-bit Barrier for Ternary LLMs
By “work out” you mean no accuracy degradation? That’s a big ask - currently we can barely quantize to dynamic fp4 with small block size - still not completely lossless on all benchmarks.
kadushka··on My friends all hate AI; I just joined an AI startup
Everyone I know loves AI. It is actively making their lives better. All kinds of people, all kind of ages - universally love it and use it every day for all kind of stuff (shopping, health, learning, home improvement, etc, etc).
kadushka··on How to put 170 atoms in an atom
The article is well-written, and is very clear, thank you!
kadushka··on Qwen3.8 Max now ranked as the best overall model by agentic index
Cost is absolutely an issue here - my time is worth approximately $1000/day, so if a slightly worse model wastes one more hour of my time a day than the best model, it costs the company >$2k/mo. Fortunately my employer understands this well and encourages me to use the best models as much as I can.
kadushka··on Qwen3.8 Max now ranked as the best overall model by agentic index
Opus 5 and 5.6 Sol are definitely not smart enough to do my job. They require constant supervision. So why would I want to switch to even worse model? Even if it's just slightly worse?
kadushka··on Tennessee man jailed 37 days for Trump meme wins settlement after lawsuit
How would this work? Where do the money for their pension fund come from? Would taking money from it result in them receiving smaller pensions?
kadushka··on Advanced Quantization Algorithm for LLMs
Most quant papers I've seen usually report non-trivial degradation on standard benchmarks, like 1-10% degradation (compared to FP16/BF16). Especially when using 4 bits or lower. For example, I just opened a random paper: https://arxiv.org/pdf/2410.09426 see Table 1.

p.s. dense vs MoE: both are being released because they offer different trade-offs: at the same level of quality, MoE will use less compute, but more memory.

kadushka··on Claude Opus 4.7
models are getting weirdly good at hacking while still sort of sucking at a bunch of economically valuable tasks

like most human hackers

kadushka··on Robot Police Dogs Powered by AI Take over Atlanta's Streets
taser/pepper spray within 5 years, firearms within 10.
kadushka··on Issue: Claude Code is unusable for complex engineering tasks with Feb updates
We are surrounded by black boxes we depend on - have been for at least a century.
kadushka··on Oracle slashes 30k jobs
Probably had a lot of meetings
kadushka··on Ask HN: Release Path for 'Transformers Alternatives'?
1. If you have good results on sufficiently large models (check latest papers re: which benchmarks are still relevant), post them on Github, along will detailed instructions how to reproduce.

2. Post the link to the GH repo in "Show HN" section.

3. If results are solid, write up a paper and upload it to arxiv. Next step would be try to publish in an ML conference.

p.s. To increase your chances of anyone actually clicking on your GH link, use good old Pytorch.

kadushka··on Show HN: Most GPU Upgrades Aren't Worth It, I Built a Calculator to Prove It
I'm not interested in gaming, but if you had a version for AI, I'd be using it!
kadushka··on Yann LeCun raises $1B to build AI that understands the physical world
Diffusion models are not autoregressive but have the same limitations
kadushka··on Yann LeCun raises $1B to build AI that understands the physical world
Imagine that we made an LLM out of all dolphin songs ever recorded, would such LLM ever reach human level intelligence?

It could potentially reach super-dolphin level intelligence

kadushka··on Claude Code is being dumbed down?
I'm an employee, and my boss loves me because I deliver things he wants quickly and reliably - because I use AI tools. Guess who he will keep in the next round of layoffs?
kadushka··on Claude Code is being dumbed down?
I'm serious - the productivity boost I'm getting from using AI models is so significant, that it's absolutely worth paying even 2k/month. It saves me a lot of time, and enables me to deliver new features much faster (making me look better for my employer) - both of which would justify spending a small fraction of my own money. I don't have to, because my employer pays for it, but as I said, if I had to, I would pay.
kadushka··on Claude Code is being dumbed down?
I would probably pay $2000 a month if I had to - it's a small fraction of my salary, and the productivity boost is worth it.
kadushka··on Why vampires live forever
Is there any evidence?
kadushka··on Vitamin D and Omega-3 have a larger effect on depression than antidepressants
Sure, could be just lucky. But if there are several successful small studies, and several unsuccessful large ones (no idea if this is the case here), we should probably look for a better explanation.
kadushka··on Vitamin D and Omega-3 have a larger effect on depression than antidepressants
use a diverse population

If that's the case, we should question whether different homogeneous population groups respond differently to the substance under test. After all, we don't want to know the "average temperature of patients in a hospital", do we?

Page 1 of 12Next →