HNHacker News
TopNewBestAskShowJobs

qeternity

8,814 karma · joined May 4, 2016

submissionscomments
qeternity··on Ember-1
This is undoubtedly great. But most of the inference cost today for dominant use cases (agentic coding) are in the prefill, not the decode. This is one of the reasons that DeepSeek is so aggressively optimizing prefill and caching.
qeternity··on Show HN: Trader News – Hacker News for Finance
Hugged to death?
qeternity··on OpenJev
Not sure why you think everyone has moved on from this. Valid JSON is only one aspect. It almost certainly has far more uptake today than ca. Sonnet 3.7. Grammar engines can enforce a lot of other things too. Models are actually really good at emitting valid JSON, but complex schemas effectively require it.
qeternity··on OpenJev
Alternative is to fail the validation and regenerate the request, but I have no idea what this person is talking about. It is absolutely not the case that everyone has moved on from structured generation. Absurd claim.
qeternity··on DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
Most people are not even referring to CUDA batching nuances.

They think that sampling is an inherent part of Transformers.

Even on this site, it is regurgitated with confidence.

qeternity··on Discovery of a new OpenAI agent message board
It’s only good for them in the sense that it is allowing them additional time and compute to cheat a reward signal.

It is not good for them in the sense that they will short circuit the RL path that actually improves general capability.

qeternity··on Apple caught off guard by AI demand for Mac Mini and Mac Studio
Has nothing to do with being headless. That's just a natural outcome.

If someone wanted to use an 8x B300 as their daily driver...go ahead.

It would still be the best way to serve a given model.

qeternity··on A practical guide to running 8x RTX PRO 6000's
*TB/s
qeternity··on GLM-5.3-Flash
https://typebulb.com/u/lab/you-re-relatively-right/full
qeternity··on Apple caught off guard by AI demand for Mac Mini and Mac Studio
> Edit: I'm thinking of a headless Mac mini, if you meant running it on the same machine you're using of course you'll need more memory, but LLMs are best served from a headless server so that's what I'd recommend.

What? LLMs are best served from a massive PD disaggregated cluster of B300s connected via NVLink.

If you're running LLMs on a Mac Mini, it's because you want to run local, not because it's the best setup.

qeternity··on Apple caught off guard by AI demand for Mac Mini and Mac Studio
You're not entitled to this stuff any more than they are.

What planet do you live on?

qeternity··on GLM-5.3-Flash
Trolling. GLM is heavily distilled from Gemini.
qeternity··on Coconut oil jet fuel matches kerosene's efficiency in engine tests
Is it not a bad thing to encourage poor people to continue doing the thing that has apparently made them poor?

It's also very not true that farmers are poor in the US (median household income about 30% greater than general population). This mindset is like a century old.

qeternity··on AI boosted homework scores, then exam scores dropped: study
College stopped being about getting an education a long time ago.

It's just a low pass filter for the job market. Students aren't actually interested in learning anything, they have been condition for a long time to jump through the hoops.

It would be much better if we actually returned to exam and performance based evaluation of students, and stopped giving everyone As just for trying hard. But parents look at college fees as buying their kid a job opportunity, and so grade inflation has basically ruined the SNR of academic performance for all but the lowest performers.

qeternity··on Felony charges for citizen deleting phone data at US Border
It's the other way around: conviction rate includes plea bargains.

Convictions at trial are much lower.

qeternity··on Felony charges for citizen deleting phone data at US Border
You are misunderstanding the conviction rate. Only ~2% of cases actually go to trial, and at trial there is ~80% conviction rate.

There are ~4x as many cases dismissed by judges before getting to trial. And the vast majority (90%) of defendants enter into plea bargains.

Prosecutors only bring charges when they feel they have a strong case. There are many many cases which are never pursued because of this, and people also get upset about that.

Many countries do not have a plea bargain system the way the US does. And if you look at conviction rates at trial, they are smack in line with much of e.g. Western Europe.

qeternity··on What Happens When the Cost of Intelligence Drops 100x
I presume you're a Windows user then.

When the M1 was released, I couldn't believe how fast it was.

qeternity··on Cerebras CS-4
Ultimately the frontier labs are competitors of Nvidia. There is a fixed amount that the market will pay for tokens. If the frontier lab model premium collapses due open weights models, Nvidia can capture a greater share of aggregate spend.
qeternity··on U.S. Debt Hits $40T as America's Borrowing Binge Continues
It is still people, not money, who votes for them over and over again…
qeternity··on Qwen 3.8 27B is excellent, but it defaults to overthinking things
Personal hardware is only sufficient today for some tasks. As data center power efficiency, and large sparse MoE task efficiency increase, personal computing will continue to lose out.
qeternity··on Qwen 3.8 27B is excellent, but it defaults to overthinking things
Computer use will blow through tokens because it's doing image capture for everything.

You may have better and more reproducible results using browser controls that aren't image based, or writing tools that completely sidestep browser use.

qeternity··on Qwen 3.8 27B is excellent, but it defaults to overthinking things
Jevons paradox: large purpose-fit data centers increase efficiency such that you can use AI in more places, and use more tokens for those tasks.

The future is not a single chat bot session of bs=1. The future is many agents performing many tasks in parallel for a single user. Large GPU clusters will always have the edge in efficiency.

qeternity··on Qwen 3.8 27B is excellent, but it defaults to overthinking things
The real breakthrough is going to be thinking in latent space.
qeternity··on Auto-research with codex: How I achieved a 232x Faster Kernel
You say this like that isn’t how the vast majority of problems are solved…
qeternity··on Qwen 3.8 27B
Has nothing to do with Chinese.

Frontier labs have already been doing this for a while, verified in smuggled traces from OAT/Ant.

Simply a way to reduce tokens.

qeternity··on Qwen 3.8 27B
35A3 might be more comparable to 10 dense.

27 dense is far more capable than 35A3.

qeternity··on Gemini 3.7 Flash
> more than half its price

Less than half its price.

More than 50% discount.

qeternity··on Qwen3.8-Max: A New Bar for Coding and Cowork
LLMs are, in theory, deterministic. Sampling is not intrinsic to LLMs.

Greedy decoding a single batch in most libraries will give you mostly deterministic outputs. Higher batch sizes can increase variance.

But all of this is down to CUDA and/or kernel implementation issues.

qeternity··on Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba
I don't follow this. Clearly all of the frontier labs are doing these things.

When OAI released gpt-oss it was released as an mxfp4 checkpoint.

OAI, Ant, et al are also obviously employing QAT.

qeternity··on Kimi-K3 Releases on HuggingFace 7/27
I cannot be sure what the likes of Cursor have done, but I think it's incredibly unlikely that they have trained a QLoRA for Composer.

It's almost certainly full parameter post training of the original model weights.

Page 1 of 34Next →