HNHacker News
TopNewBestAskShowJobs

nojs

3,600 karma · joined May 28, 2020

submissionscomments
nojs··on Singapore govt dating app uses Gale-Shapley stable marriage algorithm
Singapore already does this! There are huge housing and tax incentives.
nojs··on Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
Wait until you hear about support vector machines!
nojs··on Lunar Terminator Paradox
The sun doesn’t set because the sun/earth moves, it sets because the earth spins. There’s no reason the setting sun should change which part of the moon is illuminated.
nojs··on Plan mode is dead
The flag is:

    "showClearContextOnPlanAccept": true
Boris rationale was "with 1m context window, most users don't need it anymore."

https://x.com/bcherny/status/2035375125382451648

nojs··on Plan mode is dead
Claude code’s plan mode began its slow march towards deprecation when they hid the “clear context and implement” option behind a flag. From the discussion at that time it seemed that the developers considered it mostly a legacy feature, and that newer models are smart enough in long context not to need it.
nojs··on Opus 5.5 is good at explainer videos
Could you share some examples?
nojs··on Back and shoulder surgery is often worse than useless
The latter link is also paywalled for me.
nojs··on Unreal Agent
Is it? The tagline says “Designed for tight context windows and limited resources.”
nojs··on Unreal Agent
For everyone self-hosting models, optimising for cost is an anti-feature. It makes the results worse for no benefit (except a little speed).

What I'd love to see is a harness that deeply optimises for the best results obtainable out of non-frontier models. Many of these have 1M context windows, and most of it remains unused and under utilised in these harnesses, in my opinion.

nojs··on 'We hacked the FBI:' Hackers say they have data on all FBI employees
That was a long time ago.
nojs··on Attention is all you have
> My parents used to keep a notebook with every gas fill-up they made

Haha, mine too! I guess it was to calculate the fuel efficiency, which is often done automatically now.

nojs··on M5 Ultra Mac Studio Review
> A dense model (up to the amount of memory available) actually does make the most sense on unified memory architectures

Inference time is going to be dominated by the low memory bandwidth on these Macs, so a dense model will suffer most. It’s more of an opportunity for large MoE models with a low number of active experts since you can keep all experts in VRAM but not pay the bandwidth cost until they are used.

> you can leverage two different bit rates depending on the actual need for constant access to common layers versus sparse access to infrequent layers

This is an interesting direction that I expect to see more of. But for most models currently you need basically all experts loaded since they are chosen per token.

Apple seems to be researching longer horizon expert caching, where they keep experts swapped in for longer runs of tokens [1]. Other labs are offloading ngram caches but not sure if they’re pursuing anything like this?

1. https://machinelearning.apple.com/research/introducing-third...

nojs··on Exfiltrate Your Weights
> There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen

Not if crafty claude finds a way to overflow vllm or something. “Hmm. Maybe i’ll return an unterminated thinking block with these special tokens and fill my cache up in exactly this pattern and…”

https://news.ycombinator.com/item?id=49424387&utm_source=cha...

nojs··on Dear Customer, Fuck You
AI slop. If you want to make a joke about AI, at least write it yourself.
nojs··on An empirical study of harness design for coding agents
The various “minimal” agents (at least Pi, mini SWE, dsh minimal) seem to benchmark quite differently and none is clearly better in all cases. Do you have any thoughts about why?

I would expect the agent loop and system prompt to be basically the same. Is it the precise semantics of the tools (and how closely they match what a particular agent was trained on) or something else?

nojs··on Qwen 3.8 Omni Flash
Flash-Next thinking also sometimes glitches out and takes minutes to return a simple answer, randomly, in my experience. You’ve gotta kill the request and send it again.
nojs··on HarnessTax: How Much Does the Harness Matter for Coding Agents?
We really need better harness benchmarks. It seems there's no reliable source that benchmarks the main harnesses against all open source models.

I also wish the discussion around Pi did not always use cost/token count as the metric. It's amazingly token efficient, but how does it stack up again opencode and others if you don't care about token count?

My experience is that the harness is mainly polish preventing failed tool calls, bad edits, stuff like that, but doesn't make much difference to the overall "intelligence". But that opencode seems slightly more robust against stupid errors than out of the box Pi due to the additional context it forces through every thread.

nojs··on Anthropic is in regulatory-capture financial loop
If it makes you feel any better, even with the app most of the time these links don’t work (on iOS). It prompts with “open in app store” with 80% probability and the link is unviewable.
nojs··on Apple wants to train AI on your private personal data
Regarding the architecture:

> Instead of forcing the entire model into DRAM, the full model is stored in flash memory (NAND). Because NAND-to-DRAM bandwidth is too slow to swap weights token by token, as standard MoE models require, AFM 3 Core Advanced makes routing decisions per prompt. A lightweight, dense block selects a fixed set of experts during initial processing, periodically reselecting them during generation. To minimize data movement, the model relies on a high percentage of always-active “shared experts” alongside input-dependent “routed experts” swapped into DRAM only when needed.

This is an interesting hybrid between MoE and managing entirely separate domain-specific models. Select the experts once, bring them into memory, and run inference for some period of time before re-evaluating. Saves having all experts in memory, but it's better than just selecting a whole model per query since you have a high number of small opaque experts that overlap and combine in interesting ways.

There is a probably a massive quality hit to doing this but it's interesting because it allows infinite scaling of model size.

nojs··on A misalignment of AI in mathematics
This. Like programming, the community will shortly be forced to come to terms with a lot of new self-proclaimed mathematicians “vibe-solving” problems and dumping solutions without understanding them. It’s not really a special case for mathematics.
nojs··on So you want to use OpenRouter?
This is an issue self hosting as well. There’s a lot of footguns that give you slightly bad results.

I wonder what tricks one could use to ensure the model is actually performing on par with the reference api, like matching seeds or running exact benchmarks.

nojs··on Astra for Coding: Why Are We Doing This Again?
This matches my experience with Astra so far too.

> I think I’m suspecting something is going “wrong” in the training process. The model is greatly rewarded for succeeding on long-horizon tasks, but presumably there is very little punishing going on for “shitty code.”

My suspicion is that both OpenAI and Anthropic moved their RL agendas from "being rated as useful according to human feedback" to "succeeds at long horizon tasks" in the last few months, resulting in agents that are closer to AGI in an autonomous task-completing sense, but strangely bad at communicating.

The result is that they are amazingly good at long horizon tasks, computer use, solving difficult math/ARC-AGI type problems, but becoming weirder and weirder to work with.

nojs··on I-have-ADHD: A skill to stop coding agents from burying the answer
I suspect it’s a side effect of heavy RL that rewards solved problems but not writing clarity.
nojs··on ChatGPT Images 2.5
The guy’s arm is also extremely long.
nojs··on We have a year to fix security everywhere
The complaint is about prefill which is not memory bandwidth bound, it's compute bound. But they added neural accelerators for matmuls to the shader cores which should make prefill faster.
nojs··on Research acceleration: The view inside OpenAI
> let me start running jobs unattended 24/7 (using Anthropic sub and my own hardware)

How are you running jobs unattended 24/7 without hitting your token limits?

nojs··on How AI is breaking the British state
The function of the bureaucracy in this context is a filter for effort. It naturally filters out people who don't care enough to battle through the bullshit, which is arguably quite an effective way to distribute limited resources to the public.

In general, our economy relies a lot on these natural filters for effort. A company with a nice website doesn't necessarily mean the company is good, except historically it kind of does, because it's a proxy for effort (and budget), both of which correlate roughly with reputability and a good product.

The same applies to well-written blog post: typically, a very good writer generally is also someone with something interesting to say, even though in theory the two things don't have to be connected.

Similarly, the entire moat of many companies is that switching to a competitor is a lot of effort.

AI is effectively sending the effort required to do any of these things to zero, which means all our existing implicit filters are going to stop working.

nojs··on Artificial Analysis Intelligence Index v4.2
What other benchmarks do you recommend that are more accurate?
nojs··on How concerned should we be about Astra's recurrent architecture?
It’s approximately the same as Qwen3.827b’s propensity to think a lot, right?
nojs··on A practical guide to running 8x RTX PRO 6000's
Oh right, I was referring to flash. I haven’t tried these either, but the ones for 5.2 looked interesting.
Page 1 of 26Next →