HNHacker News
TopNewBestAskShowJobs

nullc

19,137 karma · joined December 24, 2010

gmaxwell

http://nt4tn.net/

submissionscomments
nullc··on “It works better in the app”
I take "It works better in the app" as confirmation that the app steals your data, tracks your location, etc. I wasn't going to run an 'app' in any case, but pushing confirms the decision.
nullc··on Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)
Anthropic's reasoning output isn't the real model reasoning but some sloppified summary of it.
nullc··on Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (2025)
> If the intermediate tokens represent reasoning or thought, you would expect "aha" to occur after the thoughts that led to the realisation, including the thoughts encoding the explanation: they don't have any other state.

Yes they do, they have their KV caches-- it's a pure function of the input tokens, sure but that doesn't prevent it from containing latent 'insight'. LLMs can and do pre-form the tokens they're expecting to output multiple steps in the future.

I wouldn't argue that the 'aha' means anything, but the structural argument that it can't that I think you're making isn't sound.

nullc··on Don't Paste the AI, please
> I think it can be OK if AI is used to polish up one’s message

Asking it to identify problems such as spelling, grammar, or general clarity is fine.

If you tell it to "polish" many commercial LLMs will absolutely slopify your text with characteristic AI catch phrases and non-sequiturs. Beware.

nullc··on Don't paste the AI, please
They're acting as a meatbot being piloted by the AI. Could call them skroders.

It should be the other way around-- you give the AI instructions and it does stuff. Any time the AI is giving you instructions and you do stuff, that should be a big red flag.

nullc··on RISC-V: They Should Have Known Better
I wonder how many of the obvious design shortcomings in RISC-V are from IPR avoidance / making IPR problematic parts optional.
nullc··on AI Isn't Outthinking Mathematicians. It's Out-Remembering Them
Sometimes the purpose of the proof is simply to demonstrate that some construct is a safe assumption for other more interesting work-- and could still serve that purpose even if it was entirely a black box.
nullc··on Qwen 3.8 27B
On 2x RTX A6000 non-nvlink connected but communicating across the CPU, with llama cpp I get ~60 tok/s for Qwen3.8-27B-UD-Q8_K_XL without any batching.
nullc··on Qwen 3.8 27B
It's a cheap to evaluate proxy for totally broken or not, which is a good start.

It also has a lot of resolution and not a lot of noise. Better would be multi-turn benchmarks with tools but getting good precision and accuracy for that is hard and computationally expensive.

nullc··on Qwen 3.8 27B
I'm not having any looping.

> --temp 0.2

Looping is a common symptom of changing the sampler settings from what it was RL trained with.

nullc··on Qwen 3.8 27B
35Ba3b is usable in plain cpu inference on a fast server, the dense model is MUCH slower. On GPU the 35ba3b is still around 2x the tok/s single threaded, which can be a good tradeoff for some applications.
nullc··on Qwen 3.8 27B
HUH? AgentWorld is a simulation of the world (e.g. tools and programs) for use by an agent!
nullc··on Qwen 3.8 27B
> (1) make effective use of retrieval tools and (

A downside is that you can't just download a lot of that knowledge, vs with the weights the copyright infringement has been outsourced to the lab. Nor can you just search for the info because the internet as a whole is increasingly aggressive at blocking anything that looks like an AI agent.

I'd love to see more retrieval powered local AI-- I think it's an area that open source development could excel. ... but there are advantages of having the knowledge in the weights!

Perhaps what needs happen is for someone to make an "ultrapedia", an AI restatement of a huge library of reference works-- created expressly for the purpose of being a locally stored corpus for AI agents.

nullc··on Qwen 3.8 27B
Diverging from the sampler used in RL training is not good for long multi-turn results-- it's a great way to knock models into reasoning loops that wouldn't otherwise.
nullc··on Qwen 3.8 27B
I wouldn't find glimmer interesting except that it has much less memory usage per token of KV than Qwen. So I can get 24x concurrent glimmer on 2xRTXA6000 (with 128k context) where I can only get 6 Qwen 27b. This means I can get something like 4x the aggregate tokens/s out of glimmer.

For some usages that speedup more than makes up for it being inferior to Qwen intelligence wise.

nullc··on Google is making private AI practical with homomorphic encryption
Private AI is practical by running the model locally, every much more so than any homomorphic encryption scheme.

So essentially the headline sells this as work to keep your data private, but really it's work to keep the AI-- which was trained on your code and your writing-- private.

nullc··on Qwen 3.8 27B
for layer parallelism (e.g. to get more vram) the bandwidth between layers is essentially nothing (like 16kb per token I think), so I don't think x4 would even be a problem!
nullc··on GLM-5.3: Frontier coding with emergent cyber capabilities
Modern benchmarks across the board really need to start severely penalizing disobedience and hallucination. A year ago models weren't really strong enough to justify this but they are now-- the frontier isn't in squeezing out the next bit of task completion, it's in making common cases not periodically be disastrously wrong.

A lot of the total cost of AI is fixing its "truth shaped errors", particularly in the presence of models that are very "gaslighty" when corrected.

GLM-5.2 is really the only model I've spent much time using that I didn't fatigue from being regularly lied to by the model, but that might be partially luck.

nullc··on GLM-5.3: Frontier coding with emergent cyber capabilities
Make cyber not Cyber.
nullc··on GLM-5.3: Frontier coding with emergent cyber capabilities
> Labs have already used up internet-scale data

Not really, but a lot of what isn't used isn't very good.

More important is synthetic data. Use a teacher model with RAG with a huge reference library to write synthetic transcripts of idealized behavior for the model. Use models to judge and correct these transcripts. Train on the good ones. Use bad traces to train the model to correct its own errors (e.g. don't train it to produce a bad transcript but if it finds itself in the middle of one train it to self correct).

Similarly, for tasks that can be closed loop evaluated -- e.g. running computer software and programming, unlimited amounts of novel training data can be generated... including for highly original tasks: e.g. run publications in any domain through a model prompted to look for programming problems suggested by the material. Then write/judge/improve transcripts of solving those novel problems.

I expect in the future smaller models won't be directly trained on any internet data at all-- but entirely on simulations of idealized expected behavior from the model under construction. Raw internet data in that case would show up in prompts, but never in the target output (except of course for prompts that are asking it to copy the input).

nullc··on NP-Overrated
Did you know: general purpose computers are completely pointless, because programs can run forever without producing a result.
nullc··on NP-overrated
I've made comments on HN on this point a number of times e.g. https://news.ycombinator.com/item?id=44284083

I had some tedious debate on HN once where I asked if anyone had any pointers to good parallel SMT solvers, only to fall victim to someone dedicated to dying on the hill of "parallelization can never make this kind of search faster" due to (often inapplicable) complexity theory fixation.

nullc··on NP-Overrated
That is proximal to one of my peeves in this space: People who misunderstand approximation results to be meaningful when often they're not.

For example, the minimum set cover problem shows up in cases like "What minimal set of test vectors covers all the conditions in my code?". There is an obvious greedy algorithm: "Start with nothing, pick the vector that covers the most yet-uncovered cases, repeat until all are covered".

There is an approximation result that says no polynomial time algorithm can do more than a small factor better than this greedy algorithm.

But this is a _worst case_ result, and absolutely useless for any problem you will encounter in practice.

It's trivial to come up with ways of improving the greedy algorithm: First off the simple greedy algorithm will often produce output which has completely redundant elements that can just be removed, because some collection of later added items that were necessary to cover some rare cases completely cover some earlier added item. Adding a simple postprocess to remove redundant elements immediately improves the greedy solution, particularly when the frequency of elements follows something power-law ish.

You can measure the frequency of each element and weigh uncovered elements by how rare they are (E.g. using entropy). This avoids the primary cause of the above duplicate selections.

You can use lookahead e.g. pick the pair of elements that together improve the score the most but then only commit to one.

You can use rarity weighed random starts, complete using whatever search you have, then retry multiple times.

You can compute new solutions using only the results of prior attempts. etc. etc.

In my experience basically any improvement over the greedy algorithm works on real problems, even before getting to a proper ILP solver. The greedy algorithm is just pathetic and will result in solutions much worse than you get from simple elaborations.

But over and over again you can find people being told to use the greedy algorithm because no polynomial time algorithm is better -- even in instances that are small and where actually enumerating all solutions might be tractable and justified.

nullc··on Spaghettifying DRAM
You're missing that modern CPUs substantially lock the users out of control of their own computer and include things like hidden additional network connected processors that run their own full on operating systems. ... and may well be used to surveil or remotely access your computers the the behest of powers unknown.

But they still use system dram, so this approach allows looking into those parts of your own computer from which you're normally blocked. At least on some hardware...

nullc··on Spaghettifying DRAM
This is a great starting point to go looking for PSP / SMM backdoors.
nullc··on Responding to the next frontier of critical cyber capabilities
Anyone else struck by the similarities between the description an "The Cookie Monster” by Vernor Vinge?
nullc··on AMD acquires Taalas to boost inference performance by etching models in silicon
Might be an interesting motivation for looped LLMs to cut the gate count down. Perhaps even a collection of mixed programmable layers and baked layers in a loop.
nullc··on Flock – Chilling Effects: Long Island's Emerging Open-Air Prison
Broad form car insurance might accomplish what you want:

https://www.policygenius.com/auto-insurance/what-is-broad-fo...

But I believe the intersection of states with highly private registration and where broad form is available is the empty set.

> I am not convinced your insurance company would data sharing with flock or its subsidiaries if they only have your vin.

Go pull your lexis nexis report.

nullc··on Prime Agent: A self-improving RLM agent
You shouldn't link that without mentioning that it is a particularly intense S&M/rape/incest fetish piece, and that the squick content is entirely integral to the story. I'm generally dubious of "trigger warnings" but if ever there was something that needed content tags-- this is it.

If you imagine it being written for alt.sex.stories.moderated but somehow failing to be erotic by virtue of being too explicit, and being a few orders of magnitude better writing for that venue... then you wouldn't be too far off. [Started making a silly example of how some sex story could fail like that and then realized I was litterally describing a scene from prime intellect].

It's also arguably the origin of a lot of the mentally ill ai-safety hysteria. Arguably an enjoyable romp for those who understand that it's fiction, but it seems a lot of people cannot.

nullc··on The title cards in Blade Runner are amazing
Good luck even finding the non-directors cut.

I liked the film noir voice-over.

← PreviousPage 4 of 34Next →