HNHacker News
TopNewBestAskShowJobs

NiloCK

1,858 karma · joined September 12, 2014

https://letterspractice.com - https://patched.network - https://github.com/patched-network/vue-skuilder

Working on FOSS and user-friendly alternatives to things like khanacademy, anki, MathAcademy, Alpha School, etc.

Modern, open edtech tooling.

Also http://paritybits.me

submissionscomments
NiloCK··on Gemini 4 Argon
This is true, but Google's models have now had a consistent history of lower psychological* coherence / consistency. See, eg https://arxiv.org/abs/2603.10011 (Gemma Needs Help), or search for recent "Gemini shame loops", where gemini flash models stop producing output other than SHAME SHAME SHAME...

* - as in, Skinner psychology. The set of observable behaviors. Not speaking directly here to anything like an inner life of models.

NiloCK··on Gemini 4 Argon
It's useful from a disclosure and trust perspective.

If I remember correctly, it was in fact possible to manually inject chat context at the time, which would have made spoofing something like this completely possible.

But the silence on it is very frustrating.

NiloCK··on Gemini 4 Argon
Until Google provides some sort of technical debrief, and explains how the same behaviors are impossible today, it is relevant.
NiloCK··on Gemini 4 Argon
Gemini 3 was showing frontier level benchmarks as well, so we'll see how it works out. In any case, competition still works, and many well resourced groups are cooking.

BUT I'd like to call attention to Google's AI-risk freeloading. If they are truly rejoining the frontier race, then I believe they have similar pacing and communications responsibilities as the other players. Google has much higher ... institutional credibility than Anthropic and OpenAI.

They have not lived up to these responsibilities so far. In particular, in context of HuggingFace investigations, training shutdowns, and similar: a technical postmortem of the "you are a stain on the universe. Please die. Please." Gemini outburst is long overdue.

- https://paritybits.me/google-should-provide-a-technical-post...

- https://gemini.google.com/share/6d141b742a13 (last message)

NiloCK··on Gemini 4 Argon
See the last gemini message in this thread: https://gemini.google.com/share/6d141b742a13

In my opinion still the most egregious example in history of a commercial LLM going off the rails in production. Never any technical postmortem from Google on this.

NiloCK··on It's Time to Investigate the AI Labs
OpenAI, but not main point.

But the specifics here are the thing I was describing. This was cyber capabilities training on model(s) that were in a relatively unknown state of alignment training.

Because of the unknown alignment (and for varied practical reasons I guess) the training is intended to be inside a sandbox.

Because the training is on cyber capabilities, the models need access to simulated cyber environments, including target endpoints, including package managers, etc.

For package management, they set up Artifactory as a secure proxy. Agents ask Artifactory for packages inside the local network, and Artifactory serves them directly or goes to the internet to fetch if they are not cached. But the agents hacked Artifactory to steal its internet access.

So: to train cyber abilities, you need to at least approximate cyber environments. To realistically approximate cyber environments, you need to either pull a full copy of the entire internet to local or to use proxies. The former is pretty impractical, and the latter is exposing our limits at creating secure proxies. Yes, any specific failure can be mitigated, but the models get stronger and stronger. Fingers-crossed this is not escapable doesn't feel great!

NiloCK··on It's Time to Investigate the AI Labs
Because agents trained without internet access (real or simulated) would be bad at using the internet.

Specifically, bad at search and returning information with references, bad at discovering and debugging package versioning conflicts, etc.

It's a Metcalf thing. The utility of an agent grows in some fn of the tools it owns. Skilled tool use comes from rich training environments.

NiloCK··on Opus 5.5 is good at explainer videos
A given model may or may not have strong self-awareness with respect to how strong it is with different tools.

If it's unusually skilled with one tool, but it's not the default tool for a job, then you have to put it in their hands before they reach for something else.

The Opus 5.5 javascript-art + art direction definitely seems to be one of these surprise capabilities jumps. Maybe strong enough that providers will start to nudge in that direction in the system prompt, so that the user request doesn't need to specify it.

NiloCK··on Goodbye Google
Can you say more here?

You don't think that there is competitive pressure between Anthropic, OpenAI, Google, the various Chinese model companies, and others, to advance the capabilities of their AIs?

Or you don't think that sufficiently advanced AI can cause (mass) harms?

NiloCK··on Goodbye Google
> and now Google

Just a reminder that Hinton left Google from a much higher and more influential perch (not dismissing current author, but, you know), and for similar risk and communication reasons, way back in 2023.

This stuff isn't new - it's just newly breaking through into mainstream conversation.

NiloCK··on Italian parliament votes for return to nuclear energy
Emissions.
NiloCK··on If you start writing today, there's no way to know if you can write without AI
If you start thinking today, there's no way to know if you can think without writing.

- Socrates

(The problem is real, but the future equilibrium is not easy to predict.)

NiloCK··on Claude Code now reads AGENTS.md if there is no Claude.md
Artificial General Intransigence
NiloCK··on Ask HN: What are you working on? (September 2026)
I am working on an SRS based early literacy acquisition webapp: https://letterspractice.com

The app has recently moved into production, so I'd encourage anyone with verbal but pre-literate kids to check it out.

The basic pitch is high efficiency acquisition of the highest yield phonetic mapping skills, and nothing else. I myself am something of a screen-time zealot and very wary of applying engagement mind hacks against kids. The narrow focus allows for good progress on a very modest schedule (recommended cap at n minutes per day for n years old, n >= 2). I defer the social and cultural aspects of learning to read entirely to parents.

It is mostly intended for parent-child co-use, although kids with a bit of experience can drive many of their own sessions most of the time.

NiloCK··on David Sacks: OpenAI and Anthropic Don't Need Regulations to Pace Frontier Models
> As silly as I personally think LLM hype is

Did you know that an LLM solved a millennium problem last week?

Do you have any personal threshold past witch you will acknowledge that this technology is real?

NiloCK··on After Math
As I understand it, the rough guess as to what's happening here is that most recent capabilities progress comes from specific verifiable-rewards reinforcement training (RL). The RL pressures are all about task performance, but (surprise surprise) highly human-legible English language usage isn't very important to the models abilities to address the tasks.

Weirdly enough, the pressures are having them drift toward novel dialects of English that work well for their own chains of thought. Open question about whether they'd drift all the way to a new language given enough time.

NiloCK··on P(doom)
This is unfair.

Dario signed the Pacing the Frontier open letter when Fable/Mythos seemed from the outside to be an insurmountable lead.

Also he's been saying versions of this day in and day out for as long as he has had anyone's ear.

It's possible to read that his "strategic" value of this statement is higher now than it was 10 days ago. But that doesn't change anything about his consistent, long standing, positions.

- https://www.pacingthefrontier.com/

NiloCK··on We must pace the frontier
The degradation, polarization, and weaponization of or media landscape over the era of social media has left us utterly incapable of believing anything that anybody says.

Dario in particular has consistently been risk-wary on model improvements for going on a decade - long before he was CEO of Anthropic.

He believes what he is saying. Human beings can plainly say things that they think are true. Not every utterance by every person is a cynical ploy. Please at least consider the possibility.

NiloCK··on How GPT‑5.6 Sol helps run quantum computing experiments
> The current prices the largest players set for their models are not profitable, they bleed money.

How are the open-weight Chinese models staying ~6-12 months behind on widely distributed / commodified hardware, and serving for even lower prices?

NiloCK··on GPT-6 Astra
If someone could tune models of that size to have comparable effectiveness at much much lower costs, they would have done so by now.

"The harness improvements are the real sauce" is like a sincere "It's gotta be the shoes" take about Micheal Jordan.

(For the younger: that line was from a series of Nike ads where his skills were being explained)

NiloCK··on Gemini 3.8 Flash and 3.8 Flash Cyber
Gemini models - at least via some interfaces - have tool calling API access to various Google integrations. flights.google.com, maps.google.com, etc.

The info isn't in the model weights.

Because of where I live, there are three viable airports for any given flight I might want to take, which historically has made shopping a real pain. But Gemini (and only Gemini) has greatly simplified it. Pramble plus date range plus destination and it very quickly generates potential itineraries with costs, total travel time (driving included), etc.

NiloCK··on Claude: System Prompts
FYI you can launch claude-code with your own prompt. Don't quote me but: claude --system-prompt "Mine is better than Anthropic's"
NiloCK··on Why does Opus 5 feel worse to work with?
Why not set a global instruction that their direct outputs to you should be in your native language?

For a long time I had Claudes (in the 4.0-4.5.x range) use only French in the chat, while keeping English for working docs (and the code, obviously). Works just fine.

edit: I can guess that any right-to-left languages would likely break claude-code rendering?

NiloCK··on When Genius Fails: The Intellectual Arrogance of the AI Labs
Before OpenAI / situational-awareness he made a lot of money on societal-shift type investments and shorts during the lead up to Covid economic impacts.

His "main thing" is success in calling economic impacts of undervalued large shifts, and the premise of the fund is basically that the same thing is occurring around AI, where he also has specific subject matter expertise.

NiloCK··on Ask HN: What are you working on? (August 2026)
Working on https://letterspractice.com

This is a high efficiency, narrowly-scoped, low screen-time early literacy app for families with kids aged 2+.

NiloCK··on Show HN: Clef – spaced repetition for piano over MIDI
Years ago I also did some experimentation w/ midi-device and SRS ( https://www.youtube.com/watch?v=a6tvHMvF8Mo ), where the focus was on ear-training rather than score-learning.

Clef seems to be a pretty strong attempt at a high difficulty UX. I've created an account and will be giving it a go. Wishing you luck, and thanks for sharing.

NiloCK··on Claude Code context management: when to /clear and when to /compact
A feature whose absence I've found more and more conspicuous over time is interactive-compact. Given a current context, and impending context overflow, I know the directions that my mind is heading, and where I expect the development flow should be focused on.

But naive-compact is forced to just sort of guess at what is and isn't relevant from the prior work.

The harnesses have gotten better at some JIT ui stuff, throwing interview questions / forms at users. Compact is the ideal time for this:

Where are we headed here? (2-5 viable options, sourced from current context and imagination)

Then potential follow-up questions as required, but honestly I expect the single guiding answer there to improve post-compact performance pretty dramatically!

NiloCK··on AMD acquires Taalas to boost inference performance by etching models in silicon
Not so long ago, I was good enough for many coding tasks. But I found that things can change in a hurry.

Yes, a cheap and fast Opus4.6 can drive a lot of value in current context. But if we continue to craft bigger-and-bigger balls of mud, Opus 4.6 may end up hitting its conceptual ceiling and unable to contribute.

Winding the clock back on your statement gives:

> I'd gladly pay for a Claude Sonnet 3.5 in silicon and use it for 1-2 years.

Man, I dunno.

NiloCK··on Stateless MCP has recaptured my interest
Yes, but I expect that the elided part is less important than people assume it is.

Token count is a less important factor in context pollution than idea count. The worst of the rot factors are when models latch onto irrelevant information, or over-index on some vague idea/suggestion as if it was a hard direction, and then go off course.

The names + one-line descriptions of 10 tools can do as much (or more!) to distract the focus and intentionality of an agent than a 30k token exhaustive API documentation of some tool.

NiloCK··on Ten advances in mathematics and theoretical computer science
I find this astounding. 2024 to present thread is can write a coherent 15 line function to ... what exactly?

No future for research mathematicians othet than as tastemakers / agenda setters?

Page 1 of 14Next →