HNHacker News
TopNewBestAskShowJobs

lwarfield

58 karma · joined November 9, 2021

submissionscomments
lwarfield··on Show HN: Jev Plays Pokémon Red
Looking at the diagram in the gh repo, it looks like this is entirely jev. Are there any examples of people having a big model like Fable handle high level goals?
lwarfield··on GPT-6 Astra, looped transformers, and hidden reasoning
I'm kinda surprised that the mixture of depths paper didn't come up here. It approaches the other direction of sometimes dropping layers:

https://arxiv.org/abs/2404.02258

lwarfield··on Extracting Steering Vectors from J space
The using a J lens is super cheap compared to inference. You basically add a single matrix multiply per layer. You probably wouldn't even notice the overhead in a good implementation.

You should even be able to create a j lens from scratch, but it might take a while. I was able to do it in a few hours on an H100. Creating the J lens is basically the equivalent of calculating a few thousand training steps for a model (256,000 backprops in my case). I've got more details in a blog post:

https://blog.lwarfield.dev/layer-scope/

I'm currently at work and can't those matrixes up until I get home. I'll update this comment with a link later.

lwarfield··on Extracting Steering Vectors from J space
If the author would like, I self computed a j lens for the 27b version of the qwen model. I used it for my own exploration in this area, and can share it if you want.
lwarfield··on Claude Fable 5.1 and Claude Mythos 5.1
Same for me. Every single time I tried it got flagged. I think this will be my litnus test for if the safeguards are good enough for benign requests.
lwarfield··on The Connections Museum
Best museum in Seattle! It also feels a lot more "hands on", than most museums.
lwarfield··on France's tax agency got hacked (in French)
\s Take my angry upvote!
lwarfield··on Ask HN: There is current insistence upon learning what LLM's do under the hood
Personally I'm curious to the point of doing borderline LLM archecture research. This come from genuine curiosity, and not a want to use LLMs better.

I haven't gotten much out of for using LLMs though. It makes me understand the short comings of LLMs, and I feel like I got an early insight into how important context management is.

lwarfield··on GPT 5.6 Sol is the best "vision" model OpenAI ever released
I currently have fable organize a bunch of 5.6 sol agents when working on my personal projects. This makes me wonder if I should add something along the lines of "For tasks that involve visual analysis, have gemini 3.7 look at images generated."

Overall I've been hooked on using agents from different companies for what they are best at (Thanks to Theo). Fable is expensive, but unmatched for planning and top level organization of other agents. Sol is fast, will persistantly go after goals (sometimes to its detriment), and does well with computer use.

lwarfield··on Claude: System Prompts
I've always wondered why the industry relies on the giant monolithic system prompt. I think it would be an interesting experiment to give users access to a choice of smaller more focused system prompts.

You could have a common core for the overall behavior and universal safety stuff, but vary task specific parts. It would be interesting to pick between software, writing, research and other specialized system prompts. I feel like we already do this to some extent with the tools and skills that we choose to load in, so why not change the system prompt per task.

lwarfield··on Anthropic Risk August 2026 [pdf]
Yes it is:

> Same model weights as Mythos 5, deployed with higher-coverage safeguards (see Section 4.5.2.2)

lwarfield··on Anthropic Risk August 2026 [pdf]
> 6.2 [Appendix redacted] > This appendix describes the criteria for our blocking bioclassifier exemption policy, and has been redacted from the public version of this report for security reasons.

>6.3 [Appendix redacted] > This appendix, redacted from the public version of this report, details the changes made to our constitution to expand classifier coverage to harmful uses in scope for the CB-2 threat model but not the CB-1 threat model, as described in Section 4.5.2.1.

interesting...

EDIT: After reading more I'd recommend looking at Transcript 2.20.A. Its a transcript of claude going over the redactions in the report. The section says its specifically for section 2, but the transcript also mentions other sections.

lwarfield··on Anthropic Risk August 2026 [pdf]
> More capable than Mythos 5 in some areas, less capable in others; overall slightly more capable.

This sounds like it might be a Mythos finetune for some specific task.

EDIT: After reading some more reading, it looks like model 2 might be an AI research fine tune based off the section 3.4.3 CoBench

lwarfield··on Layer Scope: How I used $20 of compute to make a new way to look at LLMs
Hey hackernews, I'm a long time lurker and first time poster. I've been doing a lot of LLM research on my own, and a friend of mine mentioned I should do a writeup of a small project of mine. I mainly want to show that you don’t need to have millions of dollars and be part of a big lab to do this type of research.

Inspired by Anthropic’s newest paper on LLM interpretability, an excellent blog post series by David Noel Ng, and other research I’m currently working on, I’ve created a cool new way to read the state in the middle of LLM networks! The best part is that it’s a wonderfully simple approach:

1. Choose a transformer block and token in an LLM you want to inspect.

2. Run the last 4 layers of the LLM.

lwarfield··on 30papers.com – Ilya's 30 essential ML papers, in a beginner friendly format
Its interesting seeing how many of these researchers became the heads of frontier labs!
lwarfield··on 30papers.com – Ilya's 30 essential ML papers, in a beginner friendly format
For beginners I'd recommend the Welch Labs Illustrated Guide To AI if your not well versed in reading papers. Its a beautiful book that I've enjoyed going through. I'd recommend going through these papers after reading that to get a deep understanding.
lwarfield··on Political bias in AI: Where the AI models stand
Do they state if they used an API endpoint without a system prompt, or were these done via prompting the currently existing chatbots with a system prompt? Without a system prompt, I'd imagine there would be more variance in answers.
lwarfield··on Will It Mythos?
There's also lowering the number of experts you run in MoE models.
lwarfield··on Artificial intelligence is not conscious – Ted Chiang
>Yeah, it's kind of mind boggling that Ted Chiang (of all people!) can't imagine intelligence without a body. and the whole thing just begs a lot of questions.

Damn, what a line!

Another thing that bothered me with his baseline for consciousness was that it did not involve the ability to change one's self. A big part of being conscious in my mind is how one's experiences shape them, and how someone can shape themselves. LLMs completely lack this, their weights are static. An LLM isn't going to be molded by a bad breakup, or a relative passing away. An LLM isn't going to set up a routine to get stronger with training, nor smarter by reading up on a field.

lwarfield··on Claude Code refuses requests or charges extra if your commits mention "OpenClaw"
This is some real "There is no claw in ba sing se" stuff.
lwarfield··on OpenAI models coming to Amazon Bedrock: Interview with OpenAI and AWS CEOs
Well that didn't take long.
lwarfield··on Ask HN: Any interesting niche hobbies?
I do a lot of lindyhop Swing dancing as a hobby. It's not innovative work, but I find it rewarding.
lwarfield··on Show HN: Gemini can now natively embed video, so I built sub-second video search
Damn, I need to going with my embeddings project. I've currently got a prototype for using embeddings (not gemini in my case) for making a game that's kinda reverse connections:

collections.lwarfield.dev

lwarfield··on European Commission Trials Matrix to Replace Teams
I've always thought the really low bandwidth support they added a few years ago was to support the french subs. It matched all the requirements of VLF/ELF communications.
lwarfield··on Man still alive six months after pig kidney transplant
I ran into this guy nere Interlaken 2 days ago! Had a nice long talk at the post office over how he was going to present this in Geneva. I heard that this is result is with little or no rejection drugs as well!
lwarfield··on Broken legs and ankles heal better if you walk on them within weeks
When I broke a joint in my pinky a few years ago it was pretty easy to tell. Early on the range of motion was the limiting factor, and I'd move it back and forth as much as I could without any pain. After that I worked on strength in a similar way, do as much as I can with no pain. I went from "you'll never play an instrument again" to rock climbing and Viola practice.

Overall, seeing my strength and range of motion slowly get better was immensely satisfying and your body is pretty good at letting you know when you're getting close to a limit.