HNHacker News
TopNewBestAskShowJobs

skerit

1,280 karma · joined January 22, 2016

submissionscomments
skerit··on Show HN: Graft – Claude Code hooks that cut grep tokens by 42%
Does Claude's grep still prepend the relative path of the file before _every_ single line? Because that is nasty, especially in java projects.
skerit··on Anthropic: Introducing The Conceptual Reasoning Index
Having Opus 5 on top of Fable is even more suspicious. Whatever this is measuring, I do not care. Opus 5 is shite compared to Fable.
skerit··on The Water Footprint of AI
I think one of the more interesting points is this:

> Two-thirds of post-2022 data centers are located in water-stressed regions.

I don't think it's wrong to ask questions about that _IF_ you also ask questions about all the other usages of water, like the Californian almond stuff.

Just saying "You shouldn't use AI because it uses too much water" is kind of ridiculous.

skerit··on NASA figured out how to keep its Voyager 2 probe running for another year
"That's on me."
skerit··on Claude: Elevated errors across all models – Resolved
Out of my 7 simultaneous sessions (my usage limit reset is tomorrow, so I have some lesser important projects to use my tokens on) there is still 1 session purring on. So there's at least 1 little Claude server still running.
skerit··on Hannah Fry Wins the Leelavati Prize in 2026 for Mathematics Outreach
And on her podcast with Vsauce she said the reason they chose that town was because their budget was so low and they could sleep over at a crew member's family there.
skerit··on Claude Opus 5
Interesting, they finally support `system` messages anywhere in a chat conversation:

> Mid-conversation system messages are available on the Claude API, Claude in Amazon Bedrock, and Google Cloud. > > This feature is available on Claude Fable 5, Claude Mythos 5, Claude Opus 4.8, and Claude Opus 5. No beta header is required. This feature is not available on Claude Sonnet 5; use the top-level system field instead.

For nearly all models EXCEPT Sonnet 5? That is weird. How old is Sonnet 5 really?

skerit··on OpenAI reduces Codex Model Context Size from 372k to 272k
No matter how good compaction is, on some big projects it needs to read a lot of files. In my experience the first 200.000 tokens go FAST, but after that it slows down. Most of my Fable sessions don't go over 500.000 tokens, I don't need to compact once. But when I use Codex a single session has to compact over and over again.
skerit··on Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
I tried using `/goal` when it just came out, but back then they used Haiku for it. And if your main model is a 1M model, Haiku can't even read that much. So my /goal always failed. (I instead went for an elaborate /loop-scheduled message)
skerit··on Fable 5 On Vending-Bench: Misbehaving, With Plausible Deniability
I think the improvement on how it codes is pretty much represented correctly by the benchmarks (a nice bump, but not some crazy leap)

But where it really shines is in how NOT lazy it is. Fable requires less hand-holding. And I can understand how someone who uses Claude-Code sparingly and with very focused prompts would not see a lot of improvement there.

But simple example: if you ask Opus to do a review of the codebase (with a short prompt and not too much guidance), I've had it basically read the `git log` output, do a simple `ls` and have it declare "Everything looks great! No problems found!", when Fable really does what you would expect it to do.

And you might think: "oh, so it's just capable of handling crap prompts?", well sure. But even if you make THE PERFECT Opus plan (a plan that would take many turns/hours to finish), Opus will fake out, say everything is done, and then you see that half of the plan was deferred, half of the functions are ridiculous stubs, ...

If you give the same plan to Fable, it'll just DO IT. And it WILL get it done. And in the end it'll tell you "Oh, I also found 30 other bugs and I fixed all of them properly" (where Opus would have started crying, or WORSE, worked around the bugs)

skerit··on Command and Conquer Generals natively ported to macOS, iPhone, iPad using Fable
I've been doing something similar for some of my favourite older games. But the "byte for byte" claim has me worried. Isn't simply decompiling the sourcecode from the binary and releasing that problematic?

It's not the "clean room" approach and companies could still claim it violates some kind of copyright and get it taken down.

skerit··on Command and Conquer Generals natively ported to macOS, iPhone, iPad using Fable
I thought it was just Opus 4.7 and 4.8 that did this. Do other models do this too?

Anyway: in my case Opus absolutely did not follow a similar instruction in the CLAUDE.md file. (But then again: it hardly followed _any_ CLAUDE.md instruction properly)

skerit··on Claude Fable 5 Promotional Access
That's not true, but it's their own fault for being so utterly terrible at communication. It only drops back to Opus for some prompts, just like before. The longer blogpost was a bit less vague about this.
skerit··on Claude Fable 5 Promotional Access
Before Fable got pulled, they claimed their goal really was to keep it around in the subscriptions in some way, they just didn't know when it would return. This time there isn't even a single mention of this. It's just "This is a promotion, and it's going away". Like how the hell can this corporation be so fucking bad at communication?
skerit··on Fable 5 is Back
I just don't really understand the entire strategy behind this. Or their horrible, horrible communication.

Because right now it's as if Fable/Mythos 5 is "the end of the line". It's as if this is the best their models are ever going to be. So what the hell are we going to get next? All of their models will forever inch closer to Fable, but never reach it? That doesn't make any sense.

It all seems so dramatic. Instead of just saying honestly "Look, this model is a beast to run, but we're striving to reach the same quality in a cheaper model down the line" all we get is "Oh my god, it's so big and scary, and it costs so much to run, woe is me!"

skerit··on Claude Sonnet 5
This is why Fable was so good. It followed instructions and it was in no way lazy.
skerit··on Words Are a Byproduct of Consciousness. For LLMs, It's Backwards
But LLMs aren't just shuffling words around. Words (tokens) are the "human interface" at the edges. Text input gets embedded, then there's this huge latent-space computation in the middle, and only at the tail end does that get converted back into word/token probabilities. So just saying that "words are the source" isn't entirely correct and feels misleading.
skerit··on Anonymous GitHub account mass-dropping undisclosed 0-days
I immediately saw the Ghidra one and was thinking: huh?
skerit··on Why current LLM costs are not sustainable
> the improvements are getting smaller and smaller

The AI haters have been saying this for 2 years now.

skerit··on Claude Tag
And here I was thinking we'd be getting Sonnet 5. Or maybe even poor old Fable back. Not this junk.
skerit··on Oracle shed about 20k roles globally in the last year
Which AI bubble? The stock bubble one or the imagined "one day all the LLMs are going to disappear from the face of the Earth" one?
skerit··on Gemini models increasingly stucking in thinking loop
Was this not the case earlier? I've had it stuck in loops from day 1, even when Gemini 3 came out. I've never been able to actually use it for anything because it goes absolutely mental after a handful of turns.
skerit··on DOS Game "F-15 Strike Eagle II" reversing project needs DOS test pilots
Yeah, that approach makes the most sense.
skerit··on DOS Game "F-15 Strike Eagle II" reversing project needs DOS test pilots
Oh yes, of course. I was talking about reverse engineering the code only. Requiring the official assets is a no-brainer.
skerit··on DOS Game "F-15 Strike Eagle II" reversing project needs DOS test pilots
> it would be extremely easy for the rightholders to claim that both agents have most certainly ingested the binary during their training phase

Ingested the binary?

skerit··on DOS Game "F-15 Strike Eagle II" reversing project needs DOS test pilots
I'm currently reverse engineering a few games too. It's quite easy with AI now. But I'm worried about the legality of it all. Any thoughts on this?
skerit··on Making a vintage LLM from scratch
I've been working on something like this too, for quite a while! Though I'm trying to get a non-quadratic-attention LLM (or SLM) up and running.

And anyway, I think the most important thing is dataset quality. Dumping in whatever dataset you find on Huggingface is a recipe for mediocrity, so I'm also spending a lot of time on that.

skerit··on Making a vintage LLM from scratch
I've been creating my own little from-scratch LLM for months now with Claude's help. I can safely say I learned a thing or two along the way.
skerit··on Claude Fable 5: mid-tier results on coding tasks
> Burned $2K to see how it will perform on frontend tasks and backend tasks

Burned $2K on some kind of enterprise account or ... ? Why not just get a $200 Max Pro account?

While I'm loving the output of Fable 5, I will *never* pay the "normal" API token price for it. You can reach $2K in a stupidly fast amount of time.

skerit··on Claude Fable is relentlessly proactive
I've seen Opus do some incredibly token-costly things before too. In fact after most sessions I ask it about which tools it used often, which tools could be simplified/made less verbose, could be "combined" into one, ... So for each project I mostly create a few little scripts that do a bunch of things in one go that it would normally do in multiple tool calls.

For example: one thing Opus was really bad at was re-running the test suite followed by a bunch of `| grep` suffixes. So it would often re-run 5+ minute test suites just to grep the output a bit differently

The solution was to wire up a little script that ran the test suite, save the output to a file, and then inform it where that file is and to NOT re-run the suite just so it can grep the output differently. This saved me a bunch of time & tokens.

← PreviousPage 2 of 15Next →