HNHacker News
TopNewBestAskShowJobs

extr

5,748 karma · joined May 26, 2016

submissionscomments
extr··on Coding Is Not Solved
IDK with Opus 5.5 it does kind of feel like it's solved.
extr··on GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?
It's a great model and you're right it does feel quite natural at times while Fable 5.1 still has a claude-ish shape to it. Unfortunately I just find that it's not reliable enough as a daily driver and ends up performing specialist tasks rather than being the primary pane of glass.
extr··on GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?
I used to do this but recently I switched to having Fable 5.1 spawn forks of itself rather than Opus subagents. Yes it's more expensive but you don't pay for reads that already happened pre-fork, and you end up doing less rework since Fable agents are just much smarter.
extr··on GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?
- It doesn't write great code.

- Occasionally has strange tics around asking for permission for obvious next-steps, implied actions, etc.

- It's very expensive, both in terms of tokens and % usage on subscription plans.

- Relatedly, effort level is unintuitive. Sometimes it seems like higher effort levels are actually cheaper due to not under-thinking and needing to correct work. But other times they are overkill and send the model into rabbitholes.

That said, it's fantastic as a code-reviewer or "hunter seeker". It's better at finding bugs than Fable and "Get this well articulated task done single-mindedly" is an Astra-shaped task.

extr··on GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?
I just tell everyone to use Fable 5.1 for everything at this point. Astra is unfortunately a dud, I'm sure they will try to fix a bunch of it with GPT-6.1 but OAI has had this issue for awhile now where every other generation has some sort of strange tic, or reward hacking issue, or something. It's almost like they are balancing the RL on the tip of a needle.

Opus 5 has issues too, comment-slop, claude-ish, etc.

5.1 on the other hand can seemingly do no wrong. Easy to work with, writes human-level code. Expensive, yes, but even at Low effort it's well worth it.

extr··on NextDNS: fully configurable DNS, accepts money from customers as business model
I use NextDNS at home. It's great to have lots of control around DNS while also running a locked-down router (Eero). I personally use it in combination with a raspberry pi to run a version of this https://cazander.ca/2024/realfakeframe.io/ that redirects Fujifilm camera uploads to my Google Photos account rather than the proprietary frame.io integration. Well worth the cost!
extr··on Vomit: Clean up Claude 5's token output with a separate LLM
Sounds like you have a lot of axes to grind outside of simply "the latest models regressed on delivering concise prose".
extr··on Vomit: Clean up Claude 5's token output with a separate LLM
It's "literally unbearable" when the AI that completes software engineering tasks at 100x speed and quality from 2 years ago uses too much jargon?
extr··on Vomit: Clean up Claude 5's token output with a separate LLM
I'm sorry but the whining over LLM output styles is embarrassing. Do Claude and GPT models always respond in exactly the way my most articulate coworker would? No. The overused jargon is absolutely annoying. But these things aren't my drinking buddies, they're professional tools. It's not _literally unreadable_. It's just not ideal. Most of my tooling is "not ideal". That's okay. That's what I'm paid for. I just work around it.

For me I added some instructions to speak clearly and it helped marginally and that's fine. There will be a new model out in a few weeks where I'm sure they've laser focused on this issue since nobody can shut the fuck up about it. The same thing happened with GPT if anyone can recall the ancient period of 4-6 months ago.

extr··on Claude Code May–August 2026 weekly limits promotion
Unfortunately Sol does not compare to Fable at all.
extr··on Grok 4.6
keep in mind fable = mythos which as been "done" since february. so the gap is not 2 months, it's more like - techniques probably started "working" in late 2025, now are trickling down to 2nd tier labs 9 months later.
extr··on Grok 4.6
It's because Fable is just synthetic RL tasks + scale. The secret has been out for awhile now.
extr··on Manus will return to operating as an independent company
Manus did a lot of harness work to make up for gaps in Opus 4.5 tier models. I tried it for a time - they had a great deep research/PDF generation pipeline, parallelization, etc. The bitter lesson has now come for them: the latest models no longer have these gaps and the entire premise of having a unique product focus on this area is no longer relevant. Even Ant/OAI have let their "Deep Research" capabilities fall by the wayside, you can largely get the same result by asking for subagents or simply "keep going" style prompting.
extr··on Managing AI Coding Costs at Scale
yeah it's true, you do have to guide them. i find that the key is you have to know what's possible. you have to have the instinct for "this really shouldn't be so difficult". my junior SWE coworkers have the same trouble as your coworkers.

but the revolution is it doesn't take that long. in like 15 minutes you can chat with fable and get to the meat of whatever the issue is with repeated questioning. and then it does the solution for you. so it's not magic but it's still like a 100x speedup.

extr··on Managing AI Coding Costs at Scale
$80 is definitely low now that I look at my numbers. but not OOMs low, it's closer to like $200 on heavy days. i don't know how you're doing $3k/day, that's wild. i'm pretty aggressive about compaction and session restarts, and i reserve Fable 5/Sol XHigh for "main thread" orchestration
extr··on Managing AI Coding Costs at Scale
Yes lol. Of all things people are getting on me for it's the number of LoC x Years In Business of this startup. I don't fucking know, I didn't start the company and I wasn't here for several of those industrious years. Looking now it looks like we have slightly fewer LoC than that, I was counting some of the generated stuff.

But who cares? The point is any codebase over a few years old with lots of customers and a big surface area has lots of code, much of it "legacy" from the standpoint of a guy in 2026.

extr··on Managing AI Coding Costs at Scale
> "unguided" means "I typed a prompt into claude code and waited yolo"

Yes, this is literally what that means.

extr··on Managing AI Coding Costs at Scale
This is a great point and I agree. My own productivity varies based on what part of the codebase I'm working on. If it's "been in there before" and I know the right questions to ask, I can one-shot a good design/improvement. If I'm spending 20-30 minutes asking Fable to "draw a diagram so I can understand" - probably less so. But notably, I CAN get there in a fraction of the time it would have taken before. You can general personalized onboarding docs to ~anything.
extr··on Managing AI Coding Costs at Scale
It's a fair point, it's not truly unlimited and I do wonder how that would change my workflow. I can definitely imagine if I was inside Anthropic or OAI with unlimited "fast" tokens, you would be more tempted to hand over even more of this process. I completely understand why they talk about "graph engineering" and such, my entire workflow above could be a graph and I could try to increase my leverage even further. Realistically though I am bounded by product decision making, not code output right now.
extr··on Managing AI Coding Costs at Scale
No, actually. The point is to build a profitable business.
extr··on Managing AI Coding Costs at Scale
> unguided LLM usage

Why aren't you guiding your LLM usage? Is that what I said - to spam it and not guide anything? Or to have a careful workflow where you agree on design and maximize your human judgement/leverage?

> any state that's not explicitly being tested and verified in QA loops

As opposed to before, when engineers perfectly reasoned about code behavior from first principals and QA was unnecessary?

extr··on Managing AI Coding Costs at Scale
How is this not true? Taking a Senior SWE @ ~$200K, even just the base salary cost / 2080 working hours is $100/hr. Fully loaded employer cost + accounting for non-coding time gets you to upper 100s easily.

Even for a junior making $100K, I have a hard time believe their time is worth less than $75/hr or so.

Edit: Fine, "Senior" is not "Average". But naive salary is not the true numerator.

extr··on Managing AI Coding Costs at Scale
Yes 100%. This morning I casually prompted Codex to drive the browser to complete extensive performance testing in-situ that would have literally been weeks of work before. Probably in reality it just wouldn't have been done, and performance guarantees would have been attempted up front via more careful design.

In this case the design was also AI generated, and there were limited wins to be found because the design was already superb.

extr··on Managing AI Coding Costs at Scale
Have you worked at many startups?
extr··on Managing AI Coding Costs at Scale
This was more true a few months ago but Fable has improved the situation considerably.

Also just remember - minimalist code looks and feels great but customers do not read your code. I have caught myself many times providing "corrections" to abstractions that were already ~fine, just not perfect. The average SWE costs $200/hr. Careful you don't burn $50 worrying about code that will likely be rewritten or can be better abstracted when that's actually needed.

extr··on Managing AI Coding Costs at Scale
Disagree. I operate this way inside a multi-million line legacy codebase.
extr··on Managing AI Coding Costs at Scale
Keep the decision-making and execution separate. Use the high IQ models to chat about the design and make them drive subagents to do the actual work. "Chat" style threads are actually quite cheap. Where it gets expensive is having Fable 5 output thousands of lines of implementation where 95% of it was already overdetermined and there were only a few important judgement calls.

I actually have no doubt that I could replace my Opus 5 Low/Medium subagent profiles with Grok 4.5/GLM 5.2/Deepseek v4 Flash and perf would probably be pretty similar.

On top of that - highly recommend adding accurate cost counters to your statusline. You can't improve what you don't measure! (Or even have any intuition about).

extr··on Managing AI Coding Costs at Scale
Performance is better than ever. It's never been more practical to set up wildly complex synthetic test environments and measure perf wins. Plus the models will find every possible algorithmic/design improvement.

It actually gives me quite an uncanny feeling, bulldozing over years of human optimization work with a newer, "perfect" design. Like bringing an AK-47 back to the middle ages.

extr··on Managing AI Coding Costs at Scale
I would be really curious to hear from devs at Databricks what the experience of development is like internally. I work at a small startup with essentially unlimited AI spend budget - the entire point is that I should be turning to it at every opportunity since our human labor is so expensive relative to tokens. So generally it's like:

- Spend most time prioritizing/discussing what to do.

- Once that's agreed, use Fable 5 High + 5.6 Sol XHigh come up with a design + plan. Agree on the high level plan. (Usually this just comes down to choosing where the change belongs on the spectrum between minimal patch <-> full redesign)

- Use Opus 5 or Sol Med to execute

- Auto-fix bugs and CI until green + thermonuclear review skill x3.

- Manual interrogation of change/nits

- Come up with QA plan and have Codex Computer Use execute on it

- Manually spot check the final result (usually a sizable diff, thousands of lines, complete feature E2E, etc)

I probably spend like $80 a day at least but I produce the output of 3 or 4 2022 engineers and probably at better quality. So it's easily worth it. Would I save money by switching to GLM 5.2 and such...perhaps? IDK. At our scale it's not worth the time spent building the eval harness to actually understand the performance tradeoff.

extr··on Deflock Casa Grande
I don't really care if it was made by AI or smeared onto the keyboard by a monkey. It was effective in it's job and communicated clearly.
Page 1 of 28Next →