HNHacker News
TopNewBestAskShowJobs

criley2

4,694 karma · joined June 19, 2013

submissionscomments
criley2··on Prompting Claude Opus 5.5
Deepseek Flash v4.1 is only "40X cheaper" if you do not account for the time of the engineer reading the output. If Opus 5.5 high requires 1/2 of the actual engineer time, and the engineer costs $100-$200/hr, then Deepseek v4.1 is actually the more expensive model to use.

I have tried your workflow many times, and simply letting Opus do the implementation costs much less than wasting hundreds of millions of tokens letting deepseek and opus go back and forth and back and forth. And bonus, my project finishes in 5 minutes instead of 20.

criley2··on Prompting Claude Opus 5.5
In my experience it never works well on any real work. In fact, I'd go the opposite, plan with the dumb model and execute with the smart model because at least the model writing the code and solving the emergent problems is capable.

In my experience (and I've been trying this a bunch): smart planner + dumb executor produces worse code with higher spend than simply using the smart planner to do both.

It's easy to understand why:

- If the planner has truly thought the issue through, properly designed the solution, solved all of the emergent problems, then the final "write" of the code is just a few more output tokens.

- If the planner has NOT truly planned the issue completely, then you're letting a substantially dumber and less capable model make significant decisions, and trusting its problem solving, without having a better model check it.

If you're highly cost conscious (paying for your own tokens and not making any money) then you have no choice but to trade your time and effort for tricks like this to save money by lowering the quality of your output.

But if your employer is paying for tokens: just use the smarter model. You save your time preventing re-work and reducing code review, you save your employer money (primarily from the cost of your own labor and reduced rework), and you get a better output every time (Opus 5.5 mogs Deepseek 4.1 flash in every single way except cost).

criley2··on Meta Blocks President Lula's Facebook Page, Campaign Ads 2 Weeks from Election
Conflating "the internet" with "~Trillionaire owned, oligarchy controlled, American social media" is certainly a choice.

I'll say the opposite: Every single place that has banned American social media is better for it. Every group of people denied access to American social media are healthier and happier for it.

criley2··on The Mafia may be keeping fentanyl out of Italy
So if you bought a bottle of water, and inside of that water was a poison that kills you, intentionally added by a criminal because it makes them more profit, that's not murder?

Those people were doing a line of coke, they had no idea they were about to overdose on an opioid.

criley2··on The Mafia may be keeping fentanyl out of Italy
Tiny bits of fentanyl are added to a variety of recreational drugs used extensively by all kinds of people, rich and poor. They add it to cocaine, heroin, meth and even marijuana. They imitate oxy, percocet, vicodin, xanax and adderal using fent laced pills as well.

As much as 20% of Americans use these recreational drugs. Europeans on average use recreational drugs like half as much as Americans, maybe 10% use them.

criley2··on The Mafia may be keeping fentanyl out of Italy
Fentanyl is made by Mexican cartels from Chinese precursors. I suppose there aren't many Mexico's with massive organized crime networks in Europe to pour deadly drugs into their countries.
criley2··on The Mafia may be keeping fentanyl out of Italy
No one "does" fentanyl, it's a highly addictive additive that is secretly added to other things. Even so, 50,000 - 70,000 Americans per year die to it (more than suicide and gun violence many other things).

This is an "American" thing because fent primarily comes from Mexican cartels (Sinaloa and Jalisco New Generation).

They buy precursors from China, they produce it in labs in Mexico, and then smuggle it into America where they kill 50,000 to 70,000 Americans per year.

Not getting political here -- but what would you think if organized crime in a neighboring country was murdering 50,000 to 70,000 of your countrymen per year?

criley2··on Claude Code reads AGENTS.md only when telemetry is on [fixed]
Developer comms are very important. I've often "fixed" something that was working exactly as the designer made it. It certainly wasn't a bug in my work though. I've worked in places where we call this a "design omission" ;)
criley2··on Claude Code reads AGENTS.md only when telemetry is on [fixed]
They're not fixing it, they're "changing" it, which makes sense, because the design changed.
criley2··on Claude Code reads AGENTS.md only when telemetry is on [fixed]
I don't think "bug" is the correct term. They put a feature behind a feature flag, and feature flags don't work if you turn them off (via telemetry). That's "Working As Designed™".
criley2··on How GLM built its own inference infrastructure
I believe the that the companies who claim to not train on my data are more likely to not train on my data than the companies who refuse to even claim they won't.

Also why Meta gets a +1, just charge less money on the training path.

criley2··on GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?
I have been writing an internal code review tool that is a bit maximalist. I created subagents for many internal domains and technologies we manage, with prompts focused on best practices, common problems, owasp guidelines, etc), a separate tier of wider band subagents (design, rollout, security/privacy), and a final agent at the top orchestrating and combining. I also use adversarial validator passes against all findings.

Right now, Fable 5.1 delivers incredible reviews. Opus 5 delivers good reviews. These agents are finding really impressive issues that humans just don't have the attention span to track down. My reviewer has a very impressive signal to noise ratio at this point, after half a year of iterating and improving. (I use a lot of Opus high, Opus medium for less critical tickets/domains, Fable 5.1 high for critical domains and all of the issue validators, and even Fable 5.1 xhigh for my design agent, whose job it is to think about the project at a high level and provide the kind of high level tech design review that AI notoriously can't do well)

I'm testing Astra so I don't have strong opinions yet. I've also done extensive testing of the same skill and subagent pattern in opencode/omp using GLM 5.3, Kimi K3 max, Deepseek V4 pro, Deepseek V4.1 flash, Qwen 3.8 2.4T max, and others.

My experience is that open weights models find between 1/4 to 1/2 of what Fable/Opus stack can find, and often miss the most critical issues. I work where privacy isn't just good behavior, it's enforced by law, and the Fable/Opus stack has found privacy leaks that the openweights stacks don't find.

You can imagine that paying for these Claude runs isn't cheap, each one can eat 25-33% of my 5 hour limit. I am quite desperate for openweights models to be competitive, but at the end of the day, the biggest limit here isn't the price difference between GLM 5.3 max (my current best-in-class choice for open weights, offering Kimi k3 performance for like half the price), it's the cost to the business for shipping lower quality.

Can't wait to dig in more with Astra, I just haven't iterated much on my skill port to codex yet.

One criticsm I have for the article, that is important for my own work, is not simply comparing "bugs found" because these agents can find endless reams of lows and nitpicks that are just ~worthless hardening. I'd be much more interested to see how many critical/high/medium's each test found, not "overall bug count". I also think review is about A LOT more than "finding bugs"...

criley2··on XCancel service is suspended until further notice
Nah 10 years ago I could view the whole website logged out. Now it's been reduced to a single post with no replies.
criley2··on Fable 5.1 Solves the Cyphral Distich, a 370-year-old cipher
Solar only became "economical" (read: profitable) because a socialist economy dumped a huge amount of money into scaling it up without requiring it to be "economical". It could have been half a century or more ago if we actually cared.

Nuclear was not stopped by the environmentalists, it was stopped by the fact that it cost 4X more than coal at the time. You claim that solar wasn't profitable in the west, thus it didn't take off, surely you can also see that nuclear wasn't profitable in the west, thus it didn't take off as well.

The same country that invested the time and money to make solar profitable is also investing the time and money to make nuclear profitable, with nuclear reactors entering mass production...

It was never about "economical", it was about a system being mature enough to make long term investments. Ours simply can't do that anymore.

criley2··on Fable 5.1 Solves the Cyphral Distich, a 370-year-old cipher
It's nice to call it "politics, ignorance and greed" but those three are just "Capitalism".

Our system is doing exactly what it is designed to do. Nuclear reactors were never profitable compared to coal or gas, so it never succeeded in strongly capitalist societies, only seeing great success in socialist economies where the people can invest outside of a profit motive.

It's actually quite funny that socialism is saving the day. The mega-capitalist countries turned their backs on nuclear and solar because they were less profitable than gas and coal. But socialist China invested anyway, and now China is mass producing nuclear reactors and every layer of the solar stack. China produces 50% of all nuclear reactors, 90% of all solar panels, 90% of all battery systems for solar storage.

Now that a socialism-based society has proven market viability, suddenly the greedy capitalists want in. But they're decades behind and don't have the private debt appetite to compete.

Womp womp... At least someone is leading the energy revolution.

criley2··on GPT-6 built this earth exploration site in 5 prompts
The term that Anthropic is now using is "mannered prose". If the creator of this example simply prompted "Remove all mannered prose" then the entire experiment would suddenly become normal sounding. In the fable 5.1 prompting guide, they have a longer prompt for removing mannered prose too if needed. I find it works on Astra as well.

https://platform.claude.com/docs/en/build-with-claude/prompt...

criley2··on Don't let anyone take away your big box of cables
This post isn't convincing me. I spent so much time meticulously organizing my techno box. I bought a back of the door shoe holder for tech. Every wire, charger, usb key, web cam, airline earbud, everything. It's been beautifully organized for years.

It's been beautifully organized and useless for so many years that it's all junk now. Nothing is using USB A/B. The airline earbuds are as much trash today as they were then. The usb keys are very old and untrustworthy. Even the USB-C wires are all out of spec, won't charge modern phones, etc. The old webcams look like true trash compared to the modern ones. Etc etc

And this post is basically "hoard all this garbage because in 10 years you might need one thing"? I'm sorry, but "buy that one thing for a few bucks off amazon the one time you need it" is looking a lot more attractive than "meticulously maintain a tech hoard to save $3 on a single use wire"

In engineering terms, it feels like my tech hoard is an automation I spent two weeks writing, for code that only ever runs once every ten years.

criley2··on Gemini 3.8 Flash and 3.8 Flash Cyber
On cost per intelligence task, Gemini38flash and Sol56 trade back and forth on cost depending on effort level. https://i.imgur.com/zPaWPXx.png As seen in this image, literally: Sol56 high ranks in between Gemini 38 medium and high. The image proves it.

I also included Sol56 xhigh, which ranks above even Gemini38 high.

criley2··on Gemini 3.8 Flash and 3.8 Flash Cyber
>There are numerous benchmarks that measure cost per task, which factors out tokens entirely. Gemini 3.8 flash is significantly lower than Sol on basically all of them https://artificialanalysis.ai/#cost-tabs

Not sure if you read your own link but Sol 56 high ranks smack between Gemini 3.8 flash medium and high. Gemini 3.8 flash comes in as more expensive per task than Sol 56 high according to artificial analysis.

Luna high is literally 30X cheaper than Gemini 3.8 flash high.

You can limit the model viewer and they're getting better at testing multiple effort levels now: https://artificialanalysis.ai/?models=gpt-5-6-sol-medium%2Cg...

One reason is clear: Sol uses dramatically fewer output tokens than Gemini 38 flash https://artificialanalysis.ai/?models=gemini-3-8-flash%2Cgem...

criley2··on Judge rules Trump administration’s blacklisting of Anthropic was illegal
The government will pay because it's not his money, it's our money. He loves spending our money...
criley2··on The Harness Is the Thing
I'm sorry, but just because you achieve results you consider acceptable with this method doesn't mean everyone does.

I don't work where we can ship slop. I don't work where PRs can be merged based on what the agents say. I work where a human has to read and approve and own every single line of code. I work where the stakes are actually high, so the cost of not using the best tools in terms of human time are big. A single turn around in a PR costs more in human time than the difference between deepseek and fable in API costs.

So, when you admit "There's a huge gap between 'figured out the hard stuff' and 'rock solid'." but then claim that the cheapest/dumbest agent in your arsenal is your go-to for "rock solid", I have to question the quality of your results.

Personally, "using plan mode" is a very 2025 way of using these tools, and I wouldn't be surprised to see "plan mode" be removed from codex/claude code/et al.

Realistically, I'm using the best models to think about a domain and problem (Fable High+), and I'm using a cheap daily driver with an advisor pattern (Opus High + Fable) to iterate through POCs, and I'm using human review to guide design. None of that is "plan mode", it's actual engineering. Then we decompose the solution, we stack it, and we use only really strong agents to build, review and refine.

This obsession with cheap agents leads to low quality outcomes. "Rock solid" deserves the best tools, and the "plan" will never be good enough. I'm going to be sending fable xhigh and sol 56 xhigh et al at it in adversarial review, why the heck am I cheaping out on the actual implementation?

And finally: my time costs way more than any of this. Cheaper models are slower overall and when combined with re-work time, are dramatically slower. I'm costing my company hundreds in my time to save a few bucks on the API bills. Nonsense!

criley2··on Qwen3.8-Flash-Next
It's not free. You're paying electricity and you're ignoring the cost of the hardware. Even on electricity alone, there are cloud providers who may beat your laptop on price per million tokens. Qwen 3.8 flash is interesting in this space.

Not to say that there aren't other benefits of running models locally, I loaded Qwen 3.8 27B 6bit MLX just yesterday.

criley2··on The Harness Is the Thing
Sonnet 5 is the worst model of 2026. Literally just turn effort slider down on Opus, it's smarter, faster and cheaper than whatever Sonnet is.

Beyond that, I find this whole plan and build thing to be a pointless waste of tokens. If your planner made a detailed enough plan, then the cost of executing that plan is a just one turn more of cached tokens, and minimal time.

Meanwhile: switching agents, reloading context and building from the plan will easily balloon your token use and time. And any emergent problem that the dumb executor finds will instantly wreck the implementation because they're not competent at solving it. And if your plan is so perfect that there's no edge case then you're wasting tokens because your planner was one turn away from finishing the project via cached tokens.

criley2··on Qwen3.8-Flash-Next
Those prices are just tokens? Since each model uses different amounts of tokens to do the same thing, it's a misleading price that often makes open-weights look more competitive than they are, since most open weights models use dramatically more tokens and time to complete tasks than many frontier models.

In Artifical Analysis's cost per task, Luna(max) costs $0.05 per task, and Qwen 3.8 27B costs $0.25 per task, a 5X increase. We'll see how 3.8-flash-next does.

criley2··on I should have loved biology
I absolutely experienced this in college. I signed up as a computer science student, as one does. I took all of the freshmen classes across broad topics, and the first biology class was basically just like the article describes. Words and memorization and microbiology and boring stuff. Who cares, let's get back to the dorm room to play CounterStrike on the school LAN.

But then I took the second required biology course. And rather than be about memorizing the krebs cycle, it was macrobiology. Ecology, evolution, and everything I hadn't realized I wanted to learn about the world around me. I instantly fell in love and 99'd the class. I changed my major to Biology, went pretty heavy into biology and chemistry, and didn't take more than a few Java classes in Comp Sci. I graduated with a B.S. in Biology.

I'm still a software developer, mind you, but I wouldn't trade the Biology degree for anything. Learning statistics, experimental design and experimentation, and just having the opportunity to deep dive into the physical world from the biggest to the smallest details, it was life-changing.

I still ended up learning data models, algorithms, and probably would be a better engineer if I had dug in deep on comp sci, but I like to think the biology degree helped me be a better developer for different reasons.

criley2··on Accelerating GPT-5.6 Sol Ultrafast
I append this to many of my opus claude code prompts

`You may use a Fable subagent to answer questions, solve problems, and provide an adversarial review of your ideas and code`

You can use a similar pattern in most any harness, and you can tell them to use other harnesses. In claude you can write `Use codex cli to have Sol56 Xhigh provide an adversarial review to your plan before presenting it to me` or `Use opencode cli with GLM 5.3 to verify all code review findings before presenting` or whatever you're doing, as long as those other tools are setup and ready to be called.

IMO: This isn't useful as a token saving pattern in my experience with agentic engineering, but it is useful as a quality-enhancer.

criley2··on GLM-5.3: Frontier coding with emergent cyber capabilities
It's pretty easy for the US to functionally ban chinese models. They only have to target US firms like inference providers or the biggest users, and pretty much the whole domestic market will fall into line. They don't actually care about the final few %.

Regardless of whether or not adversaries are using them, the US has by far the most compute available, and we've now hit the line where major providers are no longer releasing their best models. The public gets the "current" level of intelligence, while the US government gets to control access to the actual frontier of non-public AI. From their perspective, their enemies using GLM5.3 while they have GPT6 and Mythos6 or whatever is a fine trade.

I don't support a ban at all, nor the US's behavior, I'm just pointing out some facts that change the argument.

criley2··on What sort of maths are LLMs good at?
I totally agree - designing a competent AI agent with a fully customized harness to successfully pull off this task is a much more challenging engineering effort than merely creating an ordinary computer program. Had OP made chatgpt write an ordinary program instead, they likely would have succeeded in their task.
criley2··on What sort of maths are LLMs good at?
What a sloppy reply. You've hijacked a thread on mathematics first to complain that your incompetent attempt to use ChatGPT to find a job failed, but it seems now that this was a ruse to instead begin arguments unrelated to the article at all where you just spam arxiv links you've never read to "prove" that AI is a scam.

This comes across, frankly, as either Dunning-Kruger (classic illusory superiority), or potentially as mental illness. The slop dump is highly reminiscent of how a schizophrenic friend of mine communicates.

Do you really think slopping down a bunch of random arxiv links "proves" that AI is a scam and you're so smart and everyone else isn't?

Most awkwardly for your arxiv slop -- most of this is irrelevant to your central claim, and you've missed papers that are much closer.

For example your LogicGraph paper: "Can't exhaustively enumerate all minimal proofs" is not "can't distinguish Ireland from London".

Or your "Do VLMs Understand 3D Scenes..." is nothing more than citation decoration, completely irrelevant to our discussion.

Or your "Frontier LLMs Still Struggle with Simple Reasoning Tasks" which is potentially your pièce de résistance, it supports brittle multi-step constraint handling, but isn't remotely an eval of a modern web-search agent.

For example, VibeSearchBench would have been far more relevant to your claims https://arxiv.org/html/2605.27882v1 (but still obviously not proof that AI is "a parlour trick")

Going further: my point that we need to discuss your beginner's approach to the harness is substantiated clearly here: https://arxiv.org/html/2605.23950v1

Finally, failure to exhibit human-like generality is not evidence of absence of intelligence. It is evidence that whatever cognitive machinery LLMs possess has a very different error distribution from ours. Your General365, LLMEval-Logic and the Reversal Curse are actually fascinating evidence for that jaggedness, rather than proof of your claim that AI is a scam.

criley2··on “Code was never the hard part” is an insult to all programmers
A business does need a small number of their most senior engineers doing high altitude work that can, at times, include helping sales estimate new features. But in my experience, it's not rocket science and a good product team can do this on their own with a quick async check over chat. At most, a single meeting is all it takes.

I've heard of Sales Engineers as well, embedding programmers directly with sales teams.

But the vast majority of programmers should not be spending any significant amount of their time on this. Their value is in building and scaling well-specified systems.

Page 1 of 34Next →