HNHacker News
TopNewBestAskShowJobs

boc

2,331 karma · joined March 31, 2020

turquoise hexagon sun
submissionscomments
boc··on Pacing the Frontier is not the actual goal for AI labs
> I still remember the day I found out game theory was a thing - it was like finding out that someone successfully systematized common sense.

Game Theory as a field is actually very complex and often reveals dominant strategies that would never occur to someone as the "common sense" approach to a problem. It was arguably created to solve for problems where there was no obvious correct answer, even for a very intelligent person.

boc··on Sonnet 5.5
I was talking about this with a friend this weekend. We both work in the field and test new models within minutes of them being released. We both immediately clocked Opus 5.5 as being cracked within the first hour. Went on HN and the launch announcement was full of people whining and pointing at cost/token charts vs Chinese models. It was like the upside-down world.

We were both sad that HN has become a negative signal news source on AI lately - you're much more likely to be misled by this website in 2026 on the topic of frontier AI. If you're reading this comment, you should do your own research vs trusting the "Astra is 1000% the best" or "Deepseek is the $/tk KING" comments swarming these announcement posts.

boc··on Classified estimates show the NSA is paying billions to test AI models
Yeah men with guns can always take it, but there's this thing called the stock market and bond market that would freak out if free enterprise was crushed by state coercion with armed men, with billions worth of IP forcibly exfiltrated by state actors. Trillions get wiped out overnight.

That massive economy is what pays the guys with guns. You can win that battle but you will absolutely kill the golden goose and lose the war.

> and can't seem to fathom that men with guns always override the opinions of nerds in SF.

Those nerds in SF are holding the proverbial gun up to the head of the economy right now. Maybe men with physical guns are bad at understand that type of thing.

boc··on Grok 4.7
The reason I did it the way I did was so that I can still use the codex and claude code subscriptions vs paying the API cost. Can you do that via OpenCode?
boc··on Claude Opus 5.5
So far in the past 20 minutes it sounds much better in my sessions. Way better than 5.0 so far.
boc··on Grok 4.7
I've been getting a ton done with Fable as the supervisor and astra as the implementer, with opus for adversarial reviews of the astra PRs. You can use terminal multiplexers with custom harnesses to allow Fable to start codex sessions and send instructions / read instructions / allow/deny actions. It's pretty cool!
boc··on US confirms for first time it has deployed space weapons
So China's brand new recon satellite that was placed specifically in orbit over Iran just happened to self-destruct weeks after China was accused of giving Iran high-res satellite imagery that was used to kill American soldiers?

And the US just publicly told everyone they have weapons in space?

Hell of a coincidence.

boc··on Claude Fable 5.1 and Claude Mythos 5.1
Small reminder that the US government rug-pulled Fable, not Dario. Lots of the safety guards that users find annoying/objectionable were the results of negotiations to get the model back online after the US government forced them to take it down.

Maybe Dario should have just "donated" $1M to Trump's inauguration fund like Altman, Meta, Amazon, Microsoft, Tim Cook, Elon, and Google. There's a reason they are the odd man out with this current Administration.

boc··on Claudette: Make Claude stop talking like a BuzzFeed article
I set my env up so I can see the exact context used in CC CLI, and then once I get over about 40% ctx used I have it handoff to a new, fresh session. Nothing good comes from running above, say 60% of your context window. Coincidently, I usually have good results with CC. I never compact a session ever.
boc··on Norway should buy OpenAI
That assumes they release their models publicly. The future is leaning towards these labs air gapping their best stuff (Model 2, etc) and using it internally to snipe their competitors and charge insane amounts for monitored use in consulting environments.

You can't distill or catch up if you can't access the models. You'll basically have a situation where nation-states will need to try and steal the models Oceans 11 style.

boc··on Qwen3.8 27B scores 52 on Artificial Analysis
> More importantly, what's the point of building monster-scale data centers on unprecedented amounts of debt when a more than good enough model runs on a GPU from a couple years ago?

Probably because the future "monster" models will be insane. 100T+ param models might be the type of things that can independently run a small business, which means anyone not using them is at a distinct disadvantage to their competitors.

The top model from 2025 looks silly compared to the top model of the first half of 2026. Do you feel like progress has stalled?

boc··on A Preview of DuckDB v2.0
Ducklake supports postgres for the catalog, so you get the postgres concurrency benefits + duckdb engine to read the parquet files in the bucket.
boc··on Going Dark, and the era of law enforcement hacking
It's not 2025 anymore my friend.
boc··on U.S. Department of Energy Launches the Genesis Open Models Initiative
Chinese open weight models are great for this turn, but American private models generate orders of magnitude more cashflow. This cashflow = investment in training future models. It's unclear how Chinese open weight companies are going to compete in future rounds if they can't raise the same capital for training runs.

The American business model is exceedingly efficient at building large businesses from zero. I wouldn't dismiss it as just a jobs creation thing.

boc··on Poles of Inaccessibility in the San Gabriel Mountains (2015)
Kansas genuinely has a ton of paved and gravel roads all across the state. It's like a gigantic grid. Oddly high amount of street-view coverage too on google maps: https://maps.app.goo.gl/EyBFj9BNQa8nhUcE6

The Kansas landscape is very underrated. Just infinite sky in all directions.

boc··on Americans are rallying against data centers. Surprisingly few are getting built
One under-discussed topic is whether Americans hate data centers in their neighborhoods, or whether they actually just hate gas turbine generators running 24/7 in their neighborhoods.
boc··on Waymo in Dallas
You could argue that cheaper/safer local rideshare would boost traffic to local bars and restaurants in denser areas though. If you want to keep the $ local, just add an AV tax that goes to the city.
boc··on Decathlon Germany adds Wero payment option to decathlon.de website
> Here (NL) contactless payments works with every debit card, physical or stored Apple Google/Pay, in every store

That's how it works in the US as well. You might be conflating Apple Pay with Apple Wallet, which stores your credit cards for tap-to-pay. Most people in the US use Apple Wallet/Google Wallet to digitally store their credit/debit cards, and use those cards to pay at the store by tapping the POS terminal.

For example, I use Apple Wallet for nearly every transaction, but I never use Apple Pay/Apple Card as the payment method.

boc··on Ask HN: What's the best AI coding tool today?
cmux + Claude Code / Codex with a custom agents.md file.

With cmux your agents can access the browser, spawn full shells, and you aren't locked-in to a specific vendor. It's really powerful but you need to update your instruction file so your coding harness knows it can use the cmux tools to test.

You can also download Handy to use voice so you can just give your feedback via voice into the CLI during the planning phase of a project.

boc··on Claude Opus 5
Seems really good so far using it in Claude Code CLI - it gave me a new flag when I asked a question:

"I don't have a reliable way to read that number, so I'd be guessing if I gave you one — and this is exactly the kind of question where a confident guess is worse than none.

What I can tell you is what I actually observe:"

I really like this update - gave me a clear sense of the facts but didn't give me a guess just for the sake of guessing.

One oddity is that it appears to only have a 200K context window right now via CC. Hopefully the 1M version will appear soon!

boc··on Show HN: Claude-thermos keeps your Claude session warm for you
Interested in how the critics of approaches like this defend an agentic session (with Fable, for example) that stops and runs a multi-hour ML training session. It's a script, so the actual LLM convo goes stale, but then when the results get returned to the main thread you get an expensive cache hit without doing anything.

You would have avoided that cache hit if the LLM session was kept "alive" for those few hours. Why not automate the part where you keep the large main thread alive until you're ready to analyze the results?

boc··on Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
The fun part about this is that we can see who is right in about a year. If the leading labs continue making progress at hardening their models against distillation, and then they start pulling away again, we see who is right. If China is able to pass the US and release an independently better model than anything the US has, then your theory is correct.

Both sides have extremely smart people. One side has more $$$ and exclusive access to the best chips. For progress to converge without a corresponding breakthrough suggests there's something else at work.

boc··on The creepiest 'sales demo' of all time
You're making the point even stronger for US companies to just block EU traffic entirely. If you have no EU customers, the mere act of letting EU visitors on your site might cause an EU country to "make an example" of you for incorrect data handling, but the alternative of just blocking all traffic would let you travel to the EU without issues.
boc··on Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
> US companies models constantly distill each other as Musk was forced to admit under oath

> This whole narrative has just been a massive cope.

So wait, US AI companies all use distillation because... it's not effective and it's all just cope? Or is distillation really powerful and they all do it, which Musk was forced to admit under oath? But when China does distillation it isn't powerful and they don't need to do it, but they do it anyway because it's fun?

Either it's powerful and everyone, including the Chinese labs, use it as a way to rapidly catch-up against the SOTA models, or it's a red herring and the huge amounts of energy spent to protect and enable distillation is all just wasted money. Which is it?

boc··on Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
Or option three is they are drafting hard off the frontier US models via distillation.
boc··on Mawlynnong, India, transformed by tourism, bans visitors on Sundays
American recycling in a lot of major cities is single-stream - aka you put all recycling together and a central plant sorts it for you. More efficient, more accurate, and it encourages more people to recycle since it's extremely easy.

https://en.wikipedia.org/wiki/Single-stream_recycling

boc··on The classifiers Anthropic puts in front of Fable are too zealous
Missed this whole discussion today because I was here in California, doing a ton of productive work, using Fable.

I think you guys are working yourself into a lather about this topic while other people are quietly getting a ton of shit done with Anthropic's models.

boc··on Claude Sonnet 5
So you're arguing it's the Yogi Berra "Nobody goes there anymore, it's too crowded" of business models?
boc··on Claude Sonnet 5
That's not how economies work. It's not a fixed pie that everyone shares.
boc··on Claude Sonnet 5
> Anthropic has little to no chance of producing a competitive business model in the long term.

Extraordinary thing to say about the fastest growing company in the history of capitalism. They will soon have access to public markets, essentially unlimited capital, and can build insanely large models that they don't have to make public... ever. They can just use those models to run their business, train better models, eat competitors, etc.

But maybe it's Anthropic that isn't thinking ahead enough - you clearly think you can see around corners with your proclamation. So why do you think they have "little to no chance" of surviving long term?

Page 1 of 14Next →