HNHacker News
TopNewBestAskShowJobs

oh_no

88 karma · joined March 23, 2026

submissionscomments
oh_no··on Gemini 4 Argon
i believe they've got inference issue. the phone app never got past 3.6, and they're going to roll this out to Ultra subscribers before other subscribers. none of this screams they're ready to flip a switch and start serving a ton of traffic as OpenAI/Anthropic routinely do.

nothing about this announcement gives me confidence that google is back on track as a model provider.

oh_no··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
3.8 flash is more expensive than sol 6.0, google uses a LOT of reasoning tokens
oh_no··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Sol 6 is in there? You may be on Enterprise where it didn't roll out by default and comes out in a week or so. (Which is a weird and bad change to their model releases.)
oh_no··on Sonnet 5.5
i'd love to see them re-enter that space but given haiku 5 never happened I wouldn't bet on it

i think they see what openai charges for luna and just don't want to try and compete

oh_no··on Sonnet 5.5
it could be that, it could also be that sonnet max looks to burn about 60% more tokens than opus max

AA intelegence index (agent harness doesn't have sonnet data yet) on max: Astra 27k Fable 5.1 78k (Sonnet 5) 118k Opus 5.5 119k Sonnet 5.5 193k

Opus 5 was previous record holder so hats off to Anthropic on blowing it away on token churn.

oh_no··on Microsoft abandons personal AI chatbot race with Copilot reboot
microsoft Product remains an unmitigated dumpster fire.

i think github copilot has actually gotten to a pretty good place, and of course the two copilot products will both have features called "autopilot" that do entirely different things

oh_no··on GPT-6 Sol and Luna
look at token use, 3.8 flash is a huge token hog compared to openai models
oh_no··on OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005
Why do you think this uncracked code was so simple to solve?

And it's very possible Terra or GLM could crack it, turn off their web access and try yourself.

oh_no··on Grok 4.7
shutoff for existing models, new models stopped as of that announcement, astra will never be on cursor.
oh_no··on Grok 4.7
which is crazy because this was grok's competitive advantage, worse than OpenAI models but better than everything else, now it's less efficient than Opus or Fable 5.1
oh_no··on Grok 4.7
the AA numbers are generationally bad. double token use (the one thing Grok was good at was low reasoning usage!) to gain 5% in the benchmark score. with reportedly a larger model. maybe it shows gains IRL but wow, I've never seen a new generation model look so underwhelming compared to the last.
oh_no··on Ask HN: Are others seeing Google's reCAPTCHA rejecting Firefox users?
It still works, bots can solve it but it probably increases the cost of that web call by 10x or 100x for that bot, so it won't bother. Had a recent bad experience with removing recaptcha.
oh_no··on Ask HN: Are others seeing Google's reCAPTCHA rejecting Firefox users?
I get this pretty frequently on windows Firefox after switching to it, note this is my work computer, Firefox works fine at home on more open network
oh_no··on GPT-6 Astra
Very nice to see that this is even more token efficient than Sol, when Fable 5.1 is less so than the already bloated token budget of Fable 5.
oh_no··on OpenAI begins rolling out GPT-6 Astra
that's all openai models but i'm very happy openai continues to focus on efficiency rather than reasoningtokenmaxxing
oh_no··on Nvidia to acquire Hugging Face
the last thing nvidia would want is to make ai look more legally risky

to say nothing of their massive investments direct and indirect in openai

oh_no··on Path to Astra: critical capabilities and frontier safeguards
with Fable 5.1 increasing token use pretty dramatically I'm again impressed that OpenAI seems like the only lab to be driving token use down. The ExploitBench Internal Port chart showing token usage is crazy impressive
oh_no··on Our decision on Cursor following its acquisition by SpaceX
you dont like giving openai money but you're fine with spacex? not criticizing, just trying to understand.
oh_no··on Our decision on Cursor following its acquisition by SpaceX
cursor adds $0.25/M to your token bill for 3rd party services. sounds small but it's insanely big on cached inputs, which are an insanely high % of use
oh_no··on Coding expertise is going to collapse from AI reliance
1. I think LLMs will end up pretty dramatically shrinking MTTR, a lot of that will be tooling to proactively resolve problems as soon as they start, but a lot of it is that agents are very good at finding and fixing problems by correlating data, something humans can do but I think we're going to be slower at 2. "The models are down" doesn't seem like a very realistic problem unless you have a single provider/set of endpoints. I do not recommend this setup but I guess if you have a single point of failure that is a risk.
oh_no··on OpenAI: GPT 5.6 Sol price reduction (until at least Nov 21)
very easy to lose money on subscription, very easy to make money on api pricing
oh_no··on OpenAI: GPT 5.6 Sol price reduction (until at least Nov 21)
what? enterprise LLM API contracts pay listed model rates. how could they lock you into pricing on models that they've not developed yet?

enterprise SaaS LLM calls do lock in rates but don't lock in models for similar reasons to the above.

oh_no··on Claude Code May–August 2026 weekly limits promotion
yes and no, anthropic and openai are losing money on people who max out their sub, but openai has a lot more room to play with with much cheaper models to serve (by all signs we have from actual api/task pricing)
oh_no··on Gemini 3.7 Flash
5.6 Luna costs far less and benchmarks far better, have you compared for this task?
oh_no··on Managing AI Coding Costs at Scale
buddy, they're on enterprise plans paying per token
oh_no··on Beating GPT-5.6 Sol on retrieval with 100x cheaper open models
i'm a little confused by this, should most piles of company documentation look pretty similar? what are we getting by tuning at the org level?

what if you have a bunch of teams or apps that have different documentation patterns?

how much does this degrade over time, it beats leading models with that static data set but clearly this edge will degrade with data drift, how quickly does that happen?

also, any "this is 100x cheaper" blogpost means nothing if not discussing TCO (I know your team didn't write this.) I don't care what inference costs are if I don't know training/overhead costs. what's the breakeven point. and again, how long is this RAG stack going to be worth keeping, you beat 5.6 Luna but at some point un-tuned models will beat you, so this is a temporary solution that needs to be re-upped at some point. benchmarks against data drift would help there

oh_no··on Thomson Reuters built its own AI model that now ranks among the best
few days old but want to flag that there are zero apples:apples comparisons on this press release. set aside the benchmark vs older models they're testing their model with access to their internal legal research data vs open internet search on other models (and reasoning turned off on 5.5 for some reason?)

a bad product and an embarrassing press release

oh_no··on Advancing the price-performance frontier with GPT‑5.6
where are you seeing cheap Kimi? pricing I've seen is the same across the board (presumably due to licensing terms) and is in the Terra range.
oh_no··on Advancing the price-performance frontier with GPT‑5.6
I pretty strongly disagree about comparing this to Kimi and GLM, 5.2 was a big price hike for Chinese models, and Kimi K3 was a big price hike to that. K3 was within spitting distance of OpenAI pricing (more expensive than short context Terra, less than long context). And that's after months of OpenAI/Anthropic prices going up.

Now we have an American lab drastically cutting a price, feels like this is the opposite of that trend.

oh_no··on The coolest use for the Vision Pro
reading some of the comments i was expecting something really high end and polished, but yeah, this is something you could with a headset 1/10th the cost
Page 1 of 2Next →