HNHacker News
TopNewBestAskShowJobs

mnicky

195 karma · joined July 4, 2011

submissionscomments
mnicky··on Siri AI
I recommend mapy.com (mobile app and web app too when on computer) - they mostly use OSM data and rendering of map tiles is great. Also offline maps etc.
mnicky··on Anthropic confidentially submits draft S-1 to the SEC
I believe it's the opposite :) All major indices (S&P500, MSCI, FTSE...) use free-float adjustments. And recently also NASDAQ - they've changed to cap of 3x the value of free-floating shares.
mnicky··on Anthropic confidentially submits draft S-1 to the SEC
Good thing is that index funds don't hold stocks at market capitalization but only at free float value. So a company whose shares are mostly held by founders, employees, and strategic investors gets a weight well below its headline valuation.
mnicky··on SpaceX S-1
Time to invest in some value-stocks index funds then :)
mnicky··on Gemini 3.5 Flash
E.g. because they are behind on research and so must compensate with size to achieve similar level of intelligence. At least this is what I heard.

For intelligence/size only OpenAI and Anthropic are the frontier. Google has more compute so it can compensate for that with size of the models...

mnicky··on How Claude Code works in large codebases
And not only startups...
mnicky··on Elevated error rates on Opus 4.7
Hmm, today's pass rate raised to 73% - interesting, are they AB-testing some new model? This is too high for Opus 4.7.
mnicky··on Hardening Firefox with Claude Mythos Preview
> Zoom also won't let you join with browser on Firefox.

FF works for me.

mnicky··on Accelerating Gemma 4: faster inference with multi-token prediction drafters
In the Dwarkesh's podcast Dylan Patel from SemiAnalysis said that Google can currently afford to have larger models than competitors, because of access to much more compute, TPUs etc.

That could explain the token usage difference because larger models usually use less tokens per the same unit of intelligence.

mnicky··on America's Expanding Domestic Surveillance
Yeah. I really like the main idea behind GDPR, which is that data containing PII is the property of the person it describes, not of the companies that process the data to provide services.

This means that I, as the owner of my data, can refuse to provide it for some use cases, request its deletion, etc. It’s my data after all.

mnicky··on AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights
Wasn't it this one?

Article: https://alignment.anthropic.com/2025/subliminal-learning/

Paper: https://arxiv.org/abs/2507.14805

mnicky··on Claude Opus 4.7
> things that used to be cached for 1 hour now only being cached for 5 min.

Doesn't this only apply to subagents, which don't have much long-time context anyway?

mnicky··on Claude Opus 4.7
Maybe this? From the article:

> Opus 4.7 is substantially better at following instructions. Interestingly, this means that prompts written for earlier models can sometimes now produce unexpected results: where previous models interpreted instructions loosely or skipped parts entirely, Opus 4.7 takes the instructions literally. Users should re-tune their prompts and harnesses accordingly.

mnicky··on Small models also found the vulnerabilities that Mythos found
Also, what is $20,000 today can be $2000 next year. Or $20...

See e.g. https://epoch.ai/data-insights/llm-inference-price-trends/

mnicky··on Gold overtakes U.S. Treasuries as the largest foreign reserve asset
This might be unconstitutional?
mnicky··on Gold overtakes U.S. Treasuries as the largest foreign reserve asset
Averages tell nothing about an average citizen.

Also, there are other measurements like inequality, healthcare cost, social securities...

mnicky··on GPT-5.4
This observation makes sense, because all models currently probably use some kind of a sparse attention architecture.

So the closer the two related pieces of information are to each other in the input context, the larger the chance their relationship will be preserved.

mnicky··on Cancel ChatGPT AI boycott surges after OpenAI pentagon military deal
He's trying to make it sound so, but in legal domain, devil lies in the details.

It seems that government wanted to use Claude for mass analysis of commercially obtained data on American people and Anthropic wouldn't let them (source: https://www.theatlantic.com/technology/2026/03/inside-anthro... ).

DoD kept asking for changes of contract where at least the legalese would be changed to somewhat more permissive but Anthropic stayed their ground.

Sam Altman probably let them do that, while using language like "we have technical means of oversight and the same red lines as Anthropic". But in reality they will allow DoD to do what Anthropic didn't.

See this for more information: https://www.lesswrong.com/posts/PBrggrw4mhgbksoYY/a-tale-of-...

mnicky··on How I use Claude Code: Separation of planning and execution
> Very often, after a correction, it will focus a lot on the correction itself making for weird-sounding/confusing statements in commit messages and comments.

I've experienced that too. Usually when I request correction, I add something like "Include only production level comments, (not changes)". Recently I also added special instruction for this to CLAUDE.md.

mnicky··on How I use Claude Code: Separation of planning and execution
Since some time, Claude Codes's plan mode also writes file with a plan that you could probably edit etc. It's located in ~/.claude/plans/ for me. Actually, there's whole history of plans there.

I sometimes reference some of them to build context, e.g. after few unsuccessful tries to implement something, so that Claude doesn't try the same thing again.

mnicky··on GPT‑5.3‑Codex‑Spark
Can you compare it to Opus 4.6 with thinking disabled? It seems to have very impressive benchmark scores. Could also be pretty fast.
mnicky··on GPT‑5.3‑Codex‑Spark
> What am I missing?

Largest production capacity maybe?

Also, market demand will be so high that every player's chips will be sold out.

mnicky··on Gemini 3 Deep Think
Well, fair comparison would be with GPT-5.x Pro, which is the same class of a model as Gemini Deep Think.
mnicky··on Gemini 3 Deep Think
> can a sufficiently large non thinking model perform the same as a smaller thinking?

Models from Anthropic have always been excellent at this. See e.g. https://imgur.com/a/EwW9H6q (top-left Opus 4.6 is without thinking).

mnicky··on GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
It is possible to think of tokens as some proxy for thinking space. At least reasoning tokens work like this.

Dollar/watt are not public and time has confounders like hardware.

mnicky··on Claude Code is being dumbed down?
At least now we also have a tracker: https://marginlab.ai/trackers/claude-code/
mnicky··on GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
What I haven't seen discussed anywhere so far is how big a lead Anthropic seems to have in intelligence per output token, e.g. if you look at [1].

We already know that intelligence scales with the log of tokens used for reasoning, but Anthropic seems to have much more powerful non-reasoning models than its competitors.

I read somewhere that they have a policy of not advancing capabilities too much, so could it be that they are sandbagging and releasing models with artificially capped reasoning to be at a similar level to their competitors?

How do you read this?

[1] https://imgur.com/a/EwW9H6q

mnicky··on GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
> I think GPT-5.3-Codex was a disappointment

Care to elaborate more?

mnicky··on Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
Evaluation than depends on your specific cost-benefit tradeoff of accuracy vs hallucinations.

For some tasks where detecting hallucinations is easy I can see it being beneficial.

In general case not so much...

mnicky··on Frontier AI agents violate ethical constraints 30–50% of time, pressured by KPIs
If you recall the context/situation at the time it was released, that might be close to the truth. Google desperately needed to show competency in improving Gemini capabilities, and other considerations could have been assigned lower priority.

So they could have paid a price in “model welfare” and released an LLM very eager to deliver.

It also shows in AA-Omniscience Hallucination Rate benchmark where Gemini has 88%, the worst from frontier models.

← PreviousPage 3 of 5Next →