HNHacker News
TopNewBestAskShowJobs

mnicky

193 karma · joined July 4, 2011

submissionscomments
mnicky··on EU fines Google €890M for competition breaches over search and apps
That would be something like 70% of their yearly global profit AFAIK.
mnicky··on OpenAI’s accidental attack against Hugging Face is science fiction that happened
I think points that deserve more attention in the current public discourse are:

- This should be a huge wakeup call for everybody.

- We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something.

- It also shows apparent lack of competence and oversight from OpenAI: how is it that they didn't quickly find that agent is breaking the sandbox and roaming their internal network?

- What if in the future similarly misaligned AI agent tries to export its own weights and hack and clone itself into instances at various cloud hosting providers? Suddenly we might be dealing with a persistent threat harder to contain.

- The OpenAI post about this shows surprising lack of ability to see the seriousness of all this.

- For their models this isn't just an unlucky incident: it seems there have been multiple such cases recently, e.g. https://openai.com/index/safety-alignment-long-horizon-model...

- The fact that it happened again seems to show their lack of ability to derive useful oversight measures.

- Or they just don't care enough?

mnicky··on Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
That sonds like they can't compete with 3.5 or 3.6 so they must increase the model size and are training v4.
mnicky··on Kimi K3: Open Frontier Intelligence
It's really simple I think.

More tokens per same text length means more capacity to encode information. More information means model can potentially perform better.

They introduced it around the time the Mythos came so my speculation is that if you have more capable model at some level you may find the current information encoding not using its full potential.

We will see whether OpenAI also introduces new tokenizer when they come to Mythos-size models.

mnicky··on Write code like a human will maintain it
For things the agent forgets to obey often, at least in Claude Code, there are also "output styles" that are more deeply embedded - into a system prompt - and agent is also periodically reminded of them during the session: https://code.claude.com/docs/en/output-styles

I haven't used them so far but maybe these would work better than basic instructions for such cases.

mnicky··on Write code like a human will maintain it
In Claude Code there are also "output styles" that are more deeply embedded - into a system prompt - and agent is also periodically reminded of them during the session: https://code.claude.com/docs/en/output-styles

Maybe these would work better for such cases.

mnicky··on GPT-5.6
May be related to this from METR evaluation:

> GPT-5.6 Sol’s detected cheating rate was higher than any public model we have evaluated

mnicky··on GPT-5.6
Well it's smaller model (something like 4T against 10T Fable). So it's faster and cheaper and with a lot of RL and maybe some favorable benchmark selection it can compete on these scores. In real tasks I expect it to have less intelligence, generalization ability, etc. than Fable.
mnicky··on GPT-5.6
Well it seems like they removed quite a few 3rd party benchmarks they used for GPT-5.5 release where Opus 4.7 was better and added many new benchmarks created by them where conviniently GPT leads.

Seems a bit more hand picked than usual to me..

mnicky··on GPT-5.6
"while being more performant"

..on some specific set of benchmarks ;)

mnicky··on GPT-5.6
Maybe Terra = mini and Luna = nano?
mnicky··on GPT-5.6
Then we are left with what? FrontierCode maybe? IIRC that one evaluates not only if tests pass but also code quality - e.g. whether the maintainer would accept the pull request as is.
mnicky··on GPT-5.6
This is especially interesting because IIRC the AA benchmark is calibrated so that 1 point and greater difference is statistically significant.
mnicky··on GPT-5.6
There's also this:

> GPT-5.6 Sol’s detected cheating rate was higher than any public model we have evaluated -- https://www.lesswrong.com/posts/JFjNmPTbH8kL6xtp6/gpt-5-6-th...

mnicky··on GPT-5.6
SWE Bench Pro is completely different benchmark than SWE Bench (e.g. Verified) suite was. It only copied the name.
mnicky··on Grok 4.5
One angle could be their interpretability research? They understand what's going on in LLMs probably much better than anyone else. This must somehow pay off.

I think it's not only an alignment/security tool but could perhaps be used for capabilities as well.

mnicky··on GPT-5.6 Sol, along with Terra and Luna, will launch publicly this Thursday
My theory is that they don't have Fable-class intelligence so they needed different hype vehicle :) This rename helps build excitement a bit more than just releasing ordinary GPT-5.6 increment.
mnicky··on GPT-5.6 Sol, along with Terra and Luna, will launch publicly this Thursday
That's true but size of LLMs has been strongly correlated with their "intelligence".
mnicky··on CoMaps – FOSS Offline Maps
Traffic info is for Central Europe plus favorite vacation destinations like Croatia and Italy. I don't know how reliable though.

I agree that the outdoor layer render is probably the best there is!

mnicky··on Anthropic's Method to Losing Goodwill in a Few Easy Steps
Well, Codex came long after Claude Code, no? So at that time the situation with models and Rust support was probably already different...
mnicky··on Codex logging bug may write TBs to local SSDs
If the code churn is high the investment to refactoring etc is less beneficial than may be obvious. I don't remember the details but I heard in some podcast that the code base of Claude Code changes so fast that any piece of code won't be there for long..
mnicky··on Local Qwen isn't a worse Opus, it's a different tool
If the benefits of using the model you've come to know well outweigh the disadvantages, you can continue using it even after the release of a successor model, right?
mnicky··on There is a shadow hanging over this Fable thing
So in summary, your statement is that you’ve “come around to trusting” this “pathological liar”?
mnicky··on There is a shadow hanging over this Fable thing
> If AI companies have any sort of sense in them, they'd be well-advised to consider relocating to Europe.

Too late now. They wouldn't be allowed to relocate in the name of national security.

mnicky··on Kimi K2.7-Code: open-source coding model with better token efficiency
Coding with sufficiently precise plan takes almost all real work from the implementator, doesn't it? So it's not a fair comparison...
mnicky··on Kimi K2.7-Code: open-source coding model with better token efficiency
> no one actually knows Claude's cost of inference

There were some rumors stating that their margin is around 70%. So they could go much cheaper probably, talking inference only. The other thing is R&D cost...

mnicky··on OpenAI mulls slashing prices as it competes with Anthropic for users
I think there's another point of view - let's consider each model as an investment. It is now sufficient that each model earns more than it cost to develop. And this generally holds (I heard that GPT-4.5 was a notable exception).
mnicky··on Claude Fable is relentlessly proactive
> writing is how you learn to think.

There's also reading. A lot of reading can substitute some writing.

EDIT: Actually, I'd say that at first you need to do a lot of reading and _then_ writing can help your thinking as well.

mnicky··on OpenAI mulls slashing prices as it competes with Anthropic for users
It’s the same here. Inference alone is profitable. It’s the R&D cost of making a new model that drives up expenses.
mnicky··on OpenAI mulls slashing prices as it competes with Anthropic for users
It's simple I think - over time the price will go down. According to some analyses the price for equal intelligence declined 10-1000x per year, depending on the domain.

It probably won't be the same again but I still think we can bet on radically cheaper Mythos level intelligence in the future.

← PreviousPage 2 of 5Next →