HNHacker News
TopNewBestAskShowJobs

mutkach

82 karma · joined July 31, 2025

https://torus.graphics
submissionscomments
mutkach··on Germany is not supporting ChatControl – blocking minority secured
Why would you really need something like that in a non-totalitarian state? Basically, it follows the russian playbook (essentially the same 'language' - safety concerns), but instead of the FSB, who is the beneficiary actor in this case?
mutkach··on Where's the shovelware? Why AI coding claims don't add up
> 55,000 hours of [IT] experience

I didn't know Mike Judge was such a polymath!

mutkach··on Updates to Consumer Terms and Privacy Policy
How would they not "share information with third parties". You need to sift through the data to make it even remotely useful for "training". You absolutely need to share it with either Amazon (for Mechanical Turk) or with Scale AI.

I am wondering how would you use a chat transcript for training? Unless it is massive, possibly private codebases that are constantly getting piped into Claude Code right now. In that case, that would make sense.

mutkach··on Are OpenAI and Anthropic losing money on inference?
I am implying that what OpenAI pays for GPU/hour is much less than $2, because of the discount. That's an assumption. It could be $1, $0.5, no?

It could still be burning money for Microsoft/Amazon

mutkach··on Are OpenAI and Anthropic losing money on inference?
Gross margins also don't tell the whole story, we don't know how much Azure and Amazon charge for the infrastructure and we have reasons to believe they are selling it at a massive discount (Microsoft definitely does that, as follows from their agreement with OpenAI). They get the model, OpenAI gets discounted infra.
mutkach··on Are OpenAI and Anthropic losing money on inference?
A full KV-cache is quite big compared to the weights of the model (depending on the context size), that should be a factor too (and basically you need to maintain a separate KV cache for each request, I think...). Also the the token/s is not uniform across the request and it's getting slower with each subsequent generated token.

On the other side, there's an insane booster of speculative decoding, that would give a semi-prefill rate for decoding, but the memory pressure is still a factor.

I would be happy to be corrected regarding both factors.

mutkach··on Are OpenAI and Anthropic losing money on inference?
There's also an estimation of how much a KV cache grows with each subsequent token. That would be roughly ~MBs/token. I think that would be the bottleneck
mutkach··on IQ tests results for AI
Exactly. Out of all possible interpretations all of them all of them kinda converged to the same conclusion right at the beginning of the "reasoning", isn't that weird. They absolutely did train the models either on existing examples or came up with their own pre-labeled datasets. Benchmaxxing is way out of control, something must be done about that
mutkach··on IQ tests results for AI
Judging from the reasoning trace for the problem of the day - almost all of the models obviously had some presence of IQ training data or at least it could be said that the models are very biased in a beneficial way. From the beginning of the trace you kinda see that the model had already "figured it out" - the reasoning is done only for applying the basic arithmetics.

None of the models did actually "reason" about what the problem could possibly be - like none of them considered that more intricate patterns are possible in a 3x3 grid (having taken this kinds of test earlier in life, I still had a few seconds of indecision, thinking whether this is the same kind of test that I've seen and not some more elaborate one), and none of them tried solving the problem column-wise (it is still possible by the way) - personally, I think that indicates a strong bias present in the pretraining. For what it's worth, I would consider a model that would come up with at least a few different interpretations of the pattern while "reasoning" to be the most intelligent one - irrespective of the correctness of the answer.

mutkach··on The GPT-5 Launch Was Concerning
Not only "AGI" is cancelled but they also sort of admitted that so-called "scaling" "laws" don't work anymore. Scaling inference kinda still works, but obviously is bounded by context size and haystack-and-needle diminishing accuracy. So the promise of even steadily moving towards AGI is dubious at best.
mutkach··on The GPT-5 Launch Was Concerning
The focus now is not the model, but the Product - "here we improve the usuability by removing the choice between models", "here is a better voice for tts", "here is a nice interface for previewing html"

Only about 5 minutes of the whole presentation are dedicated to enterprise usage (COO in an interview sort of indirectly confirms that haven't figured it out yet). And they are cutting the costs already (opaque routing between models for non-API users is a clear sign of that). The term "AGI" is dropped, no more exponential scaling bullshit - just incremental changes over the time and only over select few domains. Actually it is a more welcoming sign and not concerning at all that this technology matures and crystallizes around this point. We will charitably forget and forgive all the insane claims made by Sam Altman in the previous years. He can also forget about cutting ties with Microsoft for that same reason.

mutkach··on You Don't Need Monads
I think it is related to category theory, namely "Yoneda" and Hom functors, but that's a wild guess
mutkach··on You Don't Need Monads
> https://github.com/iokasimov/ya/blob/main/Ya/Operators/Handc...

This is the most arcane codebase I've seen. It's on par with co-dfns compiler. The frontend syntax also looks like a cross between Haskell and APL.

mutkach··on Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs
This is a good take, actually. GPT-OSS is not much of a snowflake (judging by the model's architecture card at least) but TRT-LLM treats every model like that - there is too much hardcode - which makes it very difficult to just use it out-of-the-box for the hottest SotA thing.
mutkach··on Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs
> Inspired by GPUs, we parallelized this effort across multiple engineers. One engineer tried vLLM, another SGLang, and a third worked on TensorRT-LLM. We were able to quickly get TensorRT-LLM working, which was fortunate as it is usually the most performant inference framework for LLMs.

> TensorRT-LLM

It is usually the hardest to setup correctly and is often out of the date regarding the relevant architectures. It also requires compiling the model on the exact same hardware-drivers-libraries stack as your production environment which is a great pain in the rear end to say the least. Multimodal setups also been a disaster - at least for a while - when it was near-impossible to make it work even for mainstream models - like Multimodal Llamas. The big question is whether it's worth it, since when running the GPT-OSS-120B on H100 using vLLM is flawless in comparison - and the throughput stays at 130-140 t/s for a single H100. (It's also somewhat a clickbait of a title - I was expecting to see 500t/s for a single GPU, when in fact it's just a tensor-parallel setup)

It's also funny that they went for a separate release of TRT-LLM just to make sure that gpt-oss will work correctly, TRT-LLM is a mess

mutkach··on Trust in AI coding tools is plummeting
> Respondents were recruited primarily through channels owned by Stack Overflow. The top sources of respondents were onsite messaging, blog posts, email/newsletter subscribers, banner ads, and social media posts. Since respondents were recruited in this way, highly-engaged users on Stack Overflow were more likely to notice the prompts to take the survey over the duration of the collection promotion. We also recruited respondents via a Reddit ad campaign, this accounted for < 2% of total responses.
mutkach··on Trust in AI coding tools is plummeting
Growing disillusionment among programmers (whose productivity gains, by the way, represent the most successful use case for AI yet) is indeed not necessarily a bad thing.

What is concerning is that VCs seem to believe we are still in the exponential growth phase of the hype cycle. I believe the consensus among them (and the bigtech-adjacent shills) is that they are targeting a trillion-dollar market at minimum. Somehow.

mutkach··on Trust in AI coding tools is plummeting
> https://survey.stackoverflow.co/2025/ai
mutkach··on Stargate Norway
>230MW

rather small compared to what was advertised previously

>joint venture between Nscale and Aker

which seems to imply it is not The Stargate (Oracle, Softbank, MGX?)

← PreviousPage 2 of 2