82 karma · joined July 31, 2025
I didn't know Mike Judge was such a polymath!
I am wondering how would you use a chat transcript for training? Unless it is massive, possibly private codebases that are constantly getting piped into Claude Code right now. In that case, that would make sense.
It could still be burning money for Microsoft/Amazon
On the other side, there's an insane booster of speculative decoding, that would give a semi-prefill rate for decoding, but the memory pressure is still a factor.
I would be happy to be corrected regarding both factors.
None of the models did actually "reason" about what the problem could possibly be - like none of them considered that more intricate patterns are possible in a 3x3 grid (having taken this kinds of test earlier in life, I still had a few seconds of indecision, thinking whether this is the same kind of test that I've seen and not some more elaborate one), and none of them tried solving the problem column-wise (it is still possible by the way) - personally, I think that indicates a strong bias present in the pretraining. For what it's worth, I would consider a model that would come up with at least a few different interpretations of the pattern while "reasoning" to be the most intelligent one - irrespective of the correctness of the answer.
Only about 5 minutes of the whole presentation are dedicated to enterprise usage (COO in an interview sort of indirectly confirms that haven't figured it out yet). And they are cutting the costs already (opaque routing between models for non-API users is a clear sign of that). The term "AGI" is dropped, no more exponential scaling bullshit - just incremental changes over the time and only over select few domains. Actually it is a more welcoming sign and not concerning at all that this technology matures and crystallizes around this point. We will charitably forget and forgive all the insane claims made by Sam Altman in the previous years. He can also forget about cutting ties with Microsoft for that same reason.
This is the most arcane codebase I've seen. It's on par with co-dfns compiler. The frontend syntax also looks like a cross between Haskell and APL.
> TensorRT-LLM
It is usually the hardest to setup correctly and is often out of the date regarding the relevant architectures. It also requires compiling the model on the exact same hardware-drivers-libraries stack as your production environment which is a great pain in the rear end to say the least. Multimodal setups also been a disaster - at least for a while - when it was near-impossible to make it work even for mainstream models - like Multimodal Llamas. The big question is whether it's worth it, since when running the GPT-OSS-120B on H100 using vLLM is flawless in comparison - and the throughput stays at 130-140 t/s for a single H100. (It's also somewhat a clickbait of a title - I was expecting to see 500t/s for a single GPU, when in fact it's just a tensor-parallel setup)
It's also funny that they went for a separate release of TRT-LLM just to make sure that gpt-oss will work correctly, TRT-LLM is a mess
What is concerning is that VCs seem to believe we are still in the exponential growth phase of the hype cycle. I believe the consensus among them (and the bigtech-adjacent shills) is that they are targeting a trillion-dollar market at minimum. Somehow.
rather small compared to what was advertised previously
>joint venture between Nscale and Aker
which seems to imply it is not The Stargate (Oracle, Softbank, MGX?)