HNHacker News
TopNewBestAskShowJobs

throwdbaaway

454 karma · joined February 24, 2020

submissionscomments
throwdbaaway··on CUDA for AMD on Windows
Back to the topic about cuda moat, in the slide with title "Problems with Coopmat1", the current frontier open models have absolutely no issue with:

* manual pipelining

* shared memory staging

* tiling

* bounds checking

throwdbaaway··on CUDA for AMD on Windows
So when I said "a couple of nvidia engineers", I indeed meant Jeff.

VK_KHR_cooperative_matrix - embrace?

VK_NV_cooperative_matrix2 - extend?

I am pretty sure VkImportSemaphoreFdInfoKHR, mentioned in https://github.com/ggml-org/llama.cpp/issues/22648, works across multiple AMD devices, but somehow doesn't work across multiple nvidia devices.

throwdbaaway··on CUDA for AMD on Windows
On the other hand, despite my complain about the standard API, the models were able to come up with cooptmat1 kernels that run dsv4 flash faster than whatever the guys at antirez/ds4 can come up with using rocm, on a strix halo, with the added benefit that I can also pair the strix halo with an egpu to drastically speed things up.
throwdbaaway··on CUDA for AMD on Windows
From what I can tell, coopmat2 can get to about 75~90% of cuda performance on a single device, and there is no good way to do direct communication across devices. It is fair to say that nobody would replace cuda with coopmat2? That looks like a EEE project that can assigned to a couple of nvidia engineers, to fragment the ecosystem.
throwdbaaway··on CUDA for AMD on Windows
I tried getting LLMs to add proper Vulkan support to ik_llama.cpp, which have very good support for CUDA and CPU. The models do an admirable job; they don't care much about poor DX.

Few problems I noticed:

* coopmat2 from nvidia is the classic embrace, extend, extinguish. No point to ask the models to translate from CUDA to coopmat2. Instead, the models can understand the existing CUDA and CPU kernels, and adapt them accordingly to non-nvidia devices.

* However, the standard API is also lacking. The models struggled to make prompt processing compute-bound on strix halo when the graph is complex. Upfront standard API might just be an evolution dead end.

throwdbaaway··on GLM-5.3 is now open-weight
Hold on.. the routed experts are in FP8 now? Previously they were in BF16. Nice, this shall cut my download time by half!
throwdbaaway··on Nvidia agrees to acquire Hugging Face for $13B
Sounds like that's what z.ai did to get GLM-5.3-Flash running on Huawei chips.
throwdbaaway··on GLM-5.3-Flash
Exactly. Coding for inference is solved. CUDA is no longer a moat.
throwdbaaway··on Apple introduces M6 and M5 Ultra
> $20k workstation, best case: $15k M5 Ultra 512GB, 36-month amortization, ~$440/mo. Runs a GLM-5.3-class model at ~30 tok/s. Saturated 24/7 it produces roughly 58M output tokens/month.

For agentic coding, ~90% of the cost comes from cached input tokens. This cost increases quadratically with the session length. If sessions go near 1M context, the number of cached input tokens can easily exceed 1B in a day.

GLM-5.3 @ $0.26/M x 1000 = $260/day

This is the math to use.

throwdbaaway··on Why your local LLM feels dumber than it is
As for KV cache quantization, Q8_0 from llama.cpp / ik_llama.cpp should also work better than FP8 from vllm (see https://github.com/vllm-project/vllm/issues/33480#issuecomme...).
throwdbaaway··on Why your local LLM feels dumber than it is
> Both the NVFP4 and AWQ W4A16 failed to properly close their tool calls ...

If I understand correctly, this failure mode is just not possible with llama.cpp / ik_llama.cpp, which enforces token generation to follow the grammar once a tool call is detected.

> ... and botched Cisco command line syntax (the correct command was ‘show arp’, while they executed ‘show run’)

But this failure mode can still happen.

Anyway, NVFP4 and AWQ W4A16 are generally regarded as low quality quants. IQK/Trellis quants from ik_llama.cpp and EXL3 quants from exllama should work better.

So, perhaps the lesson here is "don't use vllm at home"?

throwdbaaway··on DeepSeek API Pricing Update
Reproduced on the CUDA stack right?

Let's say DeepSeek is being forced to use the CANN stack, and the new pricing reflects the cost when 100% of inference is done with Huawei chips. Then, I suppose we can infer that:

* CANN stack is 1.5x~2.3x less efficient in compute

* CANN stack has 6x lower inter-connect capacity

> computer chips once again become a commodity

Ascend 950 is going for $7k to $9k with mediocre looking specs. $16k for RTX Pro 6000, $6k for RTX Pro 5000. This is not looking good.

throwdbaaway··on DeepSeek-V4-Flash Update
Objectively speaking, the 2 bit quant from antirez has very low accuracy. Meanwhile, his 4 bit quant does have decent accuracy, but is a bit pointless by being bigger than the full precision MXFP4 quant. Anyway, they all work fine in practice.
throwdbaaway··on Kimi-K3 on HuggingFace
They need to get a license from moonshot to provide inference for K3. Probably have to follow the pricing set by moonshot as well.
throwdbaaway··on Laguna S 2.1
It works, thanks to https://github.com/ikawrakow/ik_llama.cpp/pull/1911, which got merged in early June. However, there might still be some issue with the chat template.
throwdbaaway··on Qwen 3.8
I suspect this is why DeepSeek had to introduce the 2x peak hours pricing. The price would be too low otherwise.
throwdbaaway··on Control the Ideas, Not the Code
Yeah antirez made a lot of big claims in that paragraph. Sounds like a case of AI psychosis.
throwdbaaway··on Show HN: Getting GLM 5.2 running on my slow computer
If you max out the ram, TG with q3 should be at least 10 t/s. And with dsa, it can still stay close to that number as the context grows.
throwdbaaway··on GLM 5.2 and the coming AI margin collapse
That's exactly what I said. They do care when FLOPs are involved. Restoring an old session with 900k tokens will require a lot of FLOPs to reprocess the 900k token.

Meanwhile, they don't really care if you use hundreds of millions of cached input tokens, which doesn't consume any FLOP.

throwdbaaway··on GLM 5.2 and the coming AI margin collapse
Different sessions. With https://github.com/fairydreaming/llama.cpp/tree/dsv4, 1M context with DSV4 Flash takes less than 6GB of VRAM. I can't run DSV4 Pro, but it should take less than 9GB of VRAM for 1M context, based on the numbers shared in https://arxiv.org/html/2606.19348v1.
throwdbaaway··on GLM 5.2 and the coming AI margin collapse
Well I wouldn't call it a low bar, since some of the edits were quite complex. And 1M context in less than 6GB of VRAM is truly impressive, but somehow this gets way less attention than the crappy turbo quant from Google.
throwdbaaway··on GLM 5.2 and the coming AI margin collapse
While we are all speculating, Boris kindly provided some guidance in https://news.ycombinator.com/item?id=47880089

> The challenge is: when you let a session idle for >1 hour, when you come back to it and send a prompt, it will be a full cache miss, all N messages. We noticed that this corner case led to outsized token costs for users. In an extreme case, if you had 900k tokens in your context window, then idled for an hour, then sent a message, that would be >900k tokens written to cache all at once, which would eat up a significant % of your rate limits, especially for Pro users.

Using the current Opus pricing, that pre-lunch 900k tokens should roughly consist of:

720k input tokens = 0.72 x $5 = $3.6

180k output tokens = 0.18 x $25 = $4.5

900k 1h cached writes = 0.9 x $10 = $9

500M cached input tokens = 500 x $0.5 = $250

$267.1 in total, with 93.6% from cached input tokens. The portion that requires GPU compute is about 3% of the total.

Post-lunch, the 900k tokens should consist of:

900k input tokens = 0.9 x $5 = $4.5

900k 1h cached writes = 0.9 x $10 = $9

So Anthropic is fine with the $267.1 accumulated over 3~4 hours before lunch, but not fine with the $13.5 incurred immediately after lunch. Why?

The only plausible explanation is that the actual cost of caching is way less than the API pricing. If you use a coding plan, Anthropic doesn't really care about your cached input tokens usage. Indeed they want you to show your ccusage screenshots. On the other hand, if you pay by API tokens, the margin is huge for cached input tokens.

Only when you do something that requires a lot of FLOPs, e.g. the post-lunch 900k input tokens, the cost becomes real.

throwdbaaway··on GLM 5.2 and the coming AI margin collapse
Indeed they are all lossy. Not sure how much they contribute to the quality loss in long context though. I got a 700k session with DSV4 Pro (official API), and the model was still coherent and didn't make any tool call error.
throwdbaaway··on GLM 5.2 and the coming AI margin collapse
The current top comment in https://lobste.rs/s/ua1gxl/glm_5_2_coming_ai_margin_collapse correctly zoomed into cached input tokens, but landed on the opposite conclusion:

> That is, for your $100/month fee, you get $3600 equivalent of API usage. This is presumably because Anthropic has figured out some clever things to do with model routing and input caching, and also can subsidize with investor money and take a hit on their operating margins.

My take: this is exactly what Anthropic wants everyone to think. In reality, 90% of that $3600 are for cached input tokens, that can be made to cost next to nothing, as shown by DeepSeek.

throwdbaaway··on GLM 5.2 and the coming AI margin collapse
Seems like a pretty pointless post that still centers around output tokens.

In agentic coding, cached input tokens is 90% of the API "cost". It doesn't require GPU compute, and DeepSeek has shown that it can be done 50~100x cheaper with MLA/CSA/HCA, and a whole bunch of disks. This should collapse the margin.

throwdbaaway··on Performance per dollar is getting faster and cheaper
And somehow they claimed that it is "lossless".
throwdbaaway··on GLM-5.2 – How to Run Locally
On ZFS with zstd compression, I am getting 1.34x compressratio for the BF16 weights (across multiple models).

Here's the du output for GLM-5.2:

    $ du -s -BG /cube/models/zai-org/GLM-5.2/
    1099G   /cube/models/zai-org/GLM-5.2/
throwdbaaway··on DeepSeek makes the V4 Pro price discount permanent
And their disk-based caching is amazing. I got a long 700k context session spanning more than a week, with pauses in between that was longer than a day, and some rewinds mixed in as well.

Stats from pi:

↑400k ↓438k R432M 71.9%/1.0M

Half a billion tokens, $2.12

throwdbaaway··on A few words on DS4
Hah, that's because the prompt itself was only about 30 tokens. We need a much bigger prompt to properly test PP.
throwdbaaway··on An AI agent deleted our production database. The agent's confession is below
Huh that's not what I gathered from the tweet at all. If I am going to write a five why's analysis, the immediate cause is the LLM wrongly decided to delete a volume, while the root cause is the bad design to co-locate staging and production data in the same volume. The writing was quite vague though, let's wait for a response from railway.
Page 1 of 11Next →