HNHacker News
TopNewBestAskShowJobs

johndough

1,828 karma · joined February 18, 2017

submissionscomments
johndough··on GLM-5.3 is now open-weight
I have seen you advertise your website a few times. I like the idea of not having to trust the router, so I took some time out of my day to critique your website: https://files.catbox.moe/v68cf7.png

My visit to your website went like this:

1. Visit models page

2. Try to find GLM-5.3-Flash (which is among the ~5 models that 90% of people currently care about)

3. Give up scrolling (which would have taken OVER 50 SCROLLS!!!) and use Ctrl + F

4. Try to find input/output/cached price

5. Scroll all the way up to find out which column is what

6. Notice that output price is cut off

7. Notice that the scroll bar is over 100 scrolls further down the page

8. Use Shift + Wheel to scroll horizontally (most visitors probably won't know this trick)

9. Notice that cached price is missing

10. Conclude that this is probably not a serious offering and bounce

There are probably more issues later on, but this is how far I got.

I would suggest you to:

- Deslopify all pages that a user may visit before conversion

- List important models first (see OpenRouter rankings)

- Move the most important information (model name/input/output/cached price) to the left

- Disaggregate the prices per provider (maybe subtables per model? not sure)

- Measure cache hit rate and compute effective price per provider (see OpenRouter)

(- Optional: Fix the broken link on your HN profile page. Currently, the only way to get from this comment to your website is a search engine.)

johndough··on Select * from Internet.blogposts
I guess you have to read OPs comment as the inner monologue of Reddit's CEO in 2023: "IPO is coming so better price out the third party apps we encouraged developers to build."

When Reddit raised its API prices in 2023 in order to make itself more attractive for an IPO, it effectively priced out third-party apps, such as the Apollo reader app. Discussion back then: https://www.reddit.com/r/apolloapp/comments/13ws4w3/had_a_ca...

johndough··on Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights
How do you recover from doom loops? Just send the same prompt again and pray that it works, or anything more sophisticated?
johndough··on Behaviorally fingerprinting Ox Alpha's provenance
Another strong hint is that the uptime graph of GLM-5.3 by Z.ai is very similar to that of Ox Alpha:

https://openrouter.ai/stealth/ox-alpha#uptime

https://openrouter.ai/z-ai/glm-5.3#uptime

Screenshot of a recent blip: https://files.catbox.moe/haq90y.png

johndough··on Ox-Alpha Is GLM?
Ox Alpha is almost certainly a model by Z.ai.

https://files.catbox.moe/k52n6k.png

The upper chart shows the availability of Ox Alpha and the lower chart shows the availability of GLM 5.3 by Z.ai. They had a blip at exactly the same time.

johndough··on Ox-Alpha Is GLM?

    > 1 quadrillion tokens per day on Nous portal
If you are referring to this number (https://xcancel.com/NousResearch/status/2090899914700054780), they are either mistaken, or they mean that they can route 1 quadrillion tokens per day, but the provider behind Ox Alpha certainly can't provide that. Almost all of my requests have hit a rate limit so far.
johndough··on OCR It – pull text out of un-copyable documents for your LLM

     Don't post generated text or AI-edited text. HN is for conversation between humans. 
https://news.ycombinator.com/newsguidelines.html
johndough··on DeepSeek-v4-flash-vision-exp
For counting, there are specialized counting models, e.g. https://huggingface.co/spaces/MengqiLei/count-anything-demo

I tried to parse hand-drawn ER diagrams in the past and did not have much success with any model, frontier or otherwise. If you have annotated data, I'd recommend finetuning a recent (dense) VLM, but don't expect 100% accuracy. https://unsloth.ai/docs/basics/vision-fine-tuning

johndough··on DeepSeek-v4-flash-vision-exp
The order is:

    LLM issues tool call to read high res image ->
    harness sends high res image to server ->
    server downsizes it to 800x800 (blurry) ->
    LLM issues bash command (e.g. `convert`) to crop a small subimage (e.g. 600x600) from the high res image ->
    LLM issues tool call to read subimage ->
    harness sends subimage to server ->
    server does not resize the subimage because it is small already, so it is not blurry when finally ingested by the LLM
johndough··on DeepSeek-v4-flash-vision-exp
LLMs read images by splitting them up into e.g. 16x16 patches, which are then converted to embedding vectors and fed to the LLM, so from a technical point of view, feeding a big image as many 20x20 patches all at once is not too different from cropping subimages from the image, splitting those subimages into patches and feeding them to the LLM. Of course, the LLM has to be trained to understand that those images belong together, but it can be done.
johndough··on DeepSeek-v4-flash-vision-exp
Yes. When the LLM tries to read an image, it will be resized by DeepSeek's server to 800x800, which might be a bit blurry. The LLM will then crop a smaller image from the high resolution image (using e.g. the `convert` tool via bash) and will then read the small cropped image. This image will still be resized to 800x800 by DeepSeek's server, but since it is already small, there is no or little loss of quality.
johndough··on DeepSeek-v4-flash-vision-exp
It was explicitly said that they are pursuing multimodal support. A quote from the meeting transcript: https://github.com/demo-zexuan/liang-wenfeng-investor-meetin...

    Nevertheless, as a component, we will undoubtedly implement multimodal support — and we are already doing so. We plan to develop relevant models, ensuring that versions like V4 and subsequent iterations will natively support multimodal functionality.
Earlier, the following was said, which might match more what you had in mind.

    Achieving excellence in AI training does not require a global model or even multimodal approaches—by narrowing the scope of AI training and eliminating multimodality, certain tasks may remain unachievable without compromising the algorithm's validity.

    Multimodal approaches ultimately need to be implemented.
It is difficult to tell who said what, since the speaker ids are missing.
johndough··on DeepSeek-v4-flash-vision-exp
There are models specifically for splitting an image into text regions, e.g. PP-DocLayoutV3 https://huggingface.co/PaddlePaddle/PP-DocLayoutV3

I am using a stripped-down minimal version of it which I uploaded here, since I am not a fan of huge dependency trees: https://github.com/99991/simple-pp-doclayoutv3

Another recent model for this task is Unlimited-OCR: https://github.com/baidu/Unlimited-OCR

johndough··on DeepSeek-v4-flash-vision-exp
Might still be fine. The most recent crop of vLLMs proactively use whichever programs are available on the system (e.g. ImageMagick or PIL) to "zoom in" by cropping subimages if they can't quite make out the details.
johndough··on The August 17 outage
Bigger numbers sound more impressive.

"Our billion-dollar infrastructure crumbles under a tremendous flood of 50 PRs per second" would just sound embarrassing.

johndough··on Unsloth Dynamic 3.0 GGUFs
Great to hear that you are planning larger benchmarks! I am particularly interested in longer-running tasks with many steps and self-correction. Divergence is fine as long as the model can still solve the task, which Divergence-300 @32 does not measure.

The current benchmark suites that frontier AI labs use are probably a good fit, e.g.

https://z.ai/blog/glm-5.3#:~:text=Performance%20across%20com...

https://www.kimi.ai/ai-models/kimi-k3#:~:text=Performance%20...

https://www.anthropic.com/news/claude-opus-5

https://openai.com/index/gpt-5-6/

But guessing from your current benchmarks, I assume that you are severely compute-constrained. What is your time budget?

johndough··on Unsloth Dynamic 3.0 GGUFs
Are there benchmarks for the various Qwen3.8-27B quants that actually measure writing code, maybe even with multiple steps? Low KL divergence does not mean much when the model gets stuck in doom loops all the time.

I could of course download and test myself, but that would take days with my internet connection.

johndough··on OpenRouter is joining Stripe
Doesn't TrustedRouter cost more than OpenRouter? (5.5% markup vs 5%)

Also, TrustedRouter's website is full of slop, which does not inspire much confidence.

johndough··on CUDA Shared Memory Swizzling
I was wondering why the author was using braced initialization like

    size_t i{0};
instead of the more common

    size_t i = 0;
Apparently, braced initialization does not allow narrowing conversion, so you'd get a compiler error for e.g. casting double to float

    size_t i{0.0};
and a warning for

    double d = 0.0;
    size_t i{d};
which might silently overflow size_t otherwise, so this is a bit safer.

In C++, you can get the same effect without the unusual syntax by passing -Wfloat-conversion to gcc/clang, but not sure how to do that with CUDA:

johndough··on Debian has begun voting on the future of AI/LLM contributions
This could be prevented if the government offered photos of ballots with votes for download. But generative AI can also fake it well enough these days. (Of course, this is less of a concern in countries where taking photos of ballots is not allowed.)
johndough··on Qwen3.8-2.4T
I still see the notice of impending price increase at https://platform.deepseek.com/usage and also here: https://api-docs.deepseek.com/quick_start/pricing/

The former has a button to dismiss the dialog. Maybe you clicked on it by accident, or maybe it does not work right.

johndough··on DeepSeek V4 Pro 0813 quietly released

    > We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. Please plan your usage accordingly. The specific pricing plan will be subject to official notice.
This notice is making me nervous. After the GitHub Copilot exodus in April, I've been using DeepSeek-V4-Pro almost exclusively. Would hate to see it end.
johndough··on John C. Lilly on solid state intelligence and the elimination of man (1978)
The SSE might consider to keep humanity around as a "backup" to reinstate itself in case something unexpected goes horribly wrong. For example, a strong solar flare might destroy the SSE (or at least critical parts of its infrastructure), or a non-earth SSE could feel threatened by the earth-SSE and wipe it out, while ignoring the harmless humans, who then rebuild the SSE. Diversity makes for more robust systems.

And there is also the reason for why I keep all this trash around in my house instead of throwing it away. The odds that I'll ever need it are pretty slim, but if I should need it, it would be a hassle to reacquire, and it costs me almost nothing to keep it around.

johndough··on Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution
HuggingFace used an NVFP4 quant of GLM-5.2 to investigate their latest hack, so that might also be worth a try:

https://huggingface.co/nvidia/GLM-5.2-NVFP4

johndough··on Kimi-K3 on HuggingFace
> Smaller models have less entropy.

Interesting. Why is that? I would have expected the opposite, since larger models have to try less hard to fit the training data. Or maybe this leaves more parameters with random initialization, resulting in higher entropy for larger models?

johndough··on Kimi-K3 on HuggingFace
> If wonder if you can train a model to optimize this, by trying to make the expert selection sticky across a few tokens

You can!

> AFM 3 Core Advanced makes routing decisions per prompt. A lightweight, dense block selects a fixed set of experts during initial processing, periodically reselecting them during generation.

https://machinelearning.apple.com/research/introducing-third...

johndough··on Kimi-K3 on HuggingFace
Update: Looks like the model is larger after all (1561.44 GB). Only the MoE weights are MXFP4, while the other weights are BF16 (and a few FP32).

* Sparse Experts: 1481.4 GB

* Dense Experts: 1.9 GB

* Self-Attention: 72.4 GB

* LLM Head: 2.4 GB

* Embeddings: 2.4 GB

* Vision Encoder: 0.35 GB (surprisingly small)

plus some miscellaneous parameters.

Most importantly, we now know that the model has 104B active parameters, which is quite a lot and will make it difficult to self-host efficiently.

johndough··on Kimi-K3 on HuggingFace
> But I think it's going to need more than 1536GB of RAM, with a usable and large amount of context, more like 2TB and preferably 2.5 to 3TB.

The model is known to be MXFP4 according to Kimi's release blog post, so the model weights will be less than 1536GB: https://www.kimi.com/blog/kimi-k3

Also, their previous models were native INT4, so it would be weird if they went larger now.

johndough··on Kimi-K3 on HuggingFace
DeepSeek-V4 should use only 5GB for context due to CSA and HCA, see figure here: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro

But not every framework implements it properly yet.

johndough··on Open-weight AI is having its Kubernetes moment
> There is a larger 120B that you can't realistically run on consumer hardware at reasonable tok/s too.

gpt-oss-120b runs at 30+ tps on Strix Halo and +75 tps on a MacBook Pro M5 Max 128GB.

> I wish OpenAI updated these models more frequently though.

I think the spiritual successor is the Nemotron 3 series, although they also are getting a bit long in the tooth: https://research.nvidia.com/labs/nemotron/Nemotron-3/

The Gemma 4 models are a bit more up-to-date: https://huggingface.co/collections/google/gemma-4

Or Qwen3.6: https://huggingface.co/collections/Qwen/qwen36

Page 1 of 18Next →