HNHacker News
TopNewBestAskShowJobs

woctordho

185 karma · joined February 7, 2023

submissionscomments
woctordho··on Vibecoding Photoshop: Time and pressure
Just use it in your workflow. If you encounter anything that doesn't work, then vibe code it to make it work. I'm exactly using it in my indie game dev with heavy graphics workflow.
woctordho··on CUDA for AMD on Windows
No, 10.1 is current
woctordho··on Detecting and countering misuse of AI: September 2026
Let me put my two cents: In China we've got accustomed to the fact that every word we say will be seen by the surveillance, so it's not a big problem that Anthropic also see it. Also we know that they can see it but they can't stop it. There are all kinds of ways to work around account blocking.

As the old saying goes, communists disdain to conceal their views and aims.

woctordho··on Qwen 3.8 follows GPT-5.5 Pro reasoning prefills
Reasoning works as long as there is a consistent latent space representation. Any kind of poison will just become part of the representation. There's evidence that even directly training on encrypted reasoning traces works, because the length is already a strong signal.
woctordho··on Path to Astra: critical capabilities and frontier safeguards
There are 'transfer stations' and that's how exactly I use GPT and Claude in China. OpenAI and Anthropic do not sell in China, so we use their AI with a much lower price like 1% of the official API price. The largest transfer stations have TBs of traffic every day, and the traffic is eventually possessed by the open source community.

Subscription engineering is a deep field. Neither OpenAI nor Anthropic have any technical advantage in this field.

woctordho··on Path to Astra: critical capabilities and frontier safeguards
All the RL data are exactly public. There are huge amount of distilled data freely available, and that amount is more than enough to train a ~10T model.
woctordho··on Seedance 2.5
MiniMax H3 is going to release weights. You can locally run it with definitely less than $10k (and possibly faster than Seedance's queue), and it's fun to train it for whatever you need.
woctordho··on Kimi-K3 on HuggingFace
GGUF is at least better than bnb. From what I know, bnb does not yet find a way to quantize MoE with enough accuracy, and maintain the dequant-MoE kernels. In the age of Qwen 3.0, people tried to make some bnb '4-bit' quants of MoE models, but actually the MoE part is not quantized. It's a pity that even Unsloth gave up low-VRAM finetuning with MoE (although they're making their GGUFs for inference), and the world of local training looks stagnated for months.

GGUF is maintained by all the llama.cpp developers. There are many quantization formats and algorithms under this container format, some are optimized for MoE (such as APEX quant), some for CPU and some for GPU, some work surprisingly well below 4-bit (and even near 1-bit). It also supports recent architectures like linear attentions and mHC.

woctordho··on Kimi-K3 on HuggingFace
Speaking of finetune, currently a common practice is LoRA over bnb 4-bit base model, but I think it's time to replace bnb with GGUF as the base model format. GGUF is actively supporting new model architectures and more aggressive quantizations.

I've made some proof of concept in https://github.com/woct0rdho/transformers5-qwen3.5-recipe . We can finetune Qwen3.5-35B-A3B in 16 GiB VRAM, and DeepSeek-V4-Flash (284B-A13B) in 90 GiB VRAM, without CPU offload. This works well on unified memory machines like Strix Halo.

Even so, larger models like Kimi-K3 still require multiple GPUs and nodes, and there are a lot more to do compare to single-GPU training.

woctordho··on Petals: Run LLMs at home, BitTorrent-style
Distributed training is much harder than distributed inference but not impossible. See the recent development of DiLoCo at Nous Research and Prime Intellect.
woctordho··on Petals: Run LLMs at home, BitTorrent-style
Relevant: Why Switzerland has 25 Gbit internet and America doesn't https://news.ycombinator.com/item?id=47652400
woctordho··on Petals: Run LLMs at home, BitTorrent-style
AI Horde has some measures to prevent Sybil attack that returns wrong results, but not enforce zero data retention. Prompts belong to the whole open source community. For example https://huggingface.co/datasets/la-ji/sd-prompt-in-the-wild
woctordho··on Petals: Run LLMs at home, BitTorrent-style
Petals is from 2022. Nowadays intelligence of smaller models, quantization techs, and optimizations to run models faster on consumer GPUs have improved a lot.

For distributed inference of smaller LLMs and diffusion models that fits in one consumer GPU rather than splits on multiple machines, there are already pretty good solutions such as AI Horde (formerly Stable Horde) [0]. Notably, it's the default provider that powers SillyTavern. It also has an interesting economy model of kudos.

[0] https://stablehorde.net/

woctordho··on China’s open-weights AI strategy is winning
So is making a PR different from making the whole software. This is what an open source community is good for.
woctordho··on China’s open-weights AI strategy is winning
See the recent development of DiLoCo at Nous Research and Prime Intellect.
woctordho··on China’s open-weights AI strategy is winning
There's a lot of individual effort of improving the models. See how many finetuned models and LoRAs are there on Hugging Face.
woctordho··on Qwen 3.8
There is a forum named Zhihu. AI translation works mostly well to translate contents there into English.
woctordho··on Kimi K3: Open Frontier Intelligence
Yes in a mid-sized company. I'm exactly doing this, and what I'm competing against is the OpenAI API priced 0.2 CNY = 1 USD in China.
woctordho··on Alternative(s) to run CUDA on non-Nvidia hardware
There's nothing wrong to run CUDA on non-Nvidia hardware. CUDA has an interface that is reasonably well-designed, well-documented/reverse-engineered, and battle-tested for decades. What we need is not to invent another interface just under the name of 'open standard', but to implement the same interface. ROCm is exactly doing this, and so are other hardware SDKs such as MooreThread and Alibaba T-Head.
woctordho··on DSpark: Speculative decoding accelerates LLM inference [pdf]
And humans don't run on markets.
woctordho··on The gap between open weights LLMs and closed source LLMs
Fun fact: Hacker News is canonically banned in China, but I'm still talking here. There are plenty of techs to work around region block. The incentive to report somebody is comically called '50w' (500k CNY) and no one gives a shit about it in real life.
woctordho··on The gap between open weights LLMs and closed source LLMs
See the recent advance of DiLoCo at Nous Research and Prime Intellect.
woctordho··on U.S. allows Anthropic to release Mythos AI to ‘trusted’ US organizations
Don't trust US or China. Trust the open source community.
woctordho··on Anthropic says Alibaba illicitly extracted Claude AI model capabilities
There are lots of botnets providing home IPs.
woctordho··on Anthropic says Alibaba illicitly extracted Claude AI model capabilities
Lots of people have succeeded. Neither Anthropic nor OpenAI has any technical advantage in the field of subscription engineering.
woctordho··on Anthropic says Alibaba illicitly extracted Claude AI model capabilities
Actually nowadays LLMs are only trained with TBs rather than PBs of data, and it's not too hard to find GBs of agent traces online.
woctordho··on GLM-5.2 is the new leading open weights model on Artificial Analysis
Simple trick: Use an agentic tool like Pi or OpenCode that allows you to switch models. First do some chats with DeepSeek or GLM who shows full thinking traces, then switch to Claude or GPT and it's more likely to show full thinking traces.
woctordho··on GLM-5.2 is the new leading open weights model on Artificial Analysis
There is already a lot of effort to collect agent traces including reasonings, e.g. see the recent discussion: https://old.reddit.com/r/LocalLLaMA/comments/1u795pb/donate_...

We've been developing DataClaw for this: https://github.com/peteromallet/dataclaw

woctordho··on Lore – Open source version control system designed for scalability
There's `--filter=blob:none` and it allows to automatically fetch blobs when needed.
woctordho··on Lore – Open source version control system designed for scalability
It's 2026. Historically the way for large binaries in git was git LFS. Now the way for large binaries in git is just git.
Page 1 of 5Next →