HNHacker News
TopNewBestAskShowJobs

yzh

257 karma · joined October 24, 2013

submissionscomments
yzh··on Show HN: We quantized Qwen3.6-35B-A3B to 2-bit: 12.3 GB, 225 tok/s on a 4090
Thanks for the comment! Fair point. We do not publish that curve for this model, and the HF model card should say so more clearly.

We did run an internal, unpublished FP16-anchored ladder on Qwen3-8B previously using the same pipeline:

          MATH-500, MBPP+, commonsense-6
2.00 bpw: 68.8, 76.5, 73.1

2.44 bpw: 70.8, 76.5, 74.3

2.75 bpw: 72.8, 78.8, 74.4

3.00 bpw: 74.2, 82.0, 75.1

FP16: 75.6, 84.1, 74.1

At 2 bits, the 8B model is about 9% below FP16 on MATH-500 and MBPP+, but only about 1% below on commonsense. The loss is concentrated in generative math and code, and largely returns as bits increase.

That is one reason this release is a 35B MoE at 2 bits rather than an 8B model: there is more redundancy to exploit. On the 35B, the worst regression is LiveCodeBench at -4.4 points, or -6.6% relative. The other five benchmarks are within run-to-run noise. You can find more details on benchmark results and their harness on our HF model card page.

We do not have a BF16 arm for the 35B because it is about 70 GB and does not fit our hardware (we did this with one 4090, one L40, and tested on a bunch of other RTX GPUs :), so FP8 is the reference.

The 35B bit-width ladder is measured in KL divergence against FP8:

2.09 bpw: 0.1041

3.09 bpw: 0.0921

4.13 bpw: 0.0892

The 2-to-3-bit improvement is statistically significant. The 3-to-4-bit improvement is not.

yzh··on MyAI101: Foundational AI literacy for students, teachers and curious adults
Hi HN! We built this to make "why AI works" feel tangible without drowning in math; feedback welcome!

- What it is

We built a structured AI-literacy platform to unpack core AI concepts with: bite-sized lessons + post-lesson quizzes, dozens of in-browser visualizers (neural nets, tokenizers, CNN, GPT-2, etc.), audio and curated videos, and daily AI news feed.

- Why we built it

Most AI courses right now felt either too math-heavy, too surface-leveled, too narrow (covers only a tiny side of AI), or just too expensive. We wanted a single place that balances rigor with intuition - accessible enough for high schoolers (and possibly younger), and structured enough to give adults a solid foundation.

- Who it's for

Anyone curious about AI (what it is, how it works, why it works, and when it doesn't), and in particular high school and college students who should really learn AI fundamentals like they do math and English, teachers looking for classroom-ready materials, adults who find themselves lost in jargons and just want to make sense of it all.

- What to try

Lessons + Thinking Corner & Teacher Notes + audio clip - start with Unit 1's first 5 lessons (after sign up, free); go through the slides, check out the Thinking Corner/Teacher Notes on each slide for additional insights, and pop up the accompanying audio clip to reinforce the slide material.

AI visualizers - GPT-2 Explorer, Tokenizer Playground, Neural Network Visualizer (free). Each visualizer includes sources and a brief 'how it works' section.

Assessments - take the short check-for-understanding at the end of each lesson.

Curated videos (YouTube) - browse through a library of handpicked short videos to reinforce concepts.

AI news feed - finish the day with a 2-minute skim of today's AI headlines; it uses AI to retrieve and rank the news with source-links.

- How we built it

React front-end + Python/Flask services (parts scaffolded with Lovable). Slides via Gamma. Visualizers are a mix of in-house and open-source (credited). Audio clips are generated using NotebookLM with context.

- Limitations

Balancing layman-friendly explanations with technical accuracy was not easy - corrections are welcome.

- Roadmap / feedback

Did we get anything wrong? Where does the difficulty curve feel off? What new content or features should we add? Our current plan is to add more lessons and visualizers. We are also exploring whether to add more practical lessons like building agents, vibe-coding app development, etc. Ideas and critiques are very welcome. Thanks!

yzh··on Groq runs Mixtral 8x7B-32k with 500 T/s
Really impressive work! I wonder how easy would it be to support (a future open source version of) SORA using Groq's design. Will there be a Video Processing Unit (VPU)?
yzh··on Polylabs: AI-powered 3D assets you can use in mobile games
We have created a 3D asset marketplace that contains a set of (expanding) categories of AI generated 3D base meshes that are mobile compatible with free AI-texturing. You can change polycount and also turn mesh into voxels, pretty cool for putting together assets for Roblox and Minecraft type of games. Go check it out and we welcome any feedback!
yzh··on Are Open-Source Large Language Models Catching Up?
AFAIK, model behind yiyan is Baidu's ERNIE. Yi-34B (and Yi model family) comes from another startup created by Kai-fu Lee earlier this year: 01.ai.
yzh··on Optimization Techniques for GPU Programming [pdf]
I would recommend the course from Oxford (https://people.maths.ox.ac.uk/gilesm/cuda/). Also explore the tutorial section of cutlass (https://github.com/NVIDIA/cutlass/blob/main/media/docs/cute/...) if you want to learn more about high performance gemm. OpenAI triton is another good resource if you want to write relatively performant cuda kernels using python for deep learning (https://openai.com/research/triton)
yzh··on BindDiffusion: One Diffusion Model to Bind Them All
We have built BindDiffusion, one diffusion model to bind multi-modal embeddings. It leverages a pre-trained diffusion model to consume conditions from diverse or even mixed modalities. This design allows many novel applications, such as audio-to-image, without any additional training. This repo is still under development. Please stay tuned!
yzh··on AI Research Trends:Sort ArXiv papers by its popularity on social media sites
I agree that some people would prefer to see this as a weekly popup, but some people have the habit of taking a look at what is happening every day :-)
yzh··on AI Research Trends:Sort ArXiv papers by its popularity on social media sites
Since it uses a 10-day window to accumulate the scores.
yzh··on Google adds a guitar tuner to Search
Now I can tell if my "claimed-to-have-perfect-pitch" friend is lying anywhere.
yzh··on Circle – A C++ compiler with compile-time imperative metaprogramming
Sean is a very good engineer. Another of his work is moderngpu, a GPU primitive library. It has the same level of doc/tutorials. I really like one thing he wrote: Software is an asset, code a liability. On good days I'd add 200 or 300 lines to this repository. On great days I'd subtract 500.
yzh··on Show HN: Thinc, a new deep learning library by the makers of spaCy and FastAPI
But Keras does not support PyTorch or MXNet. I think the design of Model, Block, and Layer like this is very intuitive and shared among several frameworks. I wish it could have multi-GPU/multi-node training capability (i.e. support horovod or gloo).
yzh··on First Case of New Coronavirus Detected in U.S.
You're right that the name is 华南海鲜市场 (Huanan seafood market). But the name is actually not accurate. The market sells a wide range of wild animals spanning from rabbits to Chinese bamboo rat, which is suspect to be the source of the virus.
yzh··on First Case of New Coronavirus Detected in U.S.
It is not SARS, but it is highly possible that it was caused by people who kill and sell wild animals in the live poultry market. Where these animals like gem-faced civet got virus from bats. It is a long-held tradition for certain regions in China to eat wild animals. We even have a Chengyu called: 山珍海味, meaning precious and tasty animals from mountains and seas. Most wild animals are not evolved to be eaten by human, so not like beef or pork, they do not taste well. People only eat it because they can. Because they want to show others that they can. This is sad.
yzh··on Volkswagen exec admits full self-driving cars 'may never happen'
Experiences tell us, when someone makes a comment that something may __never__ happen, she probably is wrong. Our current transportation infra is designed for our current transportation tools. Full self-driving cars require a revolution of urban design and transportation infra. Long-term wise I think these will happen for sure.
yzh··on Donald Knuth: Algorithms, Complexity, Life, and the Art of Programming [video]
This also reflects in Von Karman's saying: Scientists discover the world that exists, engineers create the world that never was.
yzh··on Huawei’s Revenue Hits Record $122B in 2019 Despite U.S. Campaign
As a Chinese who spent 7 years in the US and now back in China. I can tell the difference between nationalistic view then (before 2010) and now. Common people used to blindly believe the superiority of the system, but now they believe it with evidence. All things considered, more people are well-educated and they do not blindly believe anything any more. So if they do like something, it must have some real benefit. Actually, Huawei has its own PR crisis just a couple of weeks ago when media exposed that one former Huawei employee was wrongfully jailed for 251 days. I have friends switch from Huawei and to Huawei with all kinds of reasons, but mostly, it is performance and price ratio.
yzh··on AI is mostly about curve fitting (2018)
Related: https://news.ycombinator.com/item?id=21459220
yzh··on Microsoft’s web-based version of Visual Studio
Reminds me of Google's cider-ide: https://code.google.com/archive/p/cider-ide/
yzh··on Few-Shot Video-to-Video Synthesis
Also film industry.
yzh··on Ask HN: What is the most beautiful piece of code you've ever read?
Lots of ioccc code are beautifully done. Especially the best self documenting ones.
yzh··on Ask HN: What is the most beautiful piece of code you've ever read?
Reminds me of a saying I read from moderngpu library's wiki page, a highly optimized yet highly readable GPU basic primitive library: "Software is an asset, code a liability. On good days I'd add 200 or 300 lines to this repository. On great days I'd subtract 500." source: https://github.com/moderngpu/moderngpu/wiki/Introduction
yzh··on What You Can't Say (2004)
Just look at all the comments that got turned to grey. I bet there are patterns, and if looked over a long period of time, would seem like fashions too.
yzh··on CUDA Grep
I remember several years ago I saw this IOCCC winning entry on regex and thought "maybe we should write a parallel version" And it's good to see someone has it implemented now! https://www.ioccc.org/2012/hou/hint.html
yzh··on Group Chats Are Making the Internet Fun Again
Wechat (from Tencent) has figured it out at least 5-6 years ago. Group chat inside wechat is everywhere in China now. People use it at works, for study groups, or SIGs.
yzh··on Google open sourced GPipe, a library for training large ML models
This is actually not about abusively using the computing power (I believe you can use it on a 8-GPU DGX or your DIY workstation), but how to use the same number of GPUs to train a larger model with a larger input size. Pipeline model parallelism seems to be a promising method to increase accuracy by enlarge the model/data size, IIUC.
yzh··on AutoML toolkit for neural architecture search and hyper-parameter tuning
I'm working on auto hyper-parameter tuning and network optimization, I always think that people have put too much focus on NAS, which aims to create a whole new network from scratch, but not nearly enough on hyper-parameter tuning and local structural optimizations for an existing network, which I think is more demanding at least in the industry. Looks less cool than NAS though, maybe that's the reason.
yzh··on UC terminates subscriptions with Elsevier in push for open access
https://openaccessmanifesto.wordpress.com/guerilla-open-acce... Aaron's manifesto is getting more and more relevant. He would be happy to see this!
yzh··on Nasa Happily Reports the Earth Is Greener
Nothing brings me more joy than seeing a greener Earth, especially the greener part is taking place in China!
yzh··on Microsoft confirms Bing is down in China
I started using this meta-search engine: searx.me It works fine for me. Integrates results from different search engines.
Page 1 of 5Next →