HNHacker News
TopNewBestAskShowJobs

liuliu

3,289 karma · joined September 10, 2008

submissionscomments
liuliu··on M5 Ultra Mac Studio Review
Both are probably single-token decode performance, which is reasonable to show. Otherwise agree RTX 5090 should shinebetter with NVFP4.
liuliu··on M5 Ultra Mac Studio Review
When people benchmark MLX related quant models, they really need to publish numbers on benchmarks. You cannot take this as it is what you get of the original models. MLX uses pretty simple quantization methods so at lower bits without QAT, it is just not as good quality as llama.cpp ones.
liuliu··on Xiaomi Mimo 2.6 live post-training dashboard
Correct. If just stopping criteria, that is less contaminated. The question gets muddier once you also use it to determine hyperparameters during small-scale runs.
liuliu··on Xiaomi Mimo 2.6 live post-training dashboard
When you run benchmarks while training, isn't that the definition of contamination? Asking because I am not sure if this is normal in big labs now.
liuliu··on I stress-tested Meta Muse until its agent control plane started timing out
I sort of understand DSH arch, and their choices, but still baffled by why you would want the agent to change its own agent loop other than the system prompt (in that extension, memory / soul / tools / skills).
liuliu··on Run Qwen3.8 27B locally: real numbers from my Mac Studio
Yes, I heard you! I think one of the issue Draw Things inherited is the baggage of supporting too much models. Once you settled on a model, then it is just "try recommended settings", and prompt.

The model part is unfortunate, but luckily converging now.

On the LLM side, hopefully it is not an issue too as long as getting aggressive at pruning models.

liuliu··on Run Qwen3.8 27B locally: real numbers from my Mac Studio
What you get with Draw Things: 1. Download the app from Mac AppStore; 2. Download the model; 3. Tap "Try recommended settings", have the guarantee that for whatever model it supports, it is the fastest in Mac ecosystem, no need to fiddle.

What you get with local LLM options: 1. Download the inference engine app from the web; 2. Download the model; 3. Configure MTP / DFlash / DSpark whatever; 4. Configure your Pi / OpenCode harness to point to this local LLM inference engine. 5. Configure tools for these harness to be effective. 6. Switching between Ollama, LM Studio, llama.cpp (DwarfStar4), oMLX, MTPLX, to see which one is fastest for your workload. 7. Again switching between different quants of the same model to see which one is less dumb.

To be honest, llama.cpp probably the closest to deliver on "just use it, don't worry about speed" if your focus is about a pure LLM inference engine.

liuliu··on Run Qwen3.8 27B locally: real numbers from my Mac Studio
The local LLM scene needs a Draw Things equivalent for Mac. Too much fiddle for things that doesn't make sense (Qwen 3.8 27B should be exactly the same speed as Qwen 3.6 27B). It feels like that I am teasing (I am the author of Draw Things) something, because it is.
liuliu··on DFlash 2: Keep Drafting Parallel
I agree. But I think the DFlash2 case is just that 1/1000 invalid syntax failure case from sampling rather than a bug.
liuliu··on DFlash 2: Keep Drafting Parallel
Correct. I am trying to explain why even it is "exact", the generated text is different from the with / without DFlash2 runs, and potentially why the DFlash2 run will contain the invalid Python syntax.
liuliu··on DFlash 2: Keep Drafting Parallel
Yes, it doesn’t impact the probability distribution due to verifier. However, remember how you use PRNG and effectively due to the drafter is sampled from a different distribution initially, a separate rejection sampling won’t be able to recover what the “old PRNG” would choose in a “without drafter” case. Hence in my original post, it is about different trajectories you will end up with, not the correctness of each stochastic sampling.
liuliu··on DFlash 2: Keep Drafting Parallel
Only if you do greedy sampling. With probabilisitic sampling (categorical sampling), you will end up with different trajectory just “mathematically equivalent”.
liuliu··on Go is an ideal language for AI-assisted software engineering
1. The syntax surface is smaller, allowing less LLM "creativity; 2. The error handling is mechanical, which LLM clearly prefers (LLM is already trigger happy about writing tons of throw / try...catch.. in other languages, doing tons of `if err` is just in it comfort-zone).
liuliu··on MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
They ship a complete checkpoint for easily management (inference & training) in their own infrastructure. Moving to a LUT would make training on these layers impossible. BTW, these are not useful for lightweight fine-tuning, but might still be useful if you do serious post-training work.

Of course, these are also not an issue for things like FLUX.2 which adopts DiT-Air arch, that doesn't have this wasted space issue.

liuliu··on MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
One thing similar would be projecting both the head.weight and the final LLM activations into a smaller vector space, since that is basically just cosine similarity ranking step (so that would reduce the head.weight size). But again, it must be tried many times and just not working as well. LLM space is pretty saturated with tricks.
liuliu··on Explanation of INT8 ConvRot (FP8 is no longer needed)
One thing is not obvious to me is how ConvRot can be applicable beyond diffusion models. Especially for LLM decoding, as each ConvRot would be more expensive for a given decoding vector, and it is required now, so you cannot easily get the benefit for prefill only, while maintaining the same decoding performance.
liuliu··on MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
It is a well-known trick, given that the timestep is between 0 to 1, you can slicing them at any resolution (1000, or 10000, give or take), and then keep a look-up table for modulation scale / bias etc for each. It is quite different from quantization and it is indeed lossless.

It is also only applicable to diffusion models as only these operates at per-timestep.

liuliu··on U.S. debt-to-GDP ratio reaches 123%
Well, it is a "platform of Balancing Budget", you don't need to actually work on that or do anything. The harder part is just letting people believing in miracles. Look no further than the current administration. Lying has no consequences so far.
liuliu··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
DS4 is designed to do real-work. Gemma 4 is not going to cut it.
liuliu··on “We have information that Moonshot distilled Fable for the development of K3”
> training an LLM takes more resources and expertise than distilling from an existing LLM

This is not automatically true. Training and distillation use the same underlying infra and method and there is no intrinsic differences in between.

liuliu··on Bonsai 27B: A 27B-Class model that runs on a phone
It is disabled because it doesn't work :) Try it and see the doom loop it gets itself in.
liuliu··on Bonsai 27B: A 27B-Class Model that runs on a phone
Note that 3.5 9B cannot do thinking (while 3.6 27B can, pretty effectively, quite verbosely).
liuliu··on Bonsai 27B: A 27B-Class model that runs on a phone
You also need to pay close attention to BFCLv3 multi-turn result, that helps you to get a sense how frequently these quants will be in a doom loop.
liuliu··on Bonsai 27B: A 27B-Class model that runs on a phone
The problem, of course, is if you run the UD_Q2 variant (Unsloth) which does only post-training, the number is pretty close to 1-bit model here and the 5% drop in tool-call is significant than it suggests in real-life use cases.
liuliu··on Grok CLI uploaded the whole home directory to GCS
This should be the first comment here.

Too many replies here are done before reading it. It is not "just another agent does the agent thing". It is a deliberate choice of the Grok Build team to have a toggle from the server to let the program to upload your entire codebase to a Google Cloud Storage bucket. It is not an agent decision, the program is written by the Grok team, can be dissembled and seeing the logic wild-open.

liuliu··on How to setup a local coding agent on macOS
Realistically, you need to experiment with any user prompt + a good amount of system prompt (at least > 1000 tokens, but realistically, in the range of 3000 tokens probably good).

llama.cpp includes tools for that, what you are looking at is to have a prefill before token generation to measure it properly. Increasingly also, measuring token generation speed at longer context (32k or 64k) is important too.

liuliu··on Speculative KV coding: losslessly compressing KV cache by up to ~4×
It is a “research note”. It might not pan out, and you might say it doesn’t deserve the attention on the internet. But it did suggest something that resembles of compression, just no experiment done for that.
liuliu··on Superintelligence: The Idea That Eats Smart People (2016)
I actually agree. At some point, a RSI system has to interact with real-world, and that imposes serialization constraints. It is harder to know how much that slow-down would be and how much speed-up we will get before that. But a RSI cannot simply be a exponential growth forever.
liuliu··on 1-Bit Bonsai Image 4B Image Generation for Local Devices
Except the two (GPT-Image-2 and Nano Banana Pro), anything displayed here can run on the 16 GiB MacBook (including the FLUX.2 [dev]): https://tests.drawthings.ai/generate
liuliu··on 1-Bit Bonsai Image 4B Image Generation for Local Devices
> To our knowledge, Bonsai Image 4B is the first image model in its parameter class to run directly on an iPhone.

This is wrong. But they worded it carefully to be not entirely wrong.

FLUX.2 [klein] 4B (the same parameter class, basically the same model) runs on iPhone through Draw Things app, with 8-bit or 6-bit quantization (hence not "directly", I guess, but that is the technicality that sounds fishy enough).

Page 1 of 34Next →