HNHacker News
TopNewBestAskShowJobs

ycui7

87 karma · joined September 21, 2011

submissionscomments
ycui7··on DeepSeek Harness Desktop for macOS and Windows
you do can change font size in settings, but not with keyboard shortcut. for some reason, they decided to limit font size to 17 max.
ycui7··on DeepSeek Elastic Compute (DSec)
click the link, or read the PDF.

all authors are listed. there is no conspiracy to hide authorship.

arxiv simply want to keep the page short not too long.

ycui7··on The new CC, an AI agent built for families
remind me in the old time, we use cc to compile code to machine code. it was cc not gcc.
ycui7··on DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
llm is not a deterministic program. the same model won't even return deterministic answer. what's the point keep the model freezed?

if you want deterministic returns, you should set the temperature to 0 to get the best possibility of deterministic.

ycui7··on Run Qwen3.8 27B locally: real numbers from my Mac Studio
the fact that this author cannot get qwen3.8-27b run at the same speed as qwen3.6-27b, says the article is not worth reading. the author does not know anything about how to run local AI. 3.8 and 3.6 are the same model with different weight.

both tg and pp speed are so terrible on author's machine.

ycui7··on GLM-5.3-Flash
because it has 1T ssd not 4T
ycui7··on Gemini 3.7 Flash
Cerebras is an uncut whole wafer. Each wafer gets you 44GB SRAM (not a typo, SRAM, not DRAM/VRAM). A few years ago, leading process node wafer from tsmc is $20K/ea without guarantee on yield.

A single full wafer likely can run qwen3.6-27b alone. But won't be enough to run bigger models, which are pretty much all popular models.

ycui7··on Qwen 3.8 27B
if you have the VRAM, use offical release. quantized model lose focus after long context and can do damages or thinking loop
ycui7··on Qwen 3.8 27B
that is why we enable web search for the agent. the memory can come from the internet.

deepseek-v4-flash needs web search to return true facts.

ycui7··on Qwen3.8-2.4T
the a little disappointing part is this is released in BF16. so i suppose no QAT was implemented.
ycui7··on DeepMind's WeatherNext model achieves breakthrough forecasting cyclones
on one side, deepmind makes a lot of advancement in science related application. but, on the commercial side, they struggle to compete with other major LLM providers.
ycui7··on DeepSeek V4 Flash 0731
it is funny when people say i am struggling to spend money.
ycui7··on California Town Says Flock Cameras Misread License Plates 71% of the Time
they need a competent gov contractor. recognizing license plate is a fully solved problem many years ago.
ycui7··on AMD acquires Taalas to boost inference performance by etching models in silicon
so qwen3.x-27b on hardware? or better deepseek-v4-flash on hardware .
ycui7··on Qwen3.8-Max: A New Bar for Coding and Cowork
Can they still go public ? MiniMax M3 Pro is also coming, then DeepSeek-v4-Pro GA, then GLM5.5. There will only be bad news for them in the coming few weeks/months.
ycui7··on DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
the rational in one’s mind is similar to buying expensive supercar but no driving it daily.

owning a few GPUs is a lot cheaper than supercars.

ycui7··on DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
Huawei Ascend NPU
ycui7··on DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
my AC is noiser than my GPU server.
ycui7··on DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
you don’t need new weight. try vllm-moet from github. it will autogenerate 2-bit plane.
ycui7··on Moonshot’s Kimi uses 20k Nvidia chip cluster from Alibaba
K3 is natively trained to mxfp4, if they cannot get a hold of Blackwell chip, it is meaningless. Hopper does not do native 4-bit floating math.

Either they have Blackwell with native 4-bit floating math, or they use have Chinese domestic NPU that support mxfp4 natively.

The article’s statement does not make sense.

ycui7··on DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
For people with single RTX PRO 6000 96GB or DGX Spark 128GB, vllm-moet is a very good engine, although lesser known. It auto generate a symmetric 2-bit plane for inference and also generate a 4-bit delta cache to recover precision. Support ssd streaming oversized weight. You pick how much VRAM to allocate to each to balance out speed vs precision. 170 tps with ds-v4-flash demonstrated.

It use the stock model, no new models requires.

Worth spend a few hours to try.

The DGX Spark requires a small hack to ignore the difference between sm120 vs sm121, but it does run on sm121.

ycui7··on A solid-state “atomic channel” for separating rare earth elements
and it was created by Chinese born Professor and Student.
ycui7··on OpenAI says its AI went rogue and launched 'unprecedented' cyber-attack
so GLM won?
ycui7··on Who's afraid of Chinese models?
discounted competitor could cheat. they can offer subpar model response and sell it as deepseek-v4. it is uneconomical to prove inference providers are cheating, so they get away with it. cheating inference provider does not care if their customers stay.
ycui7··on Xiaomi-Robotics-1
feels like the American domestic manufacturing is done, there is no hope to save it.
ycui7··on Qwen 3.8
it takes extra effort to open source a model even if you had it running internally. even traditional software takes extra effort to get released as open source.
ycui7··on Qwen 3.8
every cloud provider trains on your data, regardless of what they promise. real user interaction is the best reinforcement-learning trace.
ycui7··on Qwen 3.8
OpenAntrophic and OpenOpenAI ?
ycui7··on China recovers Long March 10B rocket
China successfully recovers Long March 10B rocket following maiden flight, marking a breakthrough in rocket reusability
ycui7··on Muse Spark 1.1
this is not subsidizing. this is way too expensive for a no-name model.
Page 1 of 2Next →