HNHacker News
TopNewBestAskShowJobs

fspeech

2,766 karma · joined January 14, 2013

submissionscomments
fspeech··on DeepSeek Open Source Optimized Parallelism Strategies, 3 repos
It appears that they have really good toolings so they can focus on hiring really smart new grads and giving them a lot of leeway on choosing what to work on without worrying about how productive they will be. One particular memorable response when the CEO Liang was asked about not using KPI goes roughly like this: since we all reviewd and worked with the candidate, if he is not productive it must be our fault to not have given him the right support to make him productive.
fspeech··on DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling
Efficiency could lead to cheaper hardware for everyone, themselves included.
fspeech··on DeepSeek Open Infra: Open-Sourcing 5 AI Repos in 5 Days
You need that to optimize load balancing. Unfortunately that gain is not available to small or individual deployment.
fspeech··on Grok3 Launch [video]
A better approach is to split the model with MOEs running on CPUs and MLAs running on GPU. See the ktransformers project: https://github.com/kvcache-ai/ktransformers/blob/main/doc/en...

This takes advantage of the sparsity of MOE and the efficient KV-cache of MLA.

fspeech··on It's not just AI. China's medicines are surprising the world, too
China proves that you don't need 100 times return on capital to get potential innovators to take risk. 100% may be enough. People say China couldn't innovate due to its weak IPR protection, but it is turning out that you don't need to be an absolutist. Sure outright theft is bad, but imitation can be good for competition. You want to give the innovators a head start so they can profit from their efforts, but you don't want to make the protection so tight that they can sit on their initial efforts and profit for years to come without further innovation.
fspeech··on Ask HN: Is there a way to use Baidu cloud legally from US/EU?
You need a Chinese number to sign up and facial verification to make payments, as far as I can tell.
fspeech··on SambaNova Cloud Launches the Fastest DeepSeek-R1 671B (198 tok/s)
The smaller cluster is not necessarily a good thing. It may not be able to take advantage of unbalanced load on experts.
fspeech··on Alibaba's Tsai Says Apple iPhone Will Use His Firm's AI in China
Deepseek doesn't have the infrastructure to support Apple. I suspect that they are also not interested in tailoring to Apple's needs given their mission and small size.
fspeech··on BYD to offer Tesla-like self-driving tech in all models for free
In a polytheistic context the invocation is much less severe, roughly equivalent to Apollo's Eyes or Zeus's Eyes.
fspeech··on BYD to offer Tesla-like self-driving tech in all models for free
That's due to the translation. The original term 天神 comes from the polytheistic Chinese folk religion so it doesn't have the same connotation.
fspeech··on Run Deepseek from fast NVMe drives
Deepseek's open source inference code, while correct, may not be fully efficient. For example the MLA needs the right associative matrix multiplication order to be efficient.
fspeech··on Run Deepseek from fast NVMe drives
Do you have any benchmark run yet? I am interested in knowing how many tokens/sec you can get to. Though in the end it should be more efficient to run the model on distributed server clusters.
fspeech··on The Race to God Mode
Thanks for the thought provoking take.
fspeech··on SemiAnalysis says DeepSeek spent more than it claims
Right. And the number is based on rental rates of GPUs so how many GPUs they own is irrelevant to the claim.
fspeech··on Understanding Reasoning LLMs
Recent developments like V3, R1 and S1 are actually clarifying and pointing towards more understandable, efficient and therefore more accessible models.
fspeech··on Trump slaps tariffs on Mexico, Canada and China
Unless the world demand for these raw material goes down or production elsewhere goes up someone will still buy these raw materials. If tarriff is high enough to force US to source from and Canada to sell to more distant places the world will be less efficient.
fspeech··on Nvidia’s $589B DeepSeek rout
The key question is: has demand elasticity increased for Nvidia cards? An increase in elasticity means people are more willing to wait for hardware price to drop because they can do more with existing hardware. Elasticity could increase even if demand is still growing. Not all growths are equally profitable. Current high prices are extremely profitable for Nvidia. If elasticity is increasing future growth may not be as profitable as the projection from when Deepseek was relatively unknown.
fspeech··on Nvidia, ASML plunge as DeepSeek triggers tech stock selloff
That's not how commodities are priced. Commodities are priced at the marginal cost of production.
fspeech··on Nvidia, ASML plunge as DeepSeek triggers tech stock selloff
Inference can be done on much more diverse hardware.
fspeech··on Nvidia, ASML plunge as DeepSeek triggers tech stock selloff
You need to also look at valuation and how much growth justifies that valuation.
fspeech··on Nvidia’s $589B DeepSeek rout
You are not taking into account why people are willing to pay exceedingly high prices for GPUs now and that the underlying reason may have been taken away.
fspeech··on The impact of competition and DeepSeek on Nvidia
It doesn't make sense to compare individual models. A better way is to look at total compute consumed, normalized by the output. In the end what counts is the cost of providing tokens.
fspeech··on The impact of competition and DeepSeek on Nvidia
But why would the customers accept the high prices and high gross margin of Nvidia if they no longer fear missing out with insufficient hardware?
fspeech··on The impact of competition and DeepSeek on Nvidia
If you normalize Nvidia's gross margin and take into account of competitors sure. But its current high margin is driven by Big Tech FOMO. Do keep in mind that 90% margin or 10x cost to 50% margin or 2x cost is a 5x price reduction.
fspeech··on DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
Since intermediate steps of reasoning are hard to verify they only award final results. Yet that produces enough signal to produce more productive reasoning over time. In a way when pigeons are virtual one can afford to have a lot more of them.
fspeech··on DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
https://www.macrotrends.net/stocks/charts/NVDA/nvidia/gross-...

GPU prices could be a lot lower and still give the manufacturer a more "normal" 50% gross margin and the average researcher could afford more compute. A 90% gross margin, for example, would imply that price is 5x the level that that would give a 50% margin.

fspeech··on DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
It does if the spend drives GPU prices so high that more researchers can't afford to use them. And DS demonstrated what a small team of researchers can do with a moderate amount of GPUs.
fspeech··on DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
Not necessarily if you are pushing against a data wall. One could ask: after adjusting for DS efficiency gains how much more compute has OpenAI spent? Is their model correspondingly better? Or even DS could easily afford more than $6 million in compute but why didn't they just push the scaling?
fspeech··on Trae: An AI-powered IDE by ByteDance
How could more competition be bad? Sure some startups' business plans no longer work and have to pivot but many more startups now have access to better products at lower cost.
fspeech··on China's grandiose plot to take on Rolls-Royce
Not sure why it's called grandiose. They are worried that supplies of jet engines could be cut off. Their own version doesn't have to be better if the alternative is no jet engine.
← PreviousPage 5 of 34Next →