1,017 karma · joined April 3, 2021
This is insanely slow given its 200+GB/s memory bandwidth. As a comparison, I've tested GPT OSS 120B on Strix Halo and it obtains 420tps prefill and >40tps decode.
"4x gated residual streams" look quite weird. Is there any paper or technique report for this?
I am not 100% sure that all of the mitigation overhead
comes from syscalls, but it stands to reason that a lot
of it arises from security hardening in user-to-kernel
and kernel-to-user transitions.
Will io_uring be also affected by Spectre mitigations given it has eliminated most kernel/user switches?And did anyone do a head-to-head comparison between io_uring and DPDK?
-vv ⇒ -w
-vvvv ⇒ -ww
Just a joke :) > The story was reported in China but is tricky to find. Sina (Chinese language), rough translation.
> ...
The original Chinese title of Sina report: > 中国承诺为乌克兰提供核保护伞?假!
Translation: > China pledges nuclear umbrella for Ukraine? It's false!
And in later paragraphs the Sina report distinguishes "security assurance" from "nuclear umbrella" by explaining the concepts of positive and negative security assurances and citing multiple documents including UNSC resolution 255.So it is reported by Chinese media like an ordinary statement on non-nuclear-weapon countries according to UN resolutions.
1. The death toll is the same, but the number of confirmed cases reported by China CDC is much smaller than WTO (380k vs 710k). I guess people caught covid with no symptoms are excluded by China CDC.
2. After subtracting death toll in Hong Kong and Taiwan, there're only 4.6k deaths in China mainland since 2020.