309 karma · joined November 7, 2025
Then good portion of those weights are n-grams (~200GB) that don't need to be in VRAM.
Then KV cache of that model is super lightweight at ~1GB per 1M tokens. If HBF succeeds, then accelerator with 16GB of VRAM and 1TB HBF/NAND is probably all you need (?).
e.g. fineweb dataset is 50TB https://huggingface.co/datasets/HuggingFaceFW/fineweb
Anthropic doesn't provide model w/o guardrails, suppliers have to use public version, newer model releases surely know about Anthropic standoff with the military.. how to prevent model from realizing that it's doing something for the military (supplier) and sandbagging and/or subtly sabotaging the results?
=====
DOD recommends sandbagging as defense mechanism against destination "attacks":
https://media.defense.gov/2026/Sep/08/2003992823/-1/-1/1/CSA...
Page 13
When suspecting a malicious distillation campaign, consider varying changes to responses across requests to complicate response quality evaluations, such that the subtle changes avoid triggering obvious alerts. Reducing reasoning depth, presenting correct information with different reasoning, or stylistic inconsistencies may evade detection while reducing training usefulness.
Avoid informing China-based AI company users suspected of distillation campaigns of a switch to a downgraded model. Informing malicious distillers would enable them to improve their defense evasions and indicate when to roll back training.
If that's seen as valid defense mechanism, then it's also a risk if used against them.By that you mean that just typing prompt into Spotlight doesn't work reliably? For me after few words "Ask Siri" is first result.
But Apple likely would reject such photo because of inappropriate depth map / LiDAR data.
So only on devices with LiDAR / that can capture depth map.
V4 Flash prices served by DeepSeek themselves:
launch pricing: $0.0028 / $0.14 / $0.28
after Aug 16th: $0.007 / $0.22 / $0.66 during off-peak.
after Sep 10th: $0.003 / $0.15 / $0.60 during off-peak.
Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, rate is doubled.https://api-docs.deepseek.com/quick_start/pricing (archive.org for old)
552B in ~FP4, 306GB.
196B of FP8 Engrams, another 204GB, not necessary to keep in RAM.
KV cache sees another 4x size reduction, just 900MB for 1M.
So 384GB needed for a chance of achieving useful speeds. Three Sparks or quad RTX PRO 6000.Original Flash good fit for dual Spark / Strix Halo machines. This one would require third party quants and even then 4 machines.
Edit: Most of added weights/size are Engrams?
> Overall, DeepSeek-V4.1-Flash has 552B backbone parameters and 196B Engram parameters, activating 8B parameters per token during prefill and 16B during decode.
Those can stay on SSD. So I guess / it possible, that non-engram portion is still FP4 of ~same size! Need to read tech report.
You're comparing different providers then. DeepSeek price on OpenRouter is $0.66 output.
Now, 4 weeks later new Flash checkpoint (0910?) is again better than existing Pro. Same situation, but Pro is taken offline this time.
You seem to equate birth rate with happiness, but ignore efficacy and availability of modern contraception.
According to the article/Samsung RAM dies inside can support way higher bandwidth than they expose, they're limited by external interface / bus width:
> Together, they can utilize the chip’s internal bandwidth across all 16 banks, which comes out to 614 GB/s. For comparison, regular DRAM accesses can hit two banks in parallel and max out at 76.8 GB/s.
And that's just for single 64-bit IC. So way faster and more power efficient.
IR blocking film/tint on outside works great if you can't get windows shaded.
But given it's only 6B, out of which only ~2.4B seem to be actually routed ("selected at random per token"), you could get reasonable performance with experts on CPU (still haven't tested, but 20-30 for dual channel DDR5 and 4 bpw quant).