HNHacker News
TopNewBestAskShowJobs

petu

309 karma · joined November 7, 2025

submissionscomments
petu··on Keyboard differences between Windows and Macs
First key on the right after space bar.
petu··on Why isn't the industry freaking out about DeepSeek 4.1 Flash?
There's no BF16, original full quality weights are quantized already and 510GB.

Then good portion of those weights are n-grams (~200GB) that don't need to be in VRAM.

Then KV cache of that model is super lightweight at ~1GB per 1M tokens. If HBF succeeds, then accelerator with 16GB of VRAM and 1TB HBF/NAND is probably all you need (?).

petu··on Beam: Reflection's 501B open-weight model
I guess public datasets on HuggingFace and some shadow libraries content is enough to start.

e.g. fineweb dataset is 50TB https://huggingface.co/datasets/HuggingFaceFW/fineweb

petu··on Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
6B activated weights per token vs 27B. Something like DGX Spark is way better suited for Flash Next.
petu··on Sonnet 5.5
That's max output tokens per response limit, separate from context length
petu··on U.S. appeals court upholds designation of Anthropic as supply chain risk
Devil's advocate, but Anthropic talks a lot about importance of alignment, so.. How can military ensure that model is NOT refusing to work just to be malicious?

Anthropic doesn't provide model w/o guardrails, suppliers have to use public version, newer model releases surely know about Anthropic standoff with the military.. how to prevent model from realizing that it's doing something for the military (supplier) and sandbagging and/or subtly sabotaging the results?

=====

DOD recommends sandbagging as defense mechanism against destination "attacks":

https://media.defense.gov/2026/Sep/08/2003992823/-1/-1/1/CSA...

Page 13

  When suspecting a malicious distillation campaign, consider varying changes to responses across requests to complicate response quality evaluations, such that the subtle changes avoid triggering obvious alerts. Reducing reasoning depth, presenting correct information with different reasoning, or stylistic inconsistencies may evade detection while reducing training usefulness.

  Avoid informing China-based AI company users suspected of distillation campaigns of a switch to a downgraded model. Informing malicious distillers would enable them to improve their defense evasions and indicate when to roll back training.
If that's seen as valid defense mechanism, then it's also a risk if used against them.
petu··on Xiaomi MiMo v2.6
n-grams can be kept on SSD, no need to hold them in any kind of RAM (at least w/o batching)
petu··on Turn off and restrict access to Apple Intelligence features on Mac
> meander our way through Spotlight

By that you mean that just typing prompt into Spotlight doesn't work reliably? For me after few words "Ask Siri" is first result.

petu··on Raspberry Pi blocks changing RAM chips
Selling modded board as official 8GB SKU with reliability/warranty expectations.
petu··on Apple Reference Image: A New Approach for Verified Photography
Was photo not authentic? Think of it "as seen by an iPhone", not "this is authentic event" verification.

But Apple likely would reject such photo because of inappropriate depth map / LiDAR data.

petu··on JetKVM Mini
Sipeed has such device https://wiki.sipeed.com/hardware/en/kvm/NanoKVM_USB/introduc...
petu··on Proof of Capture: Apple Reference Image, but open source and using steganography
This feature is Pro phones only, not Duo: https://www.apple.com/iphone/compare/ ("Apple Reference Image (Fusion Main)")

So only on devices with LiDAR / that can capture depth map.

petu··on Proof of Capture: Apple Reference Image, but open source and using steganography
Miniature dioramas wouldn't be size appropriate. Apple could detect faces/cars/other common objects of ~known size and verify -- or even just dump depth map for anyone to check.
petu··on Proof of Capture: Apple Reference Image, but open source and using steganography
It seems to be what Apple is doing, this feature is only available on the 18 Pro's (which have depth sensor on the back), but not Duo.
petu··on Hitachi launches CO2 heat pump water heaters with solar-friendly tariff controls
Aren't they meant for underfloor heating? Heat pump efficiency drops with delta T increase, but iron radiators need 60-80C supply.
petu··on DeepSeek v4.1 Flash
You're looking at third party providers.

V4 Flash prices served by DeepSeek themselves:

  launch pricing: $0.0028 / $0.14 / $0.28 
  after Aug 16th: $0.007  / $0.22 / $0.66 during off-peak.
  after Sep 10th: $0.003  / $0.15 / $0.60 during off-peak.
Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, rate is doubled.

https://api-docs.deepseek.com/quick_start/pricing (archive.org for old)

petu··on DeepSeek v4.1 Flash
It's larger than previous V4 Flash.

  552B in ~FP4, 306GB.   
  196B of FP8 Engrams, another 204GB, not necessary to keep in RAM.  
  KV cache sees another 4x size reduction, just 900MB for 1M.  
So 384GB needed for a chance of achieving useful speeds. Three Sparks or quad RTX PRO 6000.
petu··on DeepSeek v4.1 Flash
V4 Flash also was released as mostly FP4, but this one is FP8 (?). 160GB vs 510GB.

Original Flash good fit for dual Spark / Strix Halo machines. This one would require third party quants and even then 4 machines.

Edit: Most of added weights/size are Engrams?

> Overall, DeepSeek-V4.1-Flash has 552B backbone parameters and 196B Engram parameters, activating 8B parameters per token during prefill and 16B during decode.

Those can stay on SSD. So I guess / it possible, that non-engram portion is still FP4 of ~same size! Need to read tech report.

petu··on DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
> Hoje em sites como openrouter o valor é de $0.16 output .

You're comparing different providers then. DeepSeek price on OpenRouter is $0.66 output.

petu··on DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
A month ago new V4 Flash 0731 checkpoint was better than existing V4 Pro. They've kept serving Pro, it was updated 13 days later (0813 checkpoint).

Now, 4 weeks later new Flash checkpoint (0910?) is again better than existing Pro. Same situation, but Pro is taken offline this time.

petu··on DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
4.1 releases tomorrow, right now you're supposed to be served by same old model
petu··on OpenAI's GPT-6 Astra on ARC-AGI-3
OpenAI provides API key with ~unlimited use?
petu··on AnkiDroid: Google Play no longer allowing Open Collective donation link
few weeks ago
petu··on GPU World
Do you think average human cares about "our species is going extinct" stuff? If not, why couldn't they be happy at the same time?

You seem to equate birth rate with happiness, but ignore efficacy and availability of modern contraception.

petu··on Hy4 preview
Before we worry about source code, Microsoft doesn't grant me rights to modify/redistribute/sell copy of Windows I have.
petu··on Samsung's Processing-in-Memory (PIM)
Yes, but running out of RAM is impractical due to low memory bandwidth.

According to the article/Samsung RAM dies inside can support way higher bandwidth than they expose, they're limited by external interface / bus width:

> Together, they can utilize the chip’s internal bandwidth across all 16 banks, which comes out to 614 GB/s. For comparison, regular DRAM accesses can hit two banks in parallel and max out at 76.8 GB/s.

And that's just for single 64-bit IC. So way faster and more power efficient.

petu··on GLM-5.3 is now open-weight
This time they just made FP8 "default", accompanied by "-BF16" model/page (previously "-FP8" was released alongside).
petu··on GLM-5.3 is now open-weight
I assume that's about 5.3 Flash, not full?
petu··on Air Conditioning Is Not a Luxury, It Is a Necessity
Air temp is measured in the shade, sun hitting windows/interior floors would take you past that.

IR blocking film/tint on outside works great if you can't get windows shaded.

petu··on Qwen3.8-Flash-Next
You need VRAM for the whole thing for optimal performance. Activation is chosen "randomly" for each token. PCIe becomes bottleneck, so much that just doing computation on CPU is likely faster.

But given it's only 6B, out of which only ~2.4B seem to be actually routed ("selected at random per token"), you could get reasonable performance with experts on CPU (still haven't tested, but 20-30 for dual channel DDR5 and 4 bpw quant).

Page 1 of 5Next →