HNHacker News
TopNewBestAskShowJobs

snake_doc

685 karma · joined April 18, 2022

submissionscomments
snake_doc··on Secure temporary file sharing for AI agents and humans
Hmm.. why not just use the free and MIT licensed version: https://github.com/magic-wormhole/magic-wormhole

you can host the relay and mailbox servers yourself as well

snake_doc··on DeepSeek API Pricing Update
lol what no, AI competition is super competitive in China; bytedance has 50% of the inference market and mostly serves from outside of China

https://www.bloomberg.com/news/articles/2026-06-17/microsoft...

Bytedance which runs China’s most popular Doubao AI chatbot; is spending $70B in CapEx this year, most of it outside of China (Malaysia, Thailand, Brazil, etc; and they are allowed to lease NVIDIA chips). This is roughly 50% of Microsoft CapEx.

https://www.tomshardware.com/pc-components/gpus/chinas-byted...

snake_doc··on DeepMind's WeatherNext model achieves breakthrough forecasting cyclones
> We can now generate a single 15-day forecast in less than a minute on a TPU, empowering forecasters to quickly evaluate the probability distribution of potentially devastating tail-risks.

Crazy

snake_doc··on Qwen 3.8
Not directly, relevant, but Alibaba (maker of Qwen) actually owns about ~20-30% of MoonshotAI (the maker of Kimi K3).
snake_doc··on Google limits Meta's use of its Gemini AI models
Image/video understanding still quite cost effective from the Gemini flash series models?

Image generation and veo models I’d imagine quite effective for creators; new Instagram accounts with AI content that are garnering millions of followers in spans of weeks are quite common now

snake_doc··on 2025 Letter
> That’s not change unique to America though. Why would I be thinking of changes that have affected almost every country, when talking about whether Americans recognize how America has changed?

Hahahaha, this is like saying, the world wars didn’t impact Europe, because it also impacted the whole world! Europeans, the war didn’t happen!! Anyways… this entire thread is more evidence that European stereotypes are valid for the most part

snake_doc··on 2025 Letter
> Most Americans have no sense of how very different their country is now from say, the country that launched the Apollo missions.

You cannot be kidding right? Those that remember the Apollo missions will undoubtedly agree their country is different, first but not least they are most likely using a smartphone assembled and designed with technology unimaginable by NASA planning the Apollo missions; not only that, the smartphone is assembled half way around the world by a country previously in such dire poverty and famine that over 30M died due to Marxist central planning.

snake_doc··on 2025 Letter
What exactly is wrong with Americans or for the most part the rest of the world valuing economic performance as a measure of prosperity and progress?

Your comment is again another anecdote confirming European stereotypes. It’s not a “trap”, it’s a different world view.

snake_doc··on 2025 Letter
Would you be able provide some evidence to the contrary when it comes to the topics discussed in the letter?

On industrial infrastructure

On technology innovation

On internet regulation

On central planning

Otherwise, your comment becomes an anecdote supporting the common stereotypes (assuming you’re from Europe).

snake_doc··on Lessons from the PG&E outage
Cell towers just need power to keep functioning, starlink adds no utility in an urban dense environment with fiber.
snake_doc··on GPT-5.2
> Models were run with maximum available reasoning effort in our API (xhigh for GPT‑5.2 Thinking & Pro, and high for GPT‑5.1 Thinking), except for the professional evals, where GPT‑5.2 Thinking was run with reasoning effort heavy, the maximum available in ChatGPT Pro. Benchmarks were conducted in a research environment, which may provide slightly different output from production ChatGPT in some cases.

Feels like a Llama 4 type release. Benchmarks are not apples to apples. Reasoning effort is across the board higher, thus uses more compute to achieve an higher score on benchmarks.

Also notes that some may not be producible.

Also, vision benchmarks all use Python tool harness, and they exclude scores that are low without the harness.

snake_doc··on IBM CEO says there is 'no way' spending on AI data centers will pay off
China added ~90GW of utility solar per year in last 2 years. There's ~400-500GW solar+wind under construction there.

It is possible, just may be not in the U.S.

Note: given renewables can't provide base load, capacity factor is 10-30% (lower for solar, higher for wind), so actual energy generation will vary...

snake_doc··on Modular Manifolds
Wot? Is this what AI generated non-sense has come to? This is totally unrelated.
snake_doc··on Modular Manifolds
Aren’t they all optimization techniques at the end of the day? Now you’re just debating semantics
snake_doc··on Modular Manifolds
Hmmm… http://www.incompleteideas.net/IncIdeas/BitterLesson.html
snake_doc··on Modular Manifolds
Um.. the model is tiny: https://github.com/thinking-machines-lab/manifolds/blob/main...
snake_doc··on Trump to impose $100k fee for H-1B worker visas, White House says
Oh? And taxes can’t be used to buy influence and votes? How naive… Money is fungible… one pocket into another

Exhibit 1: Tariff revenues to bail out American farmers: https://www.ft.com/content/0267b431-2ec9-4ca4-9d5c-5abf61f2b...

snake_doc··on Trump to impose $100k fee for H-1B worker visas, White House says
SCMP is owned by Alibaba, which is subject to the purview of the Chinese Central Government [1].

[1]: https://www.cecc.gov/agencies-responsible-for-censorship-in-...

snake_doc··on Trump to impose $100k fee for H-1B worker visas, White House says
Mafia behavior continues… (not my observation, but the Texas senator’s Ted Cruz[1]).

$100k is a big pizzo (protection fee)!

[1]: https://www.bloomberg.com/news/articles/2025-09-19/ted-cruz-...

> “That’s right outta ‘Goodfellas,’ that’s right out of a mafioso going into a bar saying, ‘Nice bar you have here, it’d be a shame if something happened to it,’” Cruz said, using the iconic New York accent associated with the Mafia.

snake_doc··on DuckDB is probably the most important geospatial software of the last decade
Are you querying from an EC2 instance close to the S3 data? Are the CSVs partitioned into separate files? Does the machine have 500GB of memory? It’s not always duckdb fault when there can be a clear I/O bottleneck…
snake_doc··on We built a modern data stack from scratch and reduced our bill by 70%
These just seems like over engineered solutions trying to guarantee their job security. When the dataflows are so straight forward, just replicate into pick your OLAP, and transform there.
snake_doc··on DualPipe: Bidirectional pipeline parallelism algorithm
Hmm weren’t there also supposed to be the SM re-allocation, doesn’t look like it was included; I may have been mis-remembering the explanation.
snake_doc··on Auto-Differentiating Any LLM Workflow: A Farewell to Manual Prompting
Valid point, but at least that was mathematics. This paper isn’t even math, it’s a data control flow masquerading as math.
snake_doc··on Auto-Differentiating Any LLM Workflow: A Farewell to Manual Prompting
Holy unnecessary use of terminology to explain a reverse graph traversal. “Loss”, “gradients”, “differentiating”— no! stop!

This must be what AI hype actually is. Complete incoherent language to explain a very straight forward concept.

This is just: LLMs judging intermediate node outputs, and reverse traversing the graph while doing so until it modifies the original prompt.

snake_doc··on On DeepSeek and export controls
Without taking a position on unipolar vs. multi-polar:

Dario makes an astounding implicit assumption:

- China originating labs cannot acquire chips providing 80-90% similar utility without the US within the next 2-3 years.

I'll make an observation, re: DeepSeek's incentives that drove them to create the innovations from the V2 and V3 papers.

DeepSeek, compared to American AI labs, are much more compute constrained, but in a unique way. Their chips are more memory bandwidth constrained (depending on type anywhere from 50% to 80% less bandwidth).

Therefore, each dollar/hour of investment towards memory optimization is worth MORE to DeepSeek than to American labs.

In the V2/3 paper, they've demonstrated exactly that with these memory optimization techniques.

1. MLA -> reduces KV cache by nearly 80% compared to GQA. By the way, this was published in V2 in May 2024.

2. FP8 matmul (while still accumlating in FP32 gradients) without losing significant quality.

3. DualPipe scheduling and reworking of Hopper SM's allocation on communication vs. computation -> DeepSeek's V3 paper has 2 full pages of hardware suggestions for "hardware designers" (read NVIDIA)

Export controls in a global market create different incentives in parties. The resulting incentives will change, and agents (using it as an traditional economics term) will change their capital allocation strategy.

snake_doc··on Microsoft Probing If DeepSeek-Linked Group Improperly Obtained OpenAI Data
You should still be mocked.

1. ChatGPT data is widely on the internet, just google Sharegpt dataset and you can scrap 200k+ conversations with a few stroke of huggingface commands. These were then used by the open source community like Vicuña models, there was a period of several months in the open source community where RLAIF was all the rage; so this data populated the internet. So if a company is crawling and scraping the internet, this will eventually be in the dataset.

2. The v3 deepseek model was trained on 15T tokens. Please educate yourself and calculate how long (in latency, inference for 1k token output will take almost 30seconds) and cost it would be to extract 15T tokens from ChatGPT / Azure API. Granted API accounts all have spend limits, and will trip fraud detection on OAI billing, how long would the subterfuge had to take place? With which model? At what time? Wouldn’t they have to keep repeating this for subsequent generation of OAI models?

3. OAI didn’t invent MLA, they didn’t invent multi token prediction with disconnected ROPE, they didn’t invent FP8 matmul training dynamics (while accumulating in FP32) without losing significant quality.

So go away

snake_doc··on Berkeley Researchers Replicate DeepSeek R1's Core Tech for Just $30: A Small Mod
@dang please link to either the GitHub https://github.com/Jiayi-Pan/TinyZero

or the primary source twitter thread: https://x.com/jiayi_pirate/status/1882839370505621655

snake_doc··on Nvidia’s $589B DeepSeek rout
No, the only thing that matters is if the portfolio delivers returns in excess of your cost of capital.

If your portfolio is green, you can still be a poor performer.

snake_doc··on Nvidia’s $589B DeepSeek rout
The other way is certainly also true. Your short piece is rational, but lacks insight into the inference and training dynamics of ML adoption unconstrained.

The rate of ML progress is spectacularly compute constrained today. Every step in today’s scaling program is setup to de-risked the next scale up, because the opportunity cost of compute is so high. If the opportunity cost of compute is not so high, you can skip the 1B to 8B scale ups and grid search data mixes and hyperparameters.

The market/concentration risk premium drove most of the volatility today. If it was truly value driven, then this should have happened 6 months ago when DeepSeek released V2 that had the vast majority of cost optimizations.

Cloud data center CapEx is backstopped by their growth outlook driven by the technology, not by GPU manufacturers. Dollars will shift just as quickly (like how Meta literally teared down a half built data center in 2023 to restart it to meet new designs).

snake_doc··on Show HN: I made an open-source laptop from scratch
Okay, I'll help him humble brag:

Bryan is in his last year of high school.

</end>

Keep building!

Page 1 of 6Next →