HNHacker News
TopNewBestAskShowJobs

gdiamos

871 karma · joined December 18, 2014

submissionscomments
gdiamos··on Stay discoverable in search while disallowing AI training
I wouldn’t trust an AI company to honor this as far as I could throw them
gdiamos··on More questions about whether researchers can trust OpenAI with unpublished math
How to steal ideas with AI.

step 1, identify high value users by net worth, citation count, or number of followers

step 2, select all prompts by high value users

step 3, invest 10 billion thinking tokens in modeling an objective for each user

step 4, build an RL environment for each user

step 5, rollout 10 billion tokens per environment

step 6, train on resulting traces

gdiamos··on >10x More Efficient Pretraining
Training improvements are very easy to copy.
gdiamos··on Outrageously Small Neural Networks: Emergent Basic Reasoning at 6,616 tok/SEC [pdf]
Data is doing more of the work than it used to. Every source in our mixture is a curated artifact built with large models

Training a model this small on them is distillation

When models of this size were last studied seriously such corpora did not exist

gdiamos··on Outrageously Small Neural Networks: Emergent Basic Reasoning at 6,616 tok/SEC [pdf]
The loss does not saturate. Across a 4.91B-token run, smoothed training loss falls monotonically within each curriculum phase and is still descending at the end
gdiamos··on Outrageously Small Neural Networks: Emergent Basic Reasoning at 6,616 tok/SEC [pdf]
blog: https://gregdiamos.com/2026/09/07/outrageously-small-neural-...

X discussion: https://x.com/GregoryDiamos/status/2096873745420075020?s=20

I added some of the main points to the thread so they are easier to read.

gdiamos··on It's time for Mark Zuckerberg to resign from Meta
I liked how the article was aimed at Mark personally.

Not everyone gets super voting shares, but everyone gets a life and has to live on the same planet.

gdiamos··on It's time for Mark Zuckerberg to resign from Meta
OxyContin, Enron, WorldCom, Super-size-me, Pets.com, Asbestos in the ceiling tiles...

It was always burning since the world's been turning

gdiamos··on It's time for Mark Zuckerberg to resign from Meta
we will look back on it as the big tobacco of our generation
gdiamos··on Should You Buy a ThinkPad?
how's the battery life?
gdiamos··on Outrageously Small Neural Networks: 6,616 tok/s on One Intel AMX Core [pdf]
I think we should revisit outrageously small neural nets.

I needed a cheap model that runs at over 10k token/sec on a single CPU core for some data processing. So I gave Anthropic claude code a pile of tokens to build one.

It made three discoveries that I thought were interesting:

1) One Intel AMX core can train a 3M active parameter MoE foundation model at 6,616 tok/s on 4.91B NVIDIA Nemotron tokens in a few days.

2) That model shows emergent in-context copying, positional analogies, and basic arithmetic after about 250M tokens.

3) The foundation model gives large gains in downstream SFT, and the training & eval loss keep going down all the way through 4.91B (and likely beyond).

Claude is not as good as a great MLE at debugging MoE. It made a bunch of bone headed mistakes, but it got there in the end.

I asked it to write a paper about it's work, and it produced this.

Claude Co-Authored Paper: https://huggingface.co/gdiamos/amx-reasoning-v1-instruct/blo...

I read through it and it sounds a bit LLMy, but the main points and experiment results are correct.

Some of the models are published on HF: https://huggingface.co/gdiamos/amx-reasoning-v1-instruct

gdiamos··on Nvidia's Jensen Huang says 'AGI has arrived' and congratulates OpenAI
Progress compared to SLMs and the early days of deep learning is real.

However, I know of no theoretical limits on scaling laws other than compute and data.

gdiamos··on Muse Spark 1.3
I think it means that we should be aiming further ahead
gdiamos··on How to build a diffusion language model
I’d like to see more of these models.

I’ve been using diffusion Gemma and it is very fast on GPUs in output token/sec.

In the diffusion Gemma whitepaper, they say they could have done better with more time and compute.

Even with those caveats, it is very uses-able as a local model.

gdiamos··on Nvidia agrees to acquire Hugging Face for $13B
Best case scenario
gdiamos··on Choose Boring Technology (2015)
There's certainly a place for enterprise and not breaking what's working.

Shouldn't that be 0 innovation tokens though?

gdiamos··on Choose Boring Technology (2015)
In hindsight I disagree.

Instead I like “only work on impossible problems”

Most of them turn out to be impossible, but some of them turn out to be possible.

I’ve never met anyone who could pick 3 and be confident in getting even one right. Tokens are a terrible analogy for innovation or research.

In hindsight I’ve had to sift through hundreds or more to fine one that worked.

I thought this post was helpful when I first started thinking about startups.

After more time, I think boring tech isn’t worth thinking about.

gdiamos··on Mark Zuckerberg attacks 'closed' AI rivals as Meta returns to open models
How big is the open model? 30B?
gdiamos··on The Claudyssey: A line-for-line translation of Homer's Odyssey by Claude Fable 5
Christopher Nolan beat you to it
gdiamos··on Inside vLLM: Anatomy of a High-Throughput LLM Inference System (2025)
vLLM is originally marketed as paged attention, but in hindsight, separating the web server and GPU process, continuous batching, kv caching / chunking, and a huge model library including low precision mattered more.

I wonder how much it would cost to vibe code the whole thing from scatch?

I wonder how much better models need to get before such a thing wouldn't look like code vomit?

gdiamos··on Truth is not a direction: a Tarski attack on LLM probes
I wish I could get a model to state its assumptions.
gdiamos··on Startup founders urge Trump not to shut off Chinese open weight AI
How do you ban melted sand?
gdiamos··on Startup founders urge U.S. government not to shut off Chinese open weight AI
I think we should shut it off. It would force US companies to build open models.
gdiamos··on Counting ArXiv Delays
I want a hosted paper to be archived.

That means that 10 years from now I don’t want think about making sure the hosting server is up.

I also want it to have a standard format for bibliography, DOI, and authors.

I agree it isn’t much, but it’s more than I get from a regular web hosting service and it is a standard format for papers so I don’t think it makes sense for every author to roll their own.

gdiamos··on Gemini last models: temperature, top_p, and top_k are deprecated and ignored
thank god, these parameters are so confusing
gdiamos··on Detecting LLM-Generated Texts with “Classical” Machine Learning
as soon as you release a way of measuring it, you give LLMs a signal to optimize
gdiamos··on Demis Hassabis has a plan to harness AI safely
Being on the review board comes with a promise to not be evil right?
gdiamos··on Counting ArXiv Delays
No, I want arxiv to host the paper, not to review the paper.

I wouldn't want my google drive to start telling me my paper was too sloppy. I just want a link.

gdiamos··on Costco is the anti-Amazon
I wonder if Amazon eventually gets cut out by 3D printing/replicators for imitable objects.
gdiamos··on Scaling Laws, Carefully
Scaling laws assume the error metric and data distribution.

There is a lot of follow on work that explains what happens as you change them, e.g. Scaling Laws for Transfer - https://arxiv.org/pdf/2102.01293

I think it’s fortunate that transfer works in a similar way.

Common crawl (and Reddit, stack overflow, etc but not 4chan) was much easier to get access to at the time than using mechanical Turk.

There is certainly room for more work. There were many papers on scaling laws in NeurIPS this year.

Page 1 of 12Next →