HNHacker News
TopNewBestAskShowJobs

lostmsu

6,610 karma · joined January 17, 2014

Threw hundreds of baby transformers into water because their bits-per-byte were too high.

Some fun stuff:

https://borgcloud.org/speech-to-text at $0.06/h

Roxy: iOS hands-free voice AI: https://itunes.apple.com/app/id6737482921?mt=8

Turing Test Battle Royale: https://trashtalk.borg.games

meet.hn/city/43.6534817,-79.3839347/Toronto

Socials: - linkedin.com/in/victor-msu - reddit.com/user/lostmsu - github.com/lostmsu

Interests: AI/ML, Gaming, Networking, Programming, Research, Science, Startups, Technology

---

submissionscomments
lostmsu··on Show HN: TurboGPT: train 22KiB transformer in 13s
Just uploaded to https://huggingface.co/datasets/lostmsu/hn1g/tree/main

But it is a byte predictor. You can train it on any file.

lostmsu··on Show HN: TurboGPT: train 22KiB transformer in 13s
> So, why do people keep making these? Asking genuinely

I extensively used minGPT for home experiments on transformer architecture. It is great for learning!

However, if you want to scale the experiments up at home you need to go faster. Karpathy made optimized https://github.com/karpathy/nanoGPT, but it is tuned for "8XA100 40GB node in about 4 days of training".

13s is a bit overkill here (my machine builds that project in 30s). But it gives some space for experimentation with architectures that don't have optimized primitives.

lostmsu··on Windows 11½
OMG that Windows Store opens instantly
lostmsu··on The Rise of Audio AR (2024)
I would love to just pay for a device like that. In fact, I am looking for a device like that. But man the "file an application and we will consider" reminds me of loot boxes in video games.
lostmsu··on SNL Weekend Update: Anthropic CEO Dario Amodei on A.I.'S Threat to Humanity [video]
Video unavailable

The uploader has not made this video available in your country

lostmsu··on Unreal Agent
Don't give the agent any kind of wait or sleep tools and deny any attempt to read status without new progress notifications.
lostmsu··on A single function Jev-like wrapper for LLMs, including vision models
Much less code https://news.ycombinator.com/item?id=49769618
lostmsu··on Tech Needs Humanists More
It fits in a paragraph, and we all read it. The rest of it is bullshit.
lostmsu··on Tech Needs Humanists More
> you wouldn't need to spend years in school

Do you?

lostmsu··on Mercury 2.5 LLM hits 770 tokens per second
Inco sucks. I tried their GLM 5.3 Flash and it was quantized to the point of hallucinating Chinese in the middle of English only agentic sessions. Never happened with any other provider.
lostmsu··on Unreal Agent
This problem has a few rather trivial solutions.
lostmsu··on Looking forward to Git 2.56 – and 3.0
How is this different from branches?
lostmsu··on Wall Street Is Growing Skeptical of the Data Center Boom
What's the output rate and config?
lostmsu··on Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM
It doesn't show any indications of solving catastrophic forgetting.
lostmsu··on Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM
This is slop. 8M parameter dense model with context length 64 that you train on enwik9 in 2h will have 1.15 bpb. This model has 1.8 (bits per byte, lower is better).
lostmsu··on Introducing System One Models and Jev
According to https://goodstartlabs.com/research/verification-is-the-bottl... it is only 2.6x cheaper than DeepSeek V4.1 Flash, and they did not test v4.0 Flash, which would have been same price.
lostmsu··on Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash
> Monarch Hadamard MLP: replaces the dense FFN with three learnable Walsh-Hadamard-initialized Kronecker (Monarch) factor pairs interleaved with per-channel diagonal scales, fixed permutations, a SiLU nonlinearity, and a rank-8 input-conditioned gate, so each token gets a fully mixed nonlinear transform of its d_model channels at O(d√d) parameters and compute instead of the O(d²) a dense 4x-expansion MLP would cost.

Wow, I was just researching W-H in transformers. Did yours seem to work? In my experiments swapping various components for W-H-like transforms caused extreme quality degradation.

UPD. according to the comments here, this model simply does not work at all, so I guess the answer is NO

lostmsu··on A heap overflow and SSO misconfiguration to compromise OpenAI internal repos
MS absolutely has a couple of stock-based incentives.
lostmsu··on Apple M6 Pro Achieves the Highest Single-Core CPU Score in Geekbench 7
> And there are bugs in any Linux distribution that are way too complex to fix even if theoretically possible. So it doesn’t make any difference.

It does now with LLMs

lostmsu··on Microsoft exec called AI scraping 'the largest theft of labor in human history'
Thank God there's the Internet Archive then. https://archive.org/donate I do.
lostmsu··on Microsoft exec called AI scraping 'the largest theft of labor in human history'
Good luck with that.
lostmsu··on Xiaomi Mimo 2.6 live post-training dashboard
That's posttraining. Pretraining is the expensive part.
lostmsu··on Why I'm still bearish on LLMs after Navier-Stokes
> just a list of 64 numbers

> remember even a few positions? Sure they could

A rough estimate of number of positions across all X move games is X^10. For 15 moves it is hopeless to remember even a relatively small part of them. Typical game has 40 turns, 1 move per player, so 80 moves.

lostmsu··on Canada welcomes EU proposal to become 'associate member'
> It's just political theater

Do you think the referendum in question is not a political theater?

I don't know much, but it stands to me that the EU association is magnitudes more serious.

lostmsu··on Why I'm still bearish on LLMs after Navier-Stokes
> prediction which is closer to memorization

> don't memorize inputs - they predict them

I feel some tension here.

> rice grains on a chess board? Sure, but this has nothing to do with chess, and nothing to do with how many games were in the LLM's training data.

> just a list of 64 numbers

> remember even a few positions? Sure they could, but that's irrelevant.

I don't think you do. Or rather you do know the legend but for some funny reason seem to be unable to apply its lesson here, because you are talking about enormous terabytes of training data.

> Intelligent humans created the training data, and the LLM attempts to predict (copy) the training data, so of course it looks intelligent.

If for you it is about intelligence, I am out of this discussion.

lostmsu··on Why I'm still bearish on LLMs after Navier-Stokes
They can't possibly remember even a few positions. Don't you know the legend about rice grains on a chess board?

The claim here is not about intelligence, it is about generality. There's no doubt for me the LLMs are intelligent.

lostmsu··on OpenAI expands ChatGPT ads with Sponsored Agents
I am surprised nobody brought up Accelerando yet.
lostmsu··on Stay discoverable in search while disallowing AI training
Papers charged per "user" since times immemorial.
lostmsu··on Why I'm still bearish on LLMs after Navier-Stokes
You are saying "No it is not" without an argument. The fact that computer systems could play chess yet not being AGI has no relevance to LLMs' ability to play chess being AGI, because the point is about G, not I. There's little doubt about A or I parts.
lostmsu··on Why I'm still bearish on LLMs after Navier-Stokes
Do they?
Page 1 of 34Next →