HNHacker News
TopNewBestAskShowJobs

gyrovagueGeist

186 karma · joined April 25, 2021

submissionscomments
gyrovagueGeist··on Qwen3.6-35B-A3B: Agentic coding power, now open to all
Electricity (on continental US) is pretty cheap assuming you already have the hardware:

Running at a full load of 1000W for every second of the year, for a model that produces 100 tps at 16 cents per kWh, is $1200 USD.

The same amount of tokens would cost at least $3,150 USD on current Claude Haiku 3.5 pricing.

gyrovagueGeist··on Mamba-3
Not wrong, but I think it's more accurate to say:

Mamba is an architecture for the middle layers of the network (the trunk) which assumes decoding takes place through an autoregressive sequence (popping out tokens in order). This is the SSM they talk about.

Diffusion is an alternative to the autoregressive approach where decoding takes place through iterative refinement on a batch of tokens (instead of one at a time processing and locking each one in only looking forward). This can require different architectures for the trunk, the output heads, and modifications to the objective to make the whole thing trainable. Could mamba like ideas be useful in diffusion networks...maybe but it's a different problem setup.

gyrovagueGeist··on TorchLean: Formalizing Neural Networks in Lean
This is just standard Fourier theory of being able to apply dense global convolutions with pointwise operations in frequency space? There’s no mystery here. It’s no different than a more general learnable parameterization of “Efficient Channel Attention (ECA)”
gyrovagueGeist··on Bye Bye Humanity: The Potential AMOC Collapse
They also didn't have nuclear weapons to use in global conflicts over resources.
gyrovagueGeist··on Scaffolding to Superhuman: How Curriculum Learning Solved 2048 and Tetris
I've always found curriculum learning incredibly hard to tune and calibrate reliably (even more so than many other RL approaches!).

Reward scales and horizon lengths may vary across tasks with different difficulty, effectively exploring policy space (keeping multimodal strategy distributions for exploration before overfitting on small problems), and catastrophic forgetting when mixing curriculum levels or when introducing them too late.

Does any reader/or the author have good heuristics for these? Or is it still so problem dependent that hyper parameter search for finding something that works in spite of these challenges is still the go to?

gyrovagueGeist··on A visualization of the RGB space covered by named colors
Thief of Time by Terry Prachett has a great minor bit about characters who are naming themselves after colors running out of human made labels, as they have to get increasingly esoteric with the names. It's fun to see that visualized.
gyrovagueGeist··on Were RNNs all we needed? A GPU programming perspective
It's from the depth of the computation, not the work
gyrovagueGeist··on Development speed is not a bottleneck
In the middle term, I almost feel less productive using modern GPT-5/Claude Sonnet 4 for software dev than prior models, precisely because they are more hands off and less supervised.

Because they generate so much code, that often passes initial tests, looks reasonable, and fails in nonhuman ways, in a pretty opinionated style tbh.

I have less context (and need to spend much more effort and supervision time to get up to speed to learn) to fix, refactor, and integrate the solutions, than if I was only trusting short few line windows at a time.

gyrovagueGeist··on Gemini with Deep Think achieves gold-medal standard at the IMO
Useful and interesting but likely still dangerous in production without connecting to formal verification tools.

I know o3 is far from state of the art these days but it's great at finding relevant literature and suggesting inequalities to consider but in actual proofs it can produce convincing looking statements that are false if you follow the details, or even just the algebra, carefully. Subtle errors like these might become harder to detect as the models get better.

gyrovagueGeist··on Multiplatform Matrix Multiplication Kernels
For people who are interested Kokkos (a C++ library for writing portable kernels) also has a naming scheme for hierarchical parallelism. They use ThreadTeam, Thread (for individual threads within a group), and ThreadVector (for per thread SIMD).

Just commenting to share, personally I have no naming preference but the hierarchal abstractions in general are incredibly useful.

gyrovagueGeist··on uv: An extremely fast Python package and project manager, written in Rust
I can't reproduce this, possibly it's another environment error or the problem has been fixed in the versions I am using. (uv 0.7.11, zsh 5.9, python 3.13.5, MacOS 15.3.1)
gyrovagueGeist··on uv: An extremely fast Python package and project manager, written in Rust
It does? What problems have you had?
gyrovagueGeist··on Peano arithmetic is enough, because Peano arithmetic encodes computation
Yep! Optimize (solve in infinite dimensions) and then discretize onto a finite basis has typically led to much better and stable methods than a discretize and then optimize approach.

Time-scale calculus is a pretty niche theoretical field that looks at blending the analysis of difference and differential equations, but I'm not aware of any algorithmic advances based on it.

gyrovagueGeist··on Magistral — the first reasoning model by Mistral AI
Does anyone know why they added minibatch advantage normalization (or when it can be useful)?

The paper they cite "What matters in on-policy RL" claims it does not lead to much difference on their suite of test problems, and (mean-of-minibatch)-normalization doesn't seem theoretically motivated for convergence to the optimal policy?

gyrovagueGeist··on Neuromorphic computing
I am not sure why HN has mostly LANL posts. Otherwise though it is a combination of things. Machine learning applications for NATSec & fundamental research have become more important (see FASST, proposed last year), the current political environment makes AI funding and applications more secure and easier to chase, and some of this is work that has already been going on but getting greater publicity for both of those reasons.
gyrovagueGeist··on Atuin Desktop: Runbooks That Run
All the problems of reproducibility in Python notebooks (https://arxiv.org/abs/2308.07333, https://leomurta.github.io/papers/pimentel2019a.pdf) with the power of a terminal.
gyrovagueGeist··on LLM-powered tools amplify developer capabilities rather than replacing them
How many people still play centaur chess?
gyrovagueGeist··on The Curious Similarity Between LLMs and Quantum Mechanics
ah, yes, the spooky similarities of hilbert spaces and probability theory
gyrovagueGeist··on The Google Willow Thing
They likely mean on any of the current era of NISQ-like devices (https://en.wikipedia.org/wiki/Noisy_intermediate-scale_quant...) like this one or quantum annealers.
gyrovagueGeist··on Google says AI weather model masters 15-day forecast
Yep, this is also the problem that self-driving cars have with the "accidents per mile" metric.
gyrovagueGeist··on Researchers use AI to turn sound recordings into street images
I love this paper, but something I think is often missed when it comes up is that you CAN hear the shape of many drums if you restrict the shape space, for example with a prior of "what a drum should look like" Zelditch proved spectral uniqueness for convex, fully connected, drums with some symmetry.
gyrovagueGeist··on SVDQuant: 4-Bit Quantization Powers 12B Flux on a 16GB 4090 GPU with 3x Speedup
This problem seems like it would be very similar to the Low-Rank + Sparse decompositions that used to be popular in audio-visual filtering.
gyrovagueGeist··on 500 Python Interpreters
These were two separate and mostly unrelated PEPs but otherwise that's correct.
gyrovagueGeist··on Spice: Fine-grained parallelism with sub-nanosecond overhead in Zig
This is neat and links to some great papers. I wish the comparison was with OpenMP tasks though; I’ve heard Rayon has a reputation for being a bit slow
gyrovagueGeist··on Fast Multidimensional Matrix Multiplication on CPU from Scratch (2022)
Typically no, although BLAS software engineers occasionally write HPC strassen's implementations and papers about them. In my opinion there's a few reasons why they're not common in BLAS libraries:

- They're a bit less numerically stable, which used to matter more for common BLAS use cases. Less so these days.

- The memory access patterns and algorithm parallelism make it much harder to reach as high a fraction of peak performance as standard GEMM. Matrix size restrictions for recursive mult algorithms is also an issue.

gyrovagueGeist··on Do not try to be the smartest in the room; try to be the kindest
- "In this world, Elwood, you must be oh so smart or oh so pleasant." Well, for years I was smart. I recommend pleasant"
gyrovagueGeist··on Ask HN: I have many PDFs – what is the best local way to leverage AI for search?
+1 for this. I use rga all the time. it's a "simple" solution but often enough for what I actually needed.
gyrovagueGeist··on OpenAI pulls Johansson soundalike Sky’s voice from ChatGPT
They genuinely do not sound that similar to me
gyrovagueGeist··on Why Fugaku, Japan's fastest supercomputer, went virtual on AWS
The interconnect and network topology is also a big component of the hardware where you can't "fake The Real Thing" in practice. You can often get fairly confident in program correctness for toy problem runs by scaling 1-~40 ranks on your local machine, but you can't tell much about the performance until you start running on a real distributed system where you can see how much your communication pattern stresses the cluster.

Or if you run into bugs / crashes that needs 1000s of processes or a full scale problem instance to reproduce, god help you and your SLURM queue times.

gyrovagueGeist··on A canonical Hamiltonian formulation of the Navier–Stokes problem
It's been commonly used as a noun in the HPC / scientific computing space for a long time (relevant to this thread) typically talking about "compute bound" and "memory bound" algorithms. It's spilled over everywhere now that number crunching is hip for machine learning.
Page 1 of 4Next →