HNHacker News
TopNewBestAskShowJobs

Scene_Cast2

2,808 karma · joined June 13, 2012

Lost my PW for my previous account, Scene_Cast Email: scene_cast2@outlook.com
submissionscomments
Scene_Cast2··on Open models by OpenAI
Should be in the RoPE paper. The OG transformers used multiplicative sinusoidal embeddings, while RoPE does a pairwise rotation.

There's also NoPE, I think SmolLM3 "uses NoPE" (aka doesn't use any positional stuff) every fourth layer.

Scene_Cast2··on A Photonic SRAM with Embedded XOR Logic for Ultra-Fast In-Memory Computing
Something I've never quite understood is where, on the spectrum of mainstream vs niche, in memory computing approaches lie. What are the proposed use cases?

I understand that you can get highly power efficient XORs, for example. But if we go down this path, would they help with a matrix multiply? Or the bias term of a FFN? Would there be any improvement (i.e. is there anything to offload) in regular business logic? Should I think of it as a more efficient but highly limited DSP? Or a fixed function accelerator replacement (e.g. "we want to encrypt this segment of memory")

Scene_Cast2··on Show HN: My GPU Fan Saga – A DIY ATX Fan Controller
The Windows FanControl software uses LibreHardwareMonitor as the sensor backend. It works pretty well in my experience.
Scene_Cast2··on Show HN: My GPU Fan Saga – A DIY ATX Fan Controller
I have a pretty custom GPU cooling setup on a few machines (I run ML workloads locally and I want stuff to be quiet).

Couple of gotchas that I ran across. I found that on Linux, desktop PC fan control support is pretty abysmal. The sensor library that everyone relies on, lm_sensors, is semi-abandoned and didn't recognize sensors on my relatively popular, 7 year old ATX motherboard and GPU. It also requires having Perl installed.

About GPU cooling in particular - modern NVidia cards in particular seem to have a built-in minimum of 30% fan speed when controlling them manually. The connectors are also a different, smaller connector (perhaps a JST PH?).

Scene_Cast2··on Perfume reviews
Oh, perfumes are a great hobby. If you're in SF or LA, definitely hit up one of the boutique perfume shops (Scent Bar and Ministry of Scent).

There are also bunch of sellers who package samples (aka "decants" - buy a 100mL, split it into smaller bottles). I found that 1-2mL is plenty to get an idea. I've had great experience with LuckyScent (mentioned in the article), Surrender to Chance, as well as random reddit swaps and highly rated Ebay sellers.

The perfume scene is super wide and diverse, and I found that although there are general trends, it's hard to even know all the popular brands, and everyone's nose is unique. Skip stuff like Aventus and Sauvage and buy some discovery sets (surrender to chance puts together some good ones).

There is definitely a spectrum between "wearable crowd-pleaser" and "avant-garde storytelling" - Afrika-Olifan comes to mind - love it for the creativity and execution, but it would be rude to go outside wearing it. There's also some storytelling - Black March, for example, starts off with grassy fresh earth after a rain, then turns into flowers.

Scene_Cast2··on PyPI Prohibits inbox.ru email domain registrations
Oh hey, I was the person who reported this.
Scene_Cast2··on NeuralOS: An operating system powered by neural networks
This is a really cool idea! I wonder if a more integrated adaptation would help render real OS UIs better / faster / prettier.
Scene_Cast2··on Show HN: Cactus – Ollama for Smartphones
Does this support openrouter?
Scene_Cast2··on New sphere-packing record stems from an unexpected source
What ended up launching is a fancy product quantization based on k-means. Some of the tricks were storing magnitude separately (i.e. removing the mean) and rearranging dimensions based on variance (and/or rotation based on PCA) for PQ to work better.

I also remember trying to fit a distribution so that I can generate synthetic data (not for a lack of data, but more for understanding the problem space better). The synthetic data quantized pretty differently - my guess is that it's because of random areas of density and sparsity.

I'm not quite following your exact rearranging idea though. Not sure if the above answers the question.

Scene_Cast2··on New sphere-packing record stems from an unexpected source
They weren't sparse, they were dense but the "density" was quite non-uniform (think typical learned ML vectors). Not too far from an N-dimensional gaussian (I ended up reading research on quantizing Gaussian distributions, but that didn't help either as we didn't have a perfectly gaussian thing).
Scene_Cast2··on New sphere-packing record stems from an unexpected source
Neat. I spent a month trying to use sphere packing approaches for a better compression algorithm (I had a large amount of vectors, they were grouped through clustering). Turned out that theoretical approaches only really work for uniform data and not any sort of real-world data.

EDIT: groped -> grouped

Scene_Cast2··on Functions Are Vectors (2023)
Same thing in video form explained by a different person - https://youtu.be/mhEFJr5qvLo
Scene_Cast2··on Alternative Layout System
I still can't read it despite trying.
Scene_Cast2··on The bitter lesson is coming for tokenization
I agree that you can encode any single concept and that the encoding space of a single top pick grows exponentially.

However, I'm talking about the probability distribution of tokens.

Scene_Cast2··on The bitter lesson is coming for tokenization
I realized that with tokenization, there's a theoretical bottleneck when predicting the next token.

Let's say that we have 15k unique tokens (going by modern open models). Let's also say that we have an embedding dimensionality of 1k. This implies that we have a maximum 1k degrees of freedom (or rank) on our output. The model is able to pick any single of the 15k tokens as the top token, but the expressivity of the _probability distribution_ is inherently limited to 1k unique linear components.

Scene_Cast2··on uv: An extremely fast Python package and project manager, written in Rust
10x is too precise.
Scene_Cast2··on Python can run Mojo now
The thing I focus on when writing compiled extensions for Python isn't the speed of the extension, but rather the overhead of the call and the overhead of moving objects from Python -> compiled and compiled -> Python.

Is there a zero-copy interface for larger objects? How do object lifetimes work in that case? Especially if this is to be used for ML, you need to haul over huge matrices. And the GIL stuff is also a thing.

I wonder how Mojo handles all that.

Scene_Cast2··on Figma Slides Is a Beautiful Disaster
I've been in this situation. I'll spend the hour watching the info, but I'll dislike the inefficiency. I consider it impolite.
Scene_Cast2··on Show HN: A Implementation of Alpha Zero for Chess in MLX
This is one of those topics that LLMs (Opus 4, Gemini 2.5 pro, etc) seem bad at explaining.

I was trying to figure out the difference between the Stockfish approach (minimax, alpha-beta pruning) versus Alpha Zero / Leela Chess Zero (MCTS). My very crude understanding is that stockfish has a very light & fast neural net and goes for a very thorough search. Meanwhile, in MCTS (which I don't really understand at this point), you eval the neural net, sample some paths based on the neural net (similar to minimax), and then pick the path you sampled the most. There's also the training vs eval aspect to it. Would love a better explanation.

Scene_Cast2··on Precision Clock Mk IV
That reminds me of production studio master clocks such as the bronics WCD-530W or the evertz equivalent. I'm more of a fan of the analog style such as the 1275T - https://evertz.com/products/12x5T.
Scene_Cast2··on Attention Wasn't All We Needed
RE: optimizer performance - any thoughts on heavyball?
Scene_Cast2··on Claude 4
Already up on openrouter. Opus 4 is giving 429 errors though.
Scene_Cast2··on That fractal that's been up on my wall for years
I wonder if something similar can be applied to get a dither pattern with built-in level of detail adjustment.
Scene_Cast2··on GitHub Copilot Coding Agent
Nope - I use a-la-carte pricing (through openrouter). I much prefer it over a subscription, as there are zero limits, I pay only for what I use, and there is much less of a walled garden (I can easily switch between Anthropic, Google, etc).
Scene_Cast2··on GitHub Copilot Coding Agent
Cline very visibly displays the ongoing cost of the task. Light edits are about 10 cents, and heavy stuff can run a couple of bucks. It's just that the tab accumulates faster than I expect.
Scene_Cast2··on GitHub Copilot Coding Agent
I tried doing some vibe coding on a greenfield project (using gemini 2.5 pro + cline). On one hand - super impressive, a major productivity booster (even compared to using a non-integrated LLM chat interface).

I noticed that LLMs need a very heavy hand in guiding the architecture, otherwise they'll add architectural tech debt. One easy example is that I noticed them breaking abstractions (putting things where they don't belong). Unfortunately, there's not that much self-retrospection on these aspects if you ask about the quality of the code or if there are any better ways of doing it. Of course, if you pick up that something is in the wrong spot and prompt better, they'll pick up on it immediately.

I also ended up blowing through $15 of LLM tokens in a single evening. (Previously, as a heavy LLM user including coding tasks, I was averaging maybe $20 a month.)

Scene_Cast2··on X X^t can be faster
I can't name any applications off the top of my head, other than iterative matrix multiplication for approximate eigenvector finding in square matrixes. But I don't know what's actually used for finding eigenvectors (or other decompositions for that matter).
Scene_Cast2··on It Awaits Your Experiments
Blindsight is known to be a slog for a lot of people including myself.

I love sci-fi, I love challenging ideas, and I really liked the concepts explored in Blindsight - except that I learned those concepts through summaries and selective reading.

Scene_Cast2··on The world could run on older hardware if software optimization was a priority
Where lack of performance costs money, optimization is quite invested in. See PyTorch (Inductor CUDA graphs), Triton, FlashAttention, Jax, etc.
Scene_Cast2··on Multiple security issues in GNU Screen
I use byobu (for the keybinds) on top of tmux. But Zellij (modern Rust-based alternative to tmux) has been looking quite interesting for a while.
← PreviousPage 6 of 29Next →