HNHacker News
TopNewBestAskShowJobs

stephantul

711 karma · joined September 3, 2024

NLP engineer
submissionscomments
stephantul··on Ask HN: Do you worry about fires in AI datacenters?
No
stephantul··on Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
The (relevant) segue into tires elevates this post so much. I don’t understand why, but it does
stephantul··on Pandas Should Go Extinct
Agreed on all counts.

In many cases I’ve found directly using python primitives to be less confusing than pandas.

Similarly, in companies I’ve worked at, the datasets just aren’t that big. Especially if you’ve got access to modern hardware.

stephantul··on The Last 24 Hours Are the Opening Scene in a Horror Movie
If you find a place where I can host a trillion parameter model without anyone finding out about it, let me know.
stephantul··on The Last 24 Hours Are the Opening Scene in a Horror Movie
But how. Models don’t have access to their own weights.
stephantul··on The Last 24 Hours Are the Opening Scene in a Horror Movie
Not the code: the weights. Are you going to host a trillion parameter model somewhere without someone noticing?
stephantul··on The Last 24 Hours Are the Opening Scene in a Horror Movie
Ok but do any of these data centers have a copy of the models that attacked hf?
stephantul··on The Last 24 Hours Are the Opening Scene in a Horror Movie
One thing I definitely do not understand about this discourse is that the models that are good enough to self-replicate can’t survive on normal machines, e.g., the models can’t hide on some random server.

So, if it is as dangerous as they say it is: there is a single physical source of this danger, which is OpenAI/Anthropic servers. If it is this dangerous, they can just turn it off.

Instead, we just keep pretending that the models that attacked HF were hosted or replicating on HF hardware. Not the case! They infiltrated it, but were hosted elsewhere.

stephantul··on RTK reports token savings, but our cost benchmarks disagree
I think anyone who is even a little bit realistic knows that most technologies overclaim, or evaluate under very favorable conditions.

This is not a good thing of course, but I also feel that acting surprised that this is going on is a little unnecessary.

Having said that: most tools are not helpful

stephantul··on Harnessing the Universal Geometry of Embeddings
I’ve never liked that this was called “the platonic representation hypothesis”. Lots of weird baggage attached and seems like a waste of a good name.
stephantul··on Benchmarking Vector Indexes
This is a super long ai generated slab of text with very little actual info. It doesn’t include any analysis of results.

It does include the phrase “worth listing”.

stephantul··on METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
Super charitable reading imo. This is like saying we can’t detect a speeding car because we can’t run as fast as a fast car.

It’s not like the humans were engaged in some kind of battle of wits with some super AI, it’s just some employee not monitoring the output of an experiment.

stephantul··on You Know GDPR Is Good Based on Who Hates It
Cookie banners are made annoying on purpose. This has nothing to do with GDPR itself.

The entities forced to show them would rather not, and thus make it as annoying as possible for you. They then use this to weaken support for the GDPR.

Shame on the people making stuff like this.

stephantul··on Andreessen Horowitz is investing billions into a bleak future
Tragedy of the commons. The person that ruins a commons first takes all the supply.

This is why regulation before, not after, the commons are plundered, is important.

stephantul··on The Benchmarkpocalypse
Unfortunately, even a holdout set doesn’t protect you from overfitting, it just takes longer.

Of course having a holdout set is better than not having one. It’s just not a silver bullet.

stephantul··on How I over-engineered my book
It is perhaps ironic that I find this post very difficult to read. I'm super interested in the content, but it reads like it is generated.
stephantul··on How to disable or avoid intrusive AI
My heart bleeds for Rovo’s product managers. They must also see that nobody actually wants or needs or even likes Rovo.
stephantul··on Manus will return to operating as an independent company
Anyone have info on whose regulatory restrictions they are referring to?
stephantul··on Show HN: Mcptoon – Token-efficient MCP CLI client
I think that some of these choices (as others have commented) show that the author has not investigated how tokenization works.

Tokenization is not some black box, you can run tokenizers and check them.

stephantul··on Honey, I shrunk the embeddings: Matryoshka vs. PCA
PCA is applied after the model, so there should be no difference in embedding throughput. Lookups in the index should be faster, but that speedup also applies equally to MRL.

So I guess the answer is: no

stephantul··on Honey, I shrunk the embeddings: Matryoshka vs. PCA
Ah I meant more to say that I was working on this as well. I haven’t published the results for this comparison specifically yet.
stephantul··on Honey, I shrunk the embeddings: Matryoshka vs. PCA
Nice! I’ve been working on something similar and found similar results.

In my experiments, I used lots of embedding models and the results were not nearly as uniform as this curve, just FYI. I didn’t use any of the API-based models though

I also wrote about this exact comparison when using PCA and MRL to quantize static models, see: https://stephantul.github.io/blog/mrl-pca/

stephantul··on Show HN: DeepSeek-V4 Latent Reasoning – moving "thinking" into latent space
LinkedIn recently added a “seems like AI slop” button. I.e.: independent of downvoting/not interested/flagging as spam/ToS violation, you can say “this is AI slop”. Maybe we need something like it here
stephantul··on Discovery Loop
That is true, I’ve seen people do biochemistry and geology work, and it did look very mind-numbing.

Then again, gassing rats and taking biopsies is not something you can do with AI.

stephantul··on Discovery Loop
I’ve always felt that the idea that science is bottlenecked and therefore needs more automation only works for a very narrow definition of what science is, and entails a very specific view on what it should be.
stephantul··on Show HN: Simple algorithm and color space to generate diverse skin tones
Love it. I especially thought the introspective “aside” section on related resources was great, and I wish more people would show this kind of reflection.
stephantul··on The Religion of Speed
I’m in a similar boat. The one thing that seems to help is recognizing when the ask from leadership is genuine, and also sensible from a product point of view. Seize that moment to go fast, and give that your full attention.

The other project will either peter out, because they made no sense from a product point of view or because leadership lost interest. If they’re simple, can likely be done quickly with the help of AI.

It’s not a pretty answer, but AI has helped me cope with this kind of situation much better than in the past.

stephantul··on Show HN: A 6M-token movable window on a single 46GB GPU
100% generated. I skimmed the paper, and came out with a feeling of still not knowing what this is about.
stephantul··on Show HN: OneCLI – OSS credential gateway that keeps secrets out of AI agents
It is hilarious to see most comments are people peddling their own products more or less directly.
stephantul··on NPM's release cooldown is security theater
They sell their products using the credentials they gained.

I’d never heard of socket until they found and reported shai hulud hiding in pytorch lightning. It pays off.

Page 1 of 7Next →