HNHacker News
TopNewBestAskShowJobs

numeri

982 karma · joined January 27, 2022

submissionscomments
numeri··on GPT-5-Codex is a better AI researcher than me
This is written by someone who's not an AI researcher, working with tiny models on toy datasets. It's at the level of a motivated undergraduate student in their first NLP course, but not much more.
numeri··on How to be a leader when the vibes are off
One sign would be occasionally changing course in response to overwhelming employee feedback. If that never or almost never happens, the feedback is being ignored, not taken constructively and not followed.
numeri··on Why language models hallucinate
This isn't right – calibration (informally, the degree to which certainty in the model's logits correlates with its chance of getting an answer correct) is well studied in LLMs of all sizes. LLMs are not (generally) well calibrated.
numeri··on Grok: Searching X for "From:Elonmusk (Israel or Palestine or Hamas or Gaza)"
I really like your posts, and they're generally very clearly written. Maybe this one's just the odd duck out, as it's hard for me to find what you actually meant (as clarified in your comment here) in this paragraph:

> This suggests that Grok may have a weird sense of identity—if asked for its own opinions it turns to search to find previous indications of opinions expressed by itself or by its ultimate owner. I think there is a good chance this behavior is unintended!

I'd say it's far more likely that:

1. Elon ordered his research scientists to "fix it" – make it agree with him

2. They did RL (probably just basic tool use training) to encourage checking for Elon's opinions

3. They did not update the UI (for whatever reason – most likely just because research scientists aren't responsible for front-end, so they forgot)

4. Elon is likely now upset that this is shown so obviously

The key difference is that I think it's incredibly unlikely that this is emergent behavior due to an "sense of identity", as opposed to direct efforts of the xAI research team. It's likely also a case of https://en.wiktionary.org/wiki/anticipatory_obedience.

numeri··on Grok: Searching X for "From:Elonmusk (Israel or Palestine or Hamas or Gaza)"
I'm a little shocked at Simon's conclusion here. We have a man who bought an social media website so he could control what's said, and founded an AI lab so he could get a bot that agrees with him, and who has publicly threatened said AI with being replaced if it doesn't change its political views/agree with him.

His company has also been caught adding specific instructions in this vein to its prompt.

And now it's searching for his tweets to guide its answers on political questions, and Simon somehow thinks it could be unintended, emergent behavior? Even if it were, calling this unintended would be completely ignoring higher order system dynamics (a behavior is still intended if models are rejected until one is found that implements the behavior) and the possibility of reinforcement learning to add this behavior.

numeri··on I do not remember my life and it's fine
That's a bold claim! Actually, there are plenty of scientific experiments that show actual differences between people who report aphantasia and those who don't, including different stress responses to frightening non-visual descriptions, different susceptibility to something called image priming, lower "cortical excitability in the primary visual cortex", and more: https://en.wikipedia.org/wiki/Aphantasia

So we know that at least the people who claim to see nothing act differently. Could it just be that people who act differently describe the sensation differently, you might ask?

No, because there are actual cases of acquired aphantasia after neurological damage. These people used to belong to the group that claimed to be able to imagine visual images, got sick, then sought medical help when they could no longer visualize. For me, at least, that's pretty cut and dry evidence that it's not just differing descriptions of the same (or similar) sensations.

numeri··on I do not remember my life and it's fine
That's the thing, some people do see things in their mind that clearly. It's about as rare as full aphantasia, but it's absolutely a spectrum.
numeri··on I do not remember my life and it's fine
I think you're assuming more people are like you than actually are.

This is part of the classic debate around aphantasia – both sides assume the other side is speaking more metaphorically, while they're speaking literally. E.g., "Surely he doesn't mean he literally can't visualize things, he just means it's not as sharp for him." or "Surely they don't literally mean they can see it, they're just imagining the list of details/attributes and pretending to see it."

numeri··on I do not remember my life and it's fine
They're definitely quite hard for me. I bet my colleagues, friends or family could answer them for me better than I can without prep (which would involve chatting with my wife). Many of the experiences in this article resonate with me, but it's definitely not quite as extreme.
numeri··on Claude Code: An Agentic cleanroom analysis
Is the analysis right, or did the LLM hallucinate this?
numeri··on Claude Code: An Agentic cleanroom analysis
Yes, so that one can use it for more creative writing exercises. It was pretty creative, I'll give it that.
numeri··on Claude Code: An Agentic cleanroom analysis
No, it's completely useless, and puts the entire rest of the analysis in a bad light.

LLMs have next to no understanding of their own internal processes. There's a significant amount of research that demonstrates this. All explanations of an internal thought process in an LLM are completely reverse engineered to fit the final answer (interestingly, humans are also prone to this – seen especially in split brain experiments).

In addition, the degree to which the author must have prompted the LLM to get it to anthropomorphize this hard makes the rest of the project suspect. How many of the results are repeated human prompting until the author liked the results, and how many come from actual LLM intelligence/analysis skill?

numeri··on LLMs can see and hear without any training
It makes quite a lot of sense juxtaposed with "train time compute". The point being made is that a set budget can be split between paying for more training or more inference _at test time_ or rather _at the time of testing_ the model. The word "time" in "inference time" plays a slightly different role grammatically (noun, not part of an adverbial phrase), but comes out to mean the same thing.
numeri··on Overengineered Anchor Links
You seem to be responding to what you think I'm saying, not what I'm saying. As far as I know, "killing children" is not a dog-whistle. No one uses the words "killing children" to e.g., secretly express support for the Holocaust.
numeri··on Overengineered Anchor Links
No, avoiding anything potentially negative is not what I'm saying. Your argument (that context always matters) leaves discourse and society highly susceptible to dog-whistles[1], by forcing all good-faith participants to interpret all communication in the most generous way possible. Bad-faith participants, on the other hand, are free to exploit that generosity.

By calling out and avoiding dog-whistles, even including accidental Nazi slogans (once pointed out), we reduce the impact of this attack on good-faith discussion and actual increase the level of openness and being up-front with our opinions.

One key difference between this and virtue signaling or thought policing is that it's the specific wording that is avoided, and not the underlying thoughts or opinions.

[1]: https://en.wikipedia.org/wiki/Dog_whistle_(politics)

numeri··on Overengineered Anchor Links
It was the term invented by the architects of the Holocaust, and I disagree that "eh, context matters".

Setting all moral arguments aside, it's important to know that similar phrases can work as dog-whistles to signal belonging to radical groups, and as such can easily give people the wrong impression about you as an author.

If I were to see a blog post titled "Work will set you free"[1] written by a peer, prospective employee/employer, colleague, etc., it would immediately set off alarm bells in my mind – even if the content of the post is a completely innocent discussion of the uplifting benefits of buckling down on one's workload. At best, it implies lack of awareness – at worst, it implies some extremely hateful beliefs and desires.

[1]: Written above the entrance to the Nazi concentration camps as a false promise encouraging prisoners often destined for death to work hard in forced labor.

numeri··on Ask HN: Any insider takes on Yann LeCun's push against current architectures?
No, the person you're responding to is absolutely right. The easy test (which has been done in papers again and again) is the ability to train linear probes (or non-linear classifier heads) on the current hidden representations to predict the nth-next token, and the fact that these probes have very high accuracy.
numeri··on Looking Back at Speculative Decoding
Do you have a good reference I can read up on? I'd love to learn a bit more and update my mental "citation"
numeri··on Looking Back at Speculative Decoding
I've been slightly annoyed by how the Speculative Decoding paper has gotten all the credit for the technique – I first learned about the technique from a paper more than a year older[1], Shallow Aggressive Decoding.

They introduce the same method, but apply it to grammatical error correction, meaning the "draft" output is just the input itself. The Speculative Decoding paper tries to emphasize differences between this and their method, saying that theirs is more general, as they apply it to more domains, allowing the draft to come from a smaller model, and extend it to allow sampling.

All of that is great, and deserves another paper, but doesn't deserve the credit for inventing and rights to rename the method, especially when they were aware of Shallow Aggressive Decoding before uploading their first draft.

[1]: https://arxiv.org/abs/2106.04970

numeri··on When imperfect systems are good: Bluesky's lossy timelines
This behavior started happening for me in the last few months. If I click on a result, then go back, I have different search results.

I've found a workaround, though – click back into the DDG search box at the top of the page and hit enter. This then returns the original search results.

numeri··on I had to take down my course-swapping site or be expelled
The First Amendement limits what the US Congress/government can do, not what a private person can do.

> Congress shall make no law respecting an establishment of religion, or prohibiting the free exercise thereof; or abridging the freedom of speech, or of the press; or the right of the people peaceably to assemble, and to petition the Government for a redress of grievances.

numeri··on Show HN: I made the slowest, most expensive GPT
But gpt-4o can already answer your question for a fraction of the price and time:

To determine which of the items could be used to make an actual fire, we need to analyze the definitions provided:

1. *Glaarg*: A wooden item.

2. *Bliirg*: A non-wooden item.

3. *Neerg*: A non-existent thing.

4. *Eeerg*: A thing that actually exists.

Now, let's look at the specific terms:

- *Bipk*: A glaarg (wooden item) that is also a neerg (non-existent thing). Since it is non-existent, it cannot be used to make a fire.

- *Vokp*: A glaarg (wooden item) that is also an eeerg (existent thing). Since it is a wooden item that exists, it can be used to make a fire.

- *Jokp*: A bliirg (non-wooden item) that is also an eeerg (existent thing). While it exists, it is non-wooden, so it may not be suitable for making a fire depending on its material.

- *Fhup*: A bliirg (non-wooden item) that is also a neerg (non-existent thing). Since it is non-existent, it cannot be used to make a fire.

Based on this analysis, the only item that can be used to make an actual fire is a *vokp*, as it is a wooden item that exists.

numeri··on Scrabble star wins Spanish world title despite not speaking Spanish
There was at least one case of someone trying this against Nigel in French, and he got them.
numeri··on Scrabble star wins Spanish world title despite not speaking Spanish
Nigel Richards usually outperforms (and is orders of magnitude faster than, at least for the complicated endgames, if I understand correctly) the best computer Scrabble programs.
numeri··on Scrabble star wins Spanish world title despite not speaking Spanish
Yes, from what I understand – and they also make mistakes.
numeri··on "This is not a joke, Funko just called my mom"
The big issue here is that they didn't issue a DMCA request, they reported them for fraud.
numeri··on Intel announces Arc B-series "Battlemage" discrete graphics with Linux support
GPU inference is always a balancing act, trying to avoid bottlenecks on memory bandwidth (loading data from the GPU's global memory/VRAM to the much smaller internal shared memory, where it can be used for calculations) and compute (once the values are loaded).

Splitting the model up between several GPUs would add a third much worse bottleneck – memory bandwidth between the GPUs. No matter how well you connect them, it'll be slower than transfer within a single GPU.

Still, the fact that you can fit an 8× larger GPU might be worth it to you. It's a trade-off that's almost universally made while training LLMs (sometimes even with the model split down both its width and length), but is much less attractive for inference.

numeri··on Senators say TSA's facial recognition program is out of control
The point of this program isn't that it makes things substantially quicker at the checkpoint – it is a minor speed-up at best. The goal is to normalize the collection of biometric data, to shift the Overton window of surveillance.
numeri··on Senators say TSA's facial recognition program is out of control
I've declined several times now, and have gotten harassed about it about 3/4 times. Whoever designed the program really did a good job getting buy-in from the lower-level employees.
numeri··on 25% of Adults Suspect Undiagnosed ADHD
My take on this: You're absolutely allowed to speak on your experience, but you should also take the feedback at face value and try to figure out whether it's useful or not. Plenty of feedback here on HN is good, but of course there's also plenty that misses the mark.

That being said, an ADHD diagnosis at the age of 2 goes against all current best practices, so it might be good to revisit the subject and ask to be evaluated again.

Even in the likely case that it's confirmed, you might get additional support, advice or diagnoses that can help you in the future.

← PreviousPage 3 of 6Next →