HNHacker News
TopNewBestAskShowJobs

numeri

985 karma · joined January 27, 2022

submissionscomments
numeri··on Intel announces Arc B-series "Battlemage" discrete graphics with Linux support
GPU inference is always a balancing act, trying to avoid bottlenecks on memory bandwidth (loading data from the GPU's global memory/VRAM to the much smaller internal shared memory, where it can be used for calculations) and compute (once the values are loaded).

Splitting the model up between several GPUs would add a third much worse bottleneck – memory bandwidth between the GPUs. No matter how well you connect them, it'll be slower than transfer within a single GPU.

Still, the fact that you can fit an 8× larger GPU might be worth it to you. It's a trade-off that's almost universally made while training LLMs (sometimes even with the model split down both its width and length), but is much less attractive for inference.

numeri··on Senators say TSA's facial recognition program is out of control
The point of this program isn't that it makes things substantially quicker at the checkpoint – it is a minor speed-up at best. The goal is to normalize the collection of biometric data, to shift the Overton window of surveillance.
numeri··on Senators say TSA's facial recognition program is out of control
I've declined several times now, and have gotten harassed about it about 3/4 times. Whoever designed the program really did a good job getting buy-in from the lower-level employees.
numeri··on 25% of Adults Suspect Undiagnosed ADHD
My take on this: You're absolutely allowed to speak on your experience, but you should also take the feedback at face value and try to figure out whether it's useful or not. Plenty of feedback here on HN is good, but of course there's also plenty that misses the mark.

That being said, an ADHD diagnosis at the age of 2 goes against all current best practices, so it might be good to revisit the subject and ask to be evaluated again.

Even in the likely case that it's confirmed, you might get additional support, advice or diagnoses that can help you in the future.

numeri··on AGI is far from inevitable
I think their "proof" would also prove that no child could ever learn to behave like an adult.
numeri··on Training Language Models to Self-Correct via Reinforcement Learning
I've got bad news for you – that term was used in deep learning research well before LLMs came on the scene. It has nothing to do with pundits trying to popularize anything or trying to justify LLMs' shortcomings, it was just a label researchers gave to a phenomenon they were trying to study.

A couple papers that use it in this way prior to LLMs:

- 2021: The Curious Case of Hallucinations in Neural Machine Translation (https://arxiv.org/abs/2104.06683)

- 2019: Identifying Fluently Inadequate Output in Neural and Statistical Machine Translation (https://aclanthology.org/W19-6623/)

numeri··on Training Language Models to Self-Correct via Reinforcement Learning
Sort of like this? It does help: Source-Aware Training Enables Knowledge Attribution in Language Models (https://arxiv.org/abs/2404.01019)

From the abstract:

> ... To give LLMs such ability, we explore source-aware training -- a recipe that involves (i) training the LLM to associate unique source document identifiers with the knowledge in each document, followed by (ii) an instruction-tuning stage to teach the LLM to cite a supporting pretraining source when prompted.

numeri··on Training Language Models to Self-Correct via Reinforcement Learning
OpenAI stated [1] that one of the breakthroughs needed for o1's train of thought to work was reinforcement learning to teach it to recover from faulty reasoning.

> Through reinforcement learning, o1 learns to hone its chain of thought and refine the strategies it uses. It learns to recognize and correct its mistakes. It learns to break down tricky steps into simpler ones. It learns to try a different approach when the current one isn’t working.

That's incredibly similar to this paper, which is discusses the difficulty in finding a training method that guides the model to learn a self-correcting technique (in which subsequent attempts learn from and improve on previous attempts), instead of just "collapsing" into a mode of trying to get the answer right with the very first try.

[1]: https://openai.com/index/learning-to-reason-with-llms/

numeri··on Mistral NeMo
SentencePiece is a tool and library for training and using tokenizers, and supports two algorithms: Byte-Pair Encoding (BPE) and Unigram. You could almost say it is the library for tokenizers, as it has been standard in research for years now.

Tiktoken is a library which only supports BPE. It has also become synonymous with the tokenizer used by GPT-3, ChatGPT and GPT-4, even though this is actually just a specific tokenizer included in tiktoken.

What Mistral is saying here (in marketing speak) is that they trained a new BPE model on data that is more balanced multilingually than their previous BPE model. It so happens that they trained one with SentencePiece and the other with tiktoken, but that really shouldn't make any difference in tokenization quality or compression efficiency. The switch to tiktoken probably had more to do with latency, or something similar.

numeri··on Vision language models are blind
I'm also confused about some of the figures' captions, which don't seem to match the results:

- "Only Sonnet-3.5 can count the squares in a majority of the images", but Sonnet-3, Gemini-1.5 and Sonnet-3.5 all have accuracy of >50%

- "Sonnet-3.5 tends to conservatively answer "No" regardless of the actual distance between the two circles.", but it somehow gets 91% accuracy? That doesn't sound like it tends to answer "No" regardless of distance.

numeri··on The Race to Seal Helium HDDs (2021)
Any career in fundamental research is more or less like that. From what I've seen personally, academics and government labs are the two biggest places you can find the most open ended roles. Each comes with their own caveats, of course.
numeri··on GPT-4.5 or GPT-5 being tested on LMSYS?
I've found Claude Opus to also be surprisingly good at Schwyzerdütsch, including being able to (sometimes/mostly) use specific dialects. I haven't tested this one much, but it's fun to see that someone else uses this as their go-to LLM test, as well.
numeri··on Why the OpenSSL punycode vulnerability was not detected by fuzz testing (2022)
Is "meringues me" a typo, or a really fun new vocab word for me?
numeri··on Llama 3 8B is almost as good as Wizard 2 8x22B
Yeah, this is correct and I'm not sure what paper GP was thinking of – Chinchilla is only about finding the point at which it would be more useful to scale the model rather than training longer.

Chinchilla optimal scaling is not useful if you want to use the model, just if you want to beat some other model on some metric for the minimal training costs.

numeri··on Notes on El Salvador
I believe this was one of Machiavelli's big arguments in The Prince – that sometimes a country in crisis needs a single strong leader/monarch/dictator, using cruelty if necessary to keep control and bring stability.
numeri··on LLMs use a surprisingly simple mechanism to retrieve some stored knowledge
I didn't imply that they know anything about where atoms are, I was just pointing out the sheer absurdity of that volume of data.

I should make it clear that my comparison there is unfair and mostly just funny – you don't need to store every possible combination of 10 tokens, because most of them will be nonsense, so you wouldn't actually need that much storage. That being said, it's been fairly solidly proven that LLMs aren't just lookup tables/stochastic parrots.

numeri··on LLMs use a surprisingly simple mechanism to retrieve some stored knowledge
Your suggested scheme (assuming a mapping from 10 tokens to 10 tokens, with each token taking 2 bytes to store) would take (32000 * 20) * 2 bytes = 2.3e78 TiB of storage, or about 250 MiB per atom in the observable universe (1e82), prior to compression.

I think it's more likely that LLMs are actually learning and understanding concepts as well as memorizing useful facts, than that LLMs have discovered a compression method with that high of a compression ratio, haha.

numeri··on LLMs use a surprisingly simple mechanism to retrieve some stored knowledge
This is research, trying to understand the fundamentals of how these models work. They weren't actually trying to find out where Bill Bradley went to university.
numeri··on DenseFormer: Enhancing Information Flow in Transformers
They try this in the appendix without success, unfortunately. It seems having this enabled early on in training is important.
numeri··on Ask HN: If you've used GPT-4-Turbo and Claude Opus, which do you prefer?
Claude Opus, by a lot. It is especially good with the few low-resource languages that I or people I know could test, including several German/Swiss German dialects and Azerbaijani!
numeri··on Claude 3 model family
That has nothing to do with the idea of ensembling multiple specialized/single-purpose models. Mixture of Experts is an method of splitting the feed-forwards in a model such that only a (hopefully) relevant subset of parameters is run for each token.

The model learns how to split them on its own, and usually splits based not on topic or domain, but on grammatical function or category of symbol (e.g., punctuation, counting words, conjunctions, proper nouns, etc.).

numeri··on Esperanto, Toki Pona, Swahili, Indonesian
Swahili and Indonesian are primarily second languages, used as lingua francas amongst large and diverse populations, so linguistic changes are mainly driven by non-native speakers, as opposed to the languages you listed
numeri··on Keep your phone number private with Signal usernames
I believe what the grandparent comment meant was that you can't run a server that participates in the public network, not that you can't run a private server. That was my prior understanding, at least.

I might very well be wrong, and if so, someone please correct me.

numeri··on RLHF a LLM in <50 lines of Python
I don't know about strictly superior. It's certainly strictly easier for people with a budget, who just need "good enough" results the first try. I don't have any evidence whatsoever, but I'd expect that enough tuning and retries can get squeeze a bit more performance out of RLHF than you can get out of DPO.
numeri··on Mistral CEO confirms 'leak' of new open source AI model nearing GPT4 performance
The linked leaderboard is actually very trustworthy, in that it consists not of scores on a test dataset, but of ELO ratings generated by actual humans' ratings of the models' responses.

You can go and enter any prompt you like, wait a bit, and then get two LLM responses back, which you can then rank or mark as tied, after which you'll be shown which model each came from. Maybe you already knew this, maybe you didn't. In any case, I don't see any real way for Goodhart's Law to apply here – the metric and the goal are the same here, i.e., human approval of answers.

numeri··on LiteLlama-460M-1T has 460M parameters trained with 1T tokens
Almost all LLM inference these days includes a repetition penalty, to help prevent models from falling into this pattern, the so-called "boredom trap" [1]. Once the model repeats itself once, that greatly increases the chances that the next best token is another repetition (for a raw model like this, at least, that hasn't been chat/instruct/RLHF fine-tuned).

I ran the same prompt with a repetition penalty of 1.1 and got this:

The mountain in front of me was a little higher than the other two, and I could see the top of the mountain.

I'm going to climb it," I said. "It's a good place for a picnic."

You're not going to climb it?"

No!" I said. "But I'll climb it if you want! I can climb it with you—if we can get out there together—"

[1] https://arxiv.org/abs/2007.14966

numeri··on Andrej Karpathy on Hallucinations
It's currently a major area of research and an unsolved problem to find out what individual weights do – and the most recent research seems to suggest that there is not a one-to-one relationship between ideas and weights. In fact, one line of research is showing that each weight encodes multiple ideas with what is called polysemanticity [1], while another line of research seems to show that individual factoids are spread across thousands of weights and possibly even multiple layers [2].

[1]: https://transformer-circuits.pub/2022/toy_model/index.html [2]: https://rome.baulab.info/

numeri··on Mistral "Mixtral" 8x7B 32k model [magnet]
You're not necessarily wrong, but I'd imagine this is almost prohibitively slow. Also, this model seems to use two experts per token.
numeri··on Oracle of Zotero: LLM QA of Your Research Library
I spent far too long trying to figure that out as well. It's a much catchier name, for sure, but sort of silly that it has so many forks itself.
numeri··on Exponentially faster language modelling
I would have to go back and reread the paper to be sure, but FF layers are applied position-wise, meaning independently and in parallel on all input tokens/positions. Because of that, I could imagine contexts where the sequence dimension isn't relevant, i.e., for computational complexity.
← PreviousPage 4 of 6Next →