HNHacker News
TopNewBestAskShowJobs

rain1

1,127 karma · joined May 15, 2018

submissionscomments
rain1··on Transformer Architecture Visualizer
I've used Google Antigravity to write scripts to download and produce architecture diagrams for various LLMs from huggingface. It's pretty useful so I thought I'd share it.

There's also a model comparison spreadsheet that you can compare sizes and such https://weavers.neocities.org/architecture-encyclopedia/mode...

If you'd like any additional models to be added I can add them in.

rain1··on How large are large language models?
The Gemma models are too small to be included in this list.

You're right the T5 stuff is very important historically but they're below 11B and I don't have much to say about them. Definitely a very interesting and important set of models though.

rain1··on How large are large language models?
Yes but just purely in terms of entropy, you can't make a model better than GPT-4 by training it on GPT-4 outputs. The limit you would converge towards is GPT-4.
rain1··on How large are large language models?
This is kind of related to the jack morris post https://blog.jxmo.io/p/there-are-no-new-ideas-in-ai-only he discusses how the big leaps in LLMs have mostly come - not so much from new training methods or arch. changes as such - but the ability of new archs. to ingest more data.
rain1··on How large are large language models?
It's extremely interesting how powerful a language model is at compression.

When you train it to be an assistant model, it's better at compressing assistant transcripts than it is general text.

There is an eval which I have a lot of interested in and respect for https://huggingface.co/spaces/Jellyfish042/UncheatableEval called UncheatableEval, which tests how good of a language model an LLM is by applying it on a range of compression tasks.

This task is essentially impossible to 'cheat'. Compression is a benchmark you cannot game!

rain1··on How large are large language models?
I think that one thing that this chart makes visually very clear is the point I about GPT-3 being such a huge leap, and there being a long gap before anybody was able to match it.
rain1··on How large are large language models?
This is really awesome. Thank you for creating that. I included a screenshot and link to the chart with credit to you in a comment to my post.
rain1··on How large are large language models?
I can correct mistakes.

> it somehow merged Llama 4 Maverick's custom Arena chatbot version with Behemoth

I can clarify this part. I wrote 'There was a scandal as facebook decided to mislead people by gaming the lmarena benchmark site - they served one version of llama-4 there and released a different model' which is true.

But it is inside the section about the llama 4 model behemoth. So I see how that could be confusing/misleading.

I could restructure that section a little to improve it.

> Llama 405B was also trained on more than 15 trillion tokens[1],

You're talking about Llama 405B instruct, I'm talking about Llama 405B base. Of course the instruct model has been traiend on more tokens.

> why is there such a focus on token training count?

I tried to include the rough training token count for each model I wrote about - plus additional details about training data mixture if available. Training data is an important part of an LLM.

rain1··on How large are large language models?
I have corrected that. It was supposed to say "None of this document was written by AI."

Thank you for spotting the error.

rain1··on Tell HN: Burnout is bad to your brain, take care
> Take care of your mental health

How?

rain1··on Llamafile Not Making Sense
todsacerdoti is a spambot btw
rain1··on Why does all() return True if the iterable is empty?
I don't understand this. Please can you point me to information about it?
rain1··on Why does all() return True if the iterable is empty?
The people that are astonished by this just need to learn why.

It's not the function that is wrong, it's those people.

rain1··on Tricking Monty Hall
This is incorrect, the goats and car are behind doors. They are not inside cardboard boxes.
rain1··on Claude 2
This is an example of hallucination.

An LLM doesn't know anything about itself - it can be pre-prompted with facts about itself, but this is going to be an example of it just making plausible text up.

rain1··on A Mechanistic Interpretability Analysis of Grokking
tell me you're posting from an armchair without telling me you're posting from an armchair
rain1··on Former Dolphin team member addresses Steam/Valve’s takedown of Dolphin emulator
what the actual fuck were they thinking uploading dolphin to steam??
rain1··on ChatGPT conversations can be shared publicly
Why don't they let us edit what the bot says? Could be useful.
rain1··on Io uring
This is the future of linux syscalls. Get on board with this or get left behind.
rain1··on PrivateGPT
so list a few known to work models and their requirements
rain1··on Ask HN: I have 176 logins/accounts. How many do you have?
I feel like this isn't healthy for a human being
rain1··on Let ChatGPT visit a website and have your email stolen
There are no solutions
rain1··on A guidance language for controlling LLMs
This looks incredible. Wow.
rain1··on A guidance language for controlling LLMs
Does this do one query per {{}} thing?
rain1··on Run Llama 13B with a 6GB graphics card
unfortunately that chip is proprietary and undocumented, it's very difficult for open source programs to make use of. I think there is some reverse engineering work being done but it's not complete.
rain1··on Run Llama 13B with a 6GB graphics card
wonderful! thank you
rain1··on Run Llama 13B with a 6GB graphics card
> 1. Download the weights for the model you want to use, e.g. gpt4-x-vicuna-13B.ggml.q5_1.bin

I think you need to quantize the model yourself from the float/huggingface versions. My understanding is that the quantization formats have changed recently. and old quantized models no longer work.

rain1··on Run Llama 13B with a 6GB graphics card
That is a crazy speedup!!
rain1··on Run Llama 13B with a 6GB graphics card
I don't think the integrated GPU on that supports CUDA. So you will need to use CPU mode only.
rain1··on Run Llama 13B with a 6GB graphics card
Tell us how it goes! Try different numbers of layers if needed.

A good place to dig for prompt structures may be the 'text-generation-webui' commit log. For example https://github.com/oobabooga/text-generation-webui/commit/33...

Page 1 of 5Next →