Qwen2.5-1M: Deploy your own Qwen with context length up to 1M tokens
qwenlm.github.io
qwenlm.github.io
(via https://news.ycombinator.com/item?id=42832838, but we merged that thread hither)
Developing aider, I've seen this problem with gpt-4o, Sonnet, DeepSeek, etc. Many aider users report this too. It's perhaps the #1 problem users have, so I created a dedicated help page [0].
Very large context may be useful for certain tasks with lots of "low value" context. But for coding, it seems to lure users into a problematic regime.
[0] https://aider.chat/docs/troubleshooting/edit-errors.html#don...
I've used the giant context in Gemini to dump a code base and say: describe the major data structures and data flows.
Things like that, overview documents, work great. It's amazing for orienting in an unfamiliar codebase.
Too much context can still confuse the LLMs in this situation, but they may be somewhat more resilient.
Many conflicting ideas are harder for models to follow than one large unified idea.
So we may have got to a local maximum regarding code helpers with LLMs and we'll have to wait for some breakthrough in the AI field before we get something better.
But the good thing is that DeepSeek proved those breakthroughs are going to happen one way or another, fast.
In one dialog, some 30k tokens later, Claude requested the contents of package.json... which was in the context window already - the whole file!
The strange thing was that after I said so, without re-inserting, Claude successfully read it from context to fill the gap in what it was trying to do.
It's as if a synopsis of what exists in-context delivered with each message would help. But that feels weird!
Most chat is just a long running prompt. LLMs have zero actual memory. You just keep feeding it history.
Maybe I misunderstood what you're saying but what you're describing is some kind of 2nd model that condenses that history and that gets fed; this has been done.
Really, what you probably need is another model managing the heap and the stack of the history and brining forward the current context.
But that's easy to say because we are humans.
I didn't 'fix' the problem by re-inserting the package.json. I just gave a reminder that it that it was already in context.
"I could confirm this root cause if I could see the contents of package.json".
"You can see it."
"Whoops, yes. Package x needs to bump from n to n+1."
The point being, even info inside the current context can actively be overlooked unless it is specifically referenced.
Imagine dumping the entire text of a large code repository in front of a human programmer, and asking them to fix a bug. Human programmers use IDEs, search through the code, flip back and forth between different functions, etc. Maybe with a better interface that the LLM could interact with, it would perform better.
Something: Filename: index.js Content: | Class Example...
Another item would be some kind of hyperlinking. Maybe you could load in a hrefs but there might be a more semantically popular way, but the data feeding these AIs just aren't constructed like that.
- Summarize topics (with references to shows) - Find quotes specific to a topic (again with references)
Anything above 32k tokens fails to have acceptable recall, across GPT-4o, Sonnet, and Google's Gemini Flash 1.5 and 2.0.
I suppose it kind of makes sense, given how large context windows are implemented via things like sparse attention etc.
That's a product choice from openAI, not a problem with LLM's long context window input capability.
YES that was it:
files-to-prompt \
~/Dropbox/Development/llm \
-e py -c | \
llm -m q1m 'describe this codebase in detail' \
-o num_ctx 80000
I was watching my memory usage and it quickly maxed out my 64GB so I hit Ctrl+C before my Mac crashed.Even though the new DCA based on which these models provide long context could be an interesting area to watch;
1M tokens will definitely require a lot of KV cache memory. One way to reduce the memory footprint is to use KV cache quantization, which has recently been added behind a flag [3] and will 1/4 the memory footprint if 4-bit KV cache quantization is used (OLLAMA_KV_CACHE_TYPE=q4_0 ollama serve)
[1] https://arxiv.org/pdf/2309.06180
[2] https://github.com/microsoft/vattention
[3] https://smcleod.net/2024/12/bringing-k/v-context-quantisatio...
In this LLM era, those are rookie numbers. It should be possible to get a Mac with a lesser processor but at least 256GB of memory for $2000. I realize part of the issue is the lead time for chip design -- since Mac memory is an integral part of the chip, and the current crop were designed before the idea of running something like an LLM locally was a real probability.
But I hope the next year or two show significant increases in the default (and possible) memory for Macs.
Apple is not known for leaving money on the table like that.
Also, projects like NVidia DIGITS ($2k for 128G) might make Apple unwilling to enter the market. As you said, Studio with 192G is $5600k. For purely AI purposes, two DIGITS' are a better choice, and non-AI usages don't need such ludicros amount of RAM (maybe for video, but those customers are willing to pay more).
True -- although I will say the M series chips were a step change in performance and efficiency from the Intel processors they replaced, and Apple didn't charge a premium for them.
I'm not suggesting that they'll stop charging more for RAM than the industry at large -- I'm hoping they'll unbundle RAM from CPU-type. A base Mac Mini goes for $600, and adding RAM costs $200 per 8GB. That's a ridiculous premium, clearly, and at that rate my proposed Mac Mini with 256GB of RAM would go for $6600 -- which would roll my eyes until they fell out of my head.
But Apple is also leaving money on the table if they're not offering a more expensive model people would buy. A 128GB Mini, let's say, for $2000, might be that machine.
All that said, it's also a heck of a future-proof machine, so maybe the designed-obsolescence crowd have an argument to make here.
https://github.com/taketwo/llm-ollama/blob/4ccd5181c099af963...
That 2k default is extremely low, and ollama *silently* discards the leading context. So users have no idea that most of their data hasn’t been provided to the model.
I’ve had to add docs [0] to aider about this, and aider overrides the default to at least 8k tokens. I’d like to do more, but unilaterally raising the context window size has performance implications for users.
Edit: Ok, aider now gives ollama users a clear warning when their chat context exceeds their ollama context window [1].
[0] https://aider.chat/docs/llms/ollama.html#setting-the-context...
[1] https://github.com/Aider-AI/aider/blob/main/aider/coders/bas...
Fortunately it's easy to create a variant of the model with increased context size using the CLI[3] and then use that variant instead.
Just be mindful that longer context means more memory required[4].
[1]: https://github.com/ollama/ollama/issues/4967
[2]: https://github.com/ollama/ollama/issues/7043
[3]: https://github.com/ollama/ollama/issues/8099#issuecomment-25...
[4]: https://www.reddit.com/r/LocalLLaMA/comments/1848puo/comment...
$ ollama run llama3.2
>>> /set parameter num_ctx 32768
Set parameter 'num_ctx' to '32768'
>>> /save llama3.2-32k
Created new model 'llama3.2-32k'
>>> /bye
$ ollama run llama3.2-32k "Summarize this file: $(cat README.md)"
...
The table in the reddit post above also shows context size vs memory requirements for Model: 01-ai/Yi-34B-200K
Params: 34.395B
Mode: infer Sequence Length vs Bit Precision Memory Requirements
SL / BP | 4 | 6 | 8 | 16
--------------------------------------------------------------
256 | 16.0GB | 24.0GB | 32.1GB | 64.1GB
512 | 16.0GB | 24.1GB | 32.1GB | 64.2GB
1024 | 16.1GB | 24.1GB | 32.2GB | 64.3GB
2048 | 16.1GB | 24.2GB | 32.3GB | 64.5GB
4096 | 16.3GB | 24.4GB | 32.5GB | 65.0GB
8192 | 16.5GB | 24.7GB | 33.0GB | 65.9GB
16384 | 17.0GB | 25.4GB | 33.9GB | 67.8GB
32768 | 17.9GB | 26.8GB | 35.8GB | 71.6GB
65536 | 19.8GB | 29.6GB | 39.5GB | 79.1GB
131072 | 23.5GB | 35.3GB | 47.0GB | 94.1GB
* 200000 | 27.5GB | 41.2GB | 54.9GB | 109.8GB
* Model Max Context Size
Code: https://gist.github.com/lapp0/d28931ebc9f59838800faa7c73e3a0...[1]: https://medium.com/@plienhar/llm-inference-series-4-kv-cachi...
Maybe they can take some of those hundreds of billions and invest in new approaches.
Because racks of H100s are not sustainable. But it's clear that increasing the amount of memory available is key to getting more intelligence or capabilities.
Maybe there is a way to connect DRAM with photonic interconnects that doesn't require much data ordering for AI if the neural network software model changes somewhat.
Is there something that has the same capabilities of a transformer but doesn't operate on sequences?
If I was a little smarter and had any math ability I feel like I could contribute.
But I am smart enough to know that just building bigger and bigger data centers is not the ideal path forward.
There's just not enough capacity to build memory fast enough right now. Everyone needs the biggest and fastest modules they can get, since it directly impacts the performance of the models.
There's still a lot of happening to improve memory, like the latest Titans paper: https://arxiv.org/abs/2501.00663
So I think until a breakthrough happens or the fabs catch up, it'll be this painful race to build more datacenters.
Huh? Racks of H100s are the most sustainable thing we can have for LLMs for now.
So it seems that we need a new paradigm of some sort.
So much investment is being announced for data centers. I assumed there would be more investments in fundamental or applied research. Such as for scaling memristors or something.
[1]: https://cerebras.ai/press-release/cerebras-systems-announces...
[2]: https://cerebras.ai/chip/announcing-the-cerebras-architectur...
https://github.com/ml-explore/mlx-examples/pull/956
edit: here's a quick example for qwen2.5-1M from a mlx dev
https://huggingface.co/ai21labs/AI21-Jamba-1.5-Mini 256k https://huggingface.co/THUDM/glm-4-9b-chat-1m 1M
and many other's that supposedly extended traditional models via finetune/rope scaling
lm_load_print_meta: general.name = Qwen2.5 7B Instruct 1M
llm_load_print_meta: BOS token = 151643 '<|endoftext|>'
llm_load_print_meta: EOS token = 151645 '<|im_end|>'
llm_load_print_meta: EOT token = 151645 '<|im_end|>'
llm_load_print_meta: PAD token = 151643 '<|endoftext|>'
llm_load_print_meta: LF token = 148848 'ÄĬ'
llm_load_print_meta: FIM PRE token = 151659 '<|fim_prefix|>'
llm_load_print_meta: FIM SUF token = 151661 '<|fim_suffix|>'
llm_load_print_meta: FIM MID token = 151660 '<|fim_middle|>'
llm_load_print_meta: FIM PAD token = 151662 '<|fim_pad|>'
llm_load_print_meta: FIM REP token = 151663 '<|repo_name|>'
llm_load_print_meta: FIM SEP token = 151664 '<|file_sep|>'
llm_load_print_meta: EOG token = 151643 '<|endoftext|>'
llm_load_print_meta: EOG token = 151645 '<|im_end|>'
llm_load_print_meta: EOG token = 151662 '<|fim_pad|>'
llm_load_print_meta: EOG token = 151663 '<|repo_name|>'
llm_load_print_meta: EOG token = 151664 '<|file_sep|>'
llm_load_print_meta: max token length = 256Running this on a M4 max
So even models like llama3 8b say they have a larger context, but they really don’t in practice. I have a hard time getting past 8k on 16gb vram (you can definitely set the context length higher, but the quality and speed degradation is obvious).
I’m curious how people are doing this on modest hardware.
All in all, not a cheap hobby ( if you are not doing it for work ).
But what about output? I want to generate a few thousand lines of code, anyone got any tips?
Now you just need to convince it to output that much :)
> have you been trained using openai?
No, I have not been trained using technology or data from OpenAI. I am an artificial intelligence model developed by Yandex, called YandexGPT. My “training” is based on proprietary algorithms and data that was collected and processed by the Yandex team. While I have similarities with other models such as those developed by OpenAI (e.g. GPT-3), my training was conducted independently and uses unique approaches specific to Yandex. This includes using Russian-language data as well as other sources to provide a better understanding of context and provide useful information to users. If you have questions on topics related to AI technologies or anything else, I'd be happy to help!
second, how does one increase the context window without requiring obscene amounts of RAM? we're really hitting the limitations of the transformer architecture's quadratic scaling...
[1] https://research.google/blog/chain-of-agents-large-language-...
Here, the prose says "nearly perfect", the graph is all green except for a little yellow section, and you have to parse a 96 cell table, having familiarity with several models and technical techniques to get the real # (84.4%, and that tops out at 128K, not anywhere near the claimed 1M)
I don't bring this up to denigrate, but rather to highlight that "nearly perfect" is quite far off still. Don't rely on long context for anything you build
> Even models trained on just 32K tokens, such as the Qwen2.5-7B-Instruct, achieve nearly perfect accuracy in passkey retrieval tasks with 1M-token contexts.
Which is pages after the graph and table you mentioned, which are clearly introduced as
(Graph)
> First off, we evaluate the Qwen2.5-1M models on the Passkey Retrieval task with a context length of 1 million tokens. The results show that these models can accurately retrieve hidden information from documents containing up to 1M tokens, with only minor errors observed in the 7B model.
(Table)
> For more complex long-context understanding tasks, we select RULER, LV-Eval, LongbenchChat used in this blog.
That you went so deep into the post to find your “clever” phrase to complain about tells me you’re probably being intentionally misleading. Most readers won’t read that far and ones that do certainly won’t leave with an impression that this is “nearly perfect” for complex tasks.
You're attempting to imply the rest of the sentence adds context that makes pulling out "nearly perfect" incorrect. Can you explain?
> ...
I'm not sure what the rest of the quotes are implying, as you just copy and paste and don't provide any indication of what you're communicating by sharing them. Can you explain more?
> That you went so deep into the post
It's the 587th word, less than 2 minutes reading at average reading speed.
> you’re probably being intentionally misleading.
!?!?!
#1) I'm certainly not intentionally misleading.
#2) What is misleading about "they say nearly perfect and then the highest # I can steelman from the table is 84%?"
#3) This is the first time in 15 years on HN that I've had someone accuse me of being intentionally misleading. Part of that is because there's numerous rules against that sort of dialogue. The remaining part is people, at least here, are usually self-interested enough to not make up motivations for other people feeling differently from them.
And I’d hardly call it obscene. You can buy a Mac Studio with 192GB of memory, that should allow you to max out the context window of the 7B model. Probably not going to be very fast though.
Technology has never been class-agnostic or universally accessible.
Even saying that, I would argue that there is more, not less, technology that is accessible to more people today than there ever has been.
And you can easily get a dev job in Norway without having to run an LLM locally on your computer.
There are many such threads on Reddit. M4 Max is incrementally faster, maybe 20%. Even if you factor in electricity costs, a 2x 3090 setup is IMO the sweet spot, cost/benefit wise.
And it’s maybe a zany line of argumentation, but 2x 3090 use 10x the power of an M4 Max. While the M4 is maybe the most efficient setup out there, it’s not nearly 10x as efficient. That’s IMO where the lack of compute power comes from.
The problem is that most ML models are released for NVIDIA CUDA. Getting them to work on macOS requires translating them, usually to either GGUF (the llama.cpp format) or MLX (using Apple's own MLX array framework).
As such, as a Mac user I remain envious of people with NVIDIA/CUDA rigs with decent amounts of VRAM.
The NVIDIA "Digits" product may change things when it ships: https://www.theverge.com/2025/1/6/24337530/nvidia-ces-digits... - it may become the new cheapest convenient way to get 128GB of GPU-accessible RAM for running models.
I bet that 5 3090s will smoke a Mac Studio. Can't find anyone in Norway with any in stock though. Or any 4090s with 24GB of memory.
You can get a nVidia RTX 5000 with 32GB of memory, there are two webshops that have those in stock. You'll need to wait though, because it looks like there might be one or maybe two in stock in total. And they are 63 000 NOK, and you need 4 of them. At that price you can buy two Mac Studios though.
I see people selling 3090s with 24GB secondhand for around 10 000 NOK each, but those have been running day in and day our for 3 years and don't come with a warranty.
BTW a Mac wouldn’t be able to run a model with 120GB requirements, 8GB for the rest is likely too tight a fit.
> For processing 1 million-token sequences:
> Qwen2.5-7B-Instruct-1M: At least 120GB VRAM (total across GPUs).
> Qwen2.5-14B-Instruct-1M: At least 320GB VRAM (total across GPUs).