Using mmap to make LLaMA load faster
justine.lol
justine.lol
Overall, these people would be better off taking their drama on Twitter or LinkedIn. ggerganov did the right thing kicking them out.
Good lord, it's terrible when the peanut gallery feels like they have to comment on development practice. Why would numbers of commits be a relevant metric in an Open Source project? Of course squashed commits are easier to handle during reabses and such, and when that work can be squashed to a single "initial mmap support" commit, then that's fine.
> @jart rewrote @slaren's code, which slaren wrote first
now this is just kindergarten level of arguments
What I'd care more about is whether this issue had harmed the project in a technical way, which, unfortunately, it apparently has.
Also note the stats on GH subscribers and stuff. This is a lolcow dossier…
Events like this make me glad I don’t contribute OSS. I’ll keep my coombots proprietary.
I am not much familiar with her work except the impressive Cosmopolitan / Redbean mentioned on HN in the past. But she seems to be quite a controversial figure that is for some weird technocracy and against democracy and leftists, despite being a leader in the zucotti park protests… in short, someone who is no stranger to drama and controversy, and actively courts it:
https://www.thedailybeast.com/articles/2014/08/01/occupying-...
And like I said, you can follow the links and see the brigade discussing something else entirely.
Personally I don’t want an OSS ecosystem that banishes trans people or people who had weird proto-alt-right politics pre-Trump. If you’re gonna banish anyone, banish the ones posing existential risks to projects by their brigading against contributors they don’t like.
I am not part of the YCombinator or West Coast ecosystem, but didn't it banish gay people with weird pro-alt-right politics pre-Trump, and then supported Trump? Like, for some reason there was a movement to banish Peter Thiel: https://mashable.com/article/peter-thiel-y-combinator
How do you feel about banishing people who simply admitted to voting for Proposition 8: https://www.latimes.com/business/technology/la-fi-tn-mozilla...
If you take a look at a larger problem, you'll see that there is a lot of inconsistency with human welfare on a far larger scale.
Because this same system takes Saudi money a lot. The only moment of self-reflection came after one guy, Kashoggi, was killed: https://www.barrons.com/articles/saudi-arabia-tech-fundraisi...
But not the situation of millions of people in Yemen: https://news.un.org/en/story/2022/03/1113852 ... https://techcrunch.com/2023/04/01/andreessen-horowitz-is-now...
The US military industrial complex was largely involved in airstrikes on Yemen, as the Washington Post revealed last year: https://www.washingtonpost.com/investigations/interactive/20...
As a country, we ignore the Yemen war, and are told to only clutch pearls about taking money from Russia due to the Ukraine war. I imagine that YC stopped taking Yuri Milner's money a decade ago, partly because of his ties to the Kremlin, but probably it was just a natural parting of ways eventually: https://news.ycombinator.com/item?id=15631084
Anyway, just saying ... ecosystems aren't always perfect.
Similarly, this GH issue response to that occurrence, despite having made valid points in a reasoned manner, also held some of the same kind of "happen", in response - which is understandable, but not diffusive.
It ultimately doesn't matter who contributes, if someone truly believes in the project, they'll be just as happy to step away from it if they're affecting its momentum, even through no fault of their own or just a misunderstanding.
Momentum is important - and at an early stage like this, when vibe is building and community is forming, it can be very. I hope ggerganov continues to make these difficult decisions characteristic of clear leadership.
As for these LLaMA changes, I ran it on my machine for fun, and it worked perfectly. I wound up re-converting my models, but it doesn't take terribly long to do so even for 65B. After that, generation starts nearly instantaneously, which is very impressive. I wouldn't be surprised if there are legitimate problems with the change. Obviously people who deleted their local copy of the original model to save disk space are probably displeased, and maybe it is a massive performance reduction in some cases.
I wish I understood, and yet I fear I don't really want to know at the same time.
edit: At least in this case, it seems like it's mostly drama around attribution and unnecessary changes. Kind of sad that an otherwise really useful code change wound up being marred by probably-avoidable drama, but such is life ¯\_(ツ)_/¯ Honestly, I don't have any input, I just hope everyone can resolve their gripes amicably in due time.
> This PR was written in collaboration with @slaren. This PR is also rebased on PR #586 so please do not squash merge! Use either merge or rebase.
jart made sure to that the other user got credit, in addition to making sure that their name was properly attributed in the commit log. Given all this, it feels like the drama--shouldn't exist? Like, if there's an issue with attribution, it's not because of bad-faith, and I feel like a good-faith conversation could have just resolved this, instead of bringing in trolls.
This is the original PR: https://github.com/ggerganov/llama.cpp/pull/586.
Jart's archived comments:
"my changes"
"Here's how folks in the community have been reacting to my work."
"I just wrote a change that's going to let your LLaMA models load instantly..."
"I'm the author"
"Author here..."
"Tragedy of the commons...We're talking to a group of people who live inside scientific papers and jupyer notebooks."
"My change helps inference go faster."
"The point of my change..."
"I stated my change offered a 2x improvement in memory usage."
"I can only take credit for a 2x recrease in RAM usage."
"I just wrote a change that's going to let your LLaMA models load instantly, thanks to custom malloc() and the power of mmap()"
slaren replied to jart on HN asking her why she was doing and saying those things, and she didn't bother to reply to him, despite replying to others in that subthread within minutes. https://archive.ph/zCfiJ
This is BillG-style product skill -- there is a ton of work that goes into representing a piece of software as something important and valuable that people should buy into.
That being said, it's important to attribute work properly. It can be easy to mix things up (eg. "my patch" is excusable) but repeatedly insisting authorship when you're not the author of the change just seems disingenuous. I'm sure it was in good faith, but since they didn't address the issue or clear anything up, it's come to this.
Dramatic, and hardly the conclusion people wanted to the story of a free performance improvement. It's not entirely contrived though, and I think the maintainer handled this exceptionally well given the circumstances.
Is this? If she so easily misrepresented slarens work as hers in this case, what other work isn't actually attributable to jart?
In this specific instance, jart had a communication error that she failed to clarify, and so things compounded from there. The part that she didn't author is clearly defined in Git, and the most-plausible explanation is an honest mistake. Assuming ill-intent requires you to ignore the original context of the disagreement and focus on the outrage, which pretty much says it all.
That being said, I'd love to hear what evidence you have to the contrary. Maybe you've got a link to an FTP server from 2001 with the Blinkenlights source code on it, I can't say for sure. A fraud probably doesn't write in-depth patch breakdowns on their personal blog for fun, though.
I read that PR (didn't click any links) and here on HN posted a "Great work" to jart. The reason I did that is precisely because those final lines in the PR came across as an upright acknowledgement that some people helped out. I also got the impression that jart was a co-owner of the project with all the "we"s that were thrown around.
If I was writing that PR, it would be something like "this PR consolidates slaren's mmap approach with additional work done for ... by myself". After hearing about the drama, actually reading slaren's PR, and reviewing jart's comments in issues and the PR and the hn show and tell, I am now convinced this is someone who wants to steal other people's thunder. Heck, even this front page article is yet another PR stunt. I suspect "faster fork of llama.cpp" posts will follow.
Giorgi Gerganov remains for me the hacker hero here as far as LLMs are concerned -- mmap is kiddie stuff to be frank, but anyone who gets whisper and llama to work on my laptop with a handful of files (many thanks to you sir) has my technical respect. And I think he has made the right call regarding the project.
Related Work on this problem: 1. https://www.mongodb.com/blog/post/getting-storage-engines-re... - talks about developments on MongoDB's backend to use mmap. 2. https://www.pdl.cmu.edu/PDL-FTP/Database/p13-crotty.pdf - Talks about some of the cons of mmap, some I think are not as prevalent due to the existence of low latency, high throughput storage devices. 3. https://www.cs.cit.tum.de/fileadmin/w00cfj/dis/_my_direct_up... - less relevant but related.
The one thing to understand is that the performance implications of mmap are subtle and only work when you have much more RAM than the files you're mapping in.
Really depends on what you're doing, like memory access patterns. I've definitely seen scenarios when mapping hundreds of gigabytes of data on dozens of gigabytes of ram where mmap has been an almost absurd performance boost over traditional I/O, both immediately but also asymptotically as all the most frequently accessed data ends up in cache and the least accessed data is paged out.
I don't disagree with the subtlety part though. It's very difficult to reason about I/O performance in general. Modern systems are like an onion of hidden performance optimization tricks and caching layers (both in software and hardware).
Aren't all the weights touched in every pass?
Using mmap, you avoid doing any work at all the 2nd time you load the file.
I'll repeat myself: mmap is subtle. If what you mmap is larger than your host RAM, only some of the pages will be loaded at any time, and depending on access patterns, can lead to significant paging.
An advantage of copying in userspace is the ability to use more performant instructions to perform the memcopy, which the kernel does not typically have access to (https://www.mongodb.com/blog/post/getting-storage-engines-re...)
There is no copy with mmap, the page is either unwritable or CoW. There's always a copy with read(). (But read() can still be faster and more memory efficient nevertheless.)
> An advantage of copying in userspace is the ability to use more performant instructions to perform the memcopy, which the kernel does not typically have access to (https://www.mongodb.com/blog/post/getting-storage-engines-re...)
Darwin kernel does though.
I believe Linux uses the builtin old memcpy instructions on Intel, just to force CPU vendors to keep them usable.
You are right, if you are directly modifying the mmaped region. I always internally model my data as staging my changes to be synchronized to the mmaped region, so thats my mistake there.
> the page is either unwritable or CoW.
This is not universally true, or maybe I'm confused on this statement. MAP_SHARED exists, but maybe you are referencing a specific kernels' implementation on how they achieve coherence between file backed shared memory regions in two processes? Im not sure.
> Darwin kernel does though.
Sure we can always point to a kernel that has has implemented some feature or another, which is why I said typically you don't see it.
It does not. Compare the implementation of _bcopyout against _platform_memmove, you'll see the difference :)
That doesn't work in every kernel because they don't want to bother saving/restoring the extra registers.
So the discussions end up gravitating towards weird drama. I wish you wouldn't have linked this thread. Theres going to be a bunch of stupid comments here as well about how great/awful jart is.
Does the change deserve a blog post or wild claims like "llama.cpp is 100x faster and uses half the memory!"? No. The original PR looks like a decent addition but the blog posts reads as incredibly narcissistic (i.e. lots of language like "We spent several weeks volunteering" and "our project") uh whatever. It also breaks a backwards compatibility when there's no technical reason it couldn't have been optional or put behind a feature flag, plus a ton of condescending language in the PR. Not really the kind of work I'd be proud of or would be advertising in a blog post.
The claim that it uses half the memory was probably a honest mistake. The ensuing disappointment that it did not in fact halve memory usage and drama attracted trolls and white knights and is icky. The discussion around nmap I suppose is subtle and when emotion abounds can no longer be had. :/
Better than most stuff I see in the corporate world.
like, wow, mmap and paging. really guys?
I maybe should not be surprised, given that we live in the era of Unity and Electron, but using mmap() to load large files should be not be seen as rocket science.
And this is basically available on almost any platform with a MMU and a kernel.
Memory mapped files have their disadvantages. The biggest disadvantage is that any disk read error (or yanking the USB drive) becomes an access violation exception (also known as a crash), just like you read from a bad pointer. You need to have robust exception handling, which is a taller order than just checking a return value.
Another disadvantage is that even when you have your pages mapped into memory, calling the page fault handler and getting your page has a cost of ~1200 CPU cycles on Windows just to do the User<->Kernel mode transition, plus the cost of actually performing the IO. "Just reading the file" skips many User<->Kernel mode transitions, so it's one per read call rather than one per page fault.
On the other hand, its only been a few weeks, so maybe I should ignore this absurdity and just wait.
mmap isn't relevant to anyone except CPU-using programmers because other hardware doesn't have virtual memory paging. Firmware programmers don't care, GPU programmers don't care.
I think the better read is that they're being adapted to new applications, constraints, and environments, all at once.
You have too much faith in unis. Mine did not teach me about mmap at all.
I say this to highlight the parent comment. I'm essentially in a computer science program and we have learned absolutely 0 about paging or memory in any of my required courses. We practically don't touch OS anything in any of the classes. That's not to say the courses for that aren't offered but they aren't part of the core curriculum and over my time in my program, they've mostly not been offered due to lack of student interest.
I did learn how to use linked lists like a champion though!
We've had regular discussions on HN about various storage engines, how the latencies are cut down, etc. I share your surprise at hearing 'wow, mmap!' and all the debates in the issues as what it actually does.
https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
Self respecting computer engineering curriculums will cover MMUs, page tables, TLBs, hardware interrupts, and page caches which once you know about mmap is fairly simple to understand.
The fundamentals really haven’t changed much in the past 40 years.
Suggesting that MMAP limits things to the size of the ram, well - no, as well - paging may happen, but then we are just back out to the file.
Honestly, some of the weird assertions wouldn't take long for people to double check an verify (or falsify).
Edit: can someone running llama.cpp ask it whether it thinks it's a good idea to concatenate a running list of vanity initials of important developers into a magic filetype constant?
> Regarding the version comment - yes, the plan was to bump versions and no the magic. But I'm ok to change the magic to commemorate the significance of this update. In fact, maybe we can make this a thing and everybody who makes a significant contribution to the project will get their initials appended to the version. What do you think? smile
[1] The kernel doesn't necessarily load the whole thing into page cache at once and keep it around indefinitely. It might have been recognizing a sequential loading pattern before and basically discarding pages almost immediately, where as now it might be keeping them for much longer. Or it might now be essentially skipping loading the whole thing in at once and doing it page-by-page on demand, which could be more RAM-efficient but slower. To some extent, you can control these behaviors with madvise, mlock, MAP_LOCKED, MAP_POPULATE, as well as various sysctls. Also, if it had to page out before, the anonymous memory was "dirty" and thus had to be swapped (written out to disk) where as the mmap()ed bytes are "clean" and can simply be discarded and (if needed to be paged back in later) reread from the existing file unchanged.
fwiw, I'm not a ML person, but it doesn't seem entirely crazy to me to think that SSDs are becoming fast enough that you could avoid keeping a huge model in RAM in some cases. Especially if "computational SSDs" (SSDs that can do some basic first-stage computation without transferring the input data over PCIe) ever become common. (I think some of the ML accelerators for sale today might be approximately this.)
I made an SSD into a spare swap device, and basically treated my system as having RAM+SSD's worth of RAM. It allowed me to finish a few big jobs (~96GB RAM) overnight that wouldn't have otherwise.
Using filebacked pages instead of anonymous memory is a real improvement because it doesn't have to get swapped out if there's memory pressure. And this program probably isn't the only thing running on the machine.
This model is really good for a (semi-)open source model. I think this may be the first locally runnable model that I will actually use for real stuff rather than just play around for fun.
It's not ChatGPT level but it's not that far behind. It will draw ASCII art HUDs for a text adventure or analyze data or recognize languages or write stories. AFAIK it's been trained on ChatGPT discussions so makes sense.
This AI still gets uppity sometimes about offensive content but unlike ChatGPT, you can edit the prompts to put words in its mouth to encourage it to answer properly.
I only got it working at all yesterday and there's no nice UX at all. Not sure I recommend trying to use this as llama.cpp will probably have this in no time with a much better user experience, although I am also trying to make it more usable.
If you follow the instructions on Vicuna page over how to apply the deltas, and you can compile the project, then you could run:
cargo run --release --features opencl -- --model-path /models/vicuna13b --param-path /models/vicuna13b/config.json --tokenizer-path /models/vicuna13b/tokenizer.model --prompt-file prompt --top-p 1.0 --top-k 20 --repetition-penalty 1 --temperature 0.9 --max-seq-len 2048 --f16 --percentage-to-gpu 0.9
Where /models/vicuna13b is the HuggingFace-compatible model. This will put 90% of weights on GPU and remaining 10% non CPU which is just barely enough to not run out of GPU memory (on a 24 gig card)
Create a text file 'prompt' with the prompt. I've been using this template:
You are a helpful and precise assistant for checking the quality of the answer.###Human: Can you explain nuclear power to me?###Assistant:
(the model seems to use ### as delimiters to distinguish Human and Assistant). The "system prompt" is whatever text is written at the beginning.
Vicuna: An open-source chatbot impressing GPT-4 with 90% ChatGPT quality - https://news.ycombinator.com/item?id=35378683 - March 2023 (167 comments)
> We release Vicuna weights as delta weights to comply with the LLaMA model license. You can add our delta to the original LLaMA weights to obtain the Vicuna weights.
So you get the LLaMA weights (somewhere), then apply the Vicuna deltas to them to end up with the Vicuna model.
The latest so far would be Vicuna, whose weights were just recently release.
I'm not sure what this means but I'm pretty sure I can name several "high level libraries" that mmap things. None of those are the STL, but it's not exactly perfect design.
mmap sits at this lovely intersection between virtual memory and the disk, and it’s been around for a long time. By now there are other means of playing within that nice intersection, but mmap is the pop classic.
It's already extremely high level from a certain perspective :)
If you send it over IPC it's nice to keep it mmapped instead of accidentally copying it too.
No wonder a lot of people see it as "mmap, nothing new". But this is not the case in a lot of libraries where the norm is to just budget for a lot of time moving things to/from gpu and just relying on someone else's code.
Instead of accumulating technical debt the owner of the repo decided to merge this, have a breaking change and move on. When there was some community backlash to the breaking changes there was a pull request trying to revert all changes instead working through the issues (it was a net win for several users but not all, some configurations with slower drives were better served by the older approach). There was an ugly back and forth and the repo owner decided to ban both the person who did the pull request and the author of this post.
This article brings the conversation back to the technical merits, the roadmap, credit the the original authors and tones down the ownership tone that may have pissed off some community members.
That pull request has now been closed by the owner of the repo. They are trying to move on and be productive, let's do the same.
Is this AI people learning that mmap exists? People simping cos' there a Justine involved?
I think `dd` in conjunction with the `oflag=direct` has this functionality. See: https://stackoverflow.com/questions/33485108/why-is-dd-with-...
Wasn’t it discovered last week that loading larger models was an error in measurement and the speed up was from keeping things in memory after the first loading?
Please do correct me if I’m wrong.
> The first time you load a model after rebooting your computer, it's still going to go slow, because it has to load the weights from disk. However each time it's loaded afterwards, it should be fast (at least until memory pressure causes your file cache to be evicted).
* 2x memory. A 20G data set requires 40G (20 for page cache and 20 for LLaMA)
* Things would be _even slower_ if they weren't in page cache after first loading. mmap is fast because it does not require a copy and reduces the working set size
And they're arguing about mmap, something that's been around forever.
It reminds me of that speedup in a package manager because they were reading uncached byte-at-a-time off of disk. You need to explicitly turn buffered reads off...but why would you do that in the first place? Unbuffered reads are almost never a good idea, ever.
It makes me wonder what other weird sub-optimal stuff is lying underneath the resource behemoth that is ML.
This is just like any other technology. Use it wrong, you will get burned. Doesn't matter how shiny the container is.
"ML stuff" also (mostly?) includes purely statistical methods from the 90s that are deterministic and arguably the best way to solve a large variety of non-generative problems.
In fact, unless generation of arbitrary output is a major objective, it's likely you can solve whatever ML task on a workstation from 2010 that uses intel integrated graphics.
> Currently THP only works for anonymous memory mappings and tmpfs/shmem.
$ cat /proc/<firefox process>/smaps
[...]
7efca62e3000-7efcaa13d000 r-xp 00ae3000 00:18 75354786 /usr/lib/libLLVM-15.so
Size: 63848 kB
KernelPageSize: 4 kB
MMUPageSize: 4 kB
Rss: 58420 kB
Pss: 19213 kB
Pss_Dirty: 0 kB
Shared_Clean: 56372 kB
Shared_Dirty: 0 kB
Private_Clean: 2048 kB
Private_Dirty: 0 kB
Referenced: 58420 kB
Anonymous: 0 kB
LazyFree: 0 kB
AnonHugePages: 0 kB
ShmemPmdMapped: 0 kB
FilePmdMapped: 57344 kB <==== huge pages
Shared_Hugetlb: 0 kB
Private_Hugetlb: 0 kB
Swap: 0 kB
SwapPss: 0 kB
Locked: 0 kB
THPeligible: 1
VmFlags: rd ex mr mw me sdI see https://docs.kernel.org/filesystems/proc.html describes FilePmdMapped as "Page cache mapped into userspace with huge pages", consistent with what you are saying. I don't fully understand the distinction between that and FileHugePages: "Memory used for filesystem data (page cache) allocated with huge pages". I wouldn't think it'd be possible to map it into userspace as huge pages if the kernel hasn't allocated it as contiguous physical memory (and consistently aligned with the userspace virtual addresses), so there's something I'm missing.
What kernel version did that output come from? Do you happen to know if Firefox did anything special to set that up? What filesystem type is this?
I also see something about MADV_COLLAPSE that supposedly supports file-backed pages. [1]
$ zgrep 'CONFIG_READ_ONLY_THP_FOR_FS' /proc/config.gz
CONFIG_READ_ONLY_THP_FOR_FS=y
> What kernel version did that output come from?kernel 6.2.8-arch1-1
> What filesystem type is this?
btrfs
> I don't fully understand the distinction between that and FileHugePages: "Memory used for filesystem data (page cache) allocated with huge pages".
Probably pages in the page cache that aren't mapped into a user process. Which is what happens when you read()
I didn't find any documentation that would indicate that explicit huge pages didn't work with on-disk filesystems, but sure enough, it doesn't seem to work on ext4.
Maybe sarcastic, but it is how things look like today.
The trick she did overriding malloc & friends to validate that the optimization would be worth doing is, in my mind, one of the high-points of the paper. It's a very clever way of making a meaningful measurement, which was the keystone of the entire change.
I've never heard of, thought of, or used that trick, and the fact she had it in her arsenal to apply to this very specific situation is pretty impressive, to me at least.
As you can see from comments up this thread (ref. various GitHub issues) - the people "gluing mmap" don't actually have a single clue. They can't properly measure memory consumption (they don't understand what the numbers they're seeing actually mean). They don't understand how paging, swapping or virtual memory work. They don't actually understand the concept of memory-mapped files, why they're there and how they work. They can't explain why their code behaves differently when using memory-mapped files.
Moments like this are here to remind you that there's actual knowledge and skill to building scalable and efficient software, and that hustling and copy-pasting StackOverflow examples will only get you so-far, as will "piecing together dataframe pipelines" in Python.
Hopefully we should publicize this kind of achievement as a way to teach more devs about mmap... (this really should be common knowledge)