Meanwhile if I use Notepad++ and load the same file windows shows it is using 586MB.
Any ideas why such high memory usage if this is a direct port to QT?
Meanwhile if I use Notepad++ and load the same file windows shows it is using 586MB.
Any ideas why such high memory usage if this is a direct port to QT?
~2x memory looks like a naive implementation of just allocating an std::string from heap for each line. Due to heap fragmentation and various overhead it would quickly blow up.
~1x memory looks like just reading the entire file into RAM (that would still be slow).
A truly efficient implementation would never need to load the entire original file in RAM. It would just need to remember the binary offset of each line in a way that combines random access and reasonably fast insertion/deletion (e.g. K-fold trees). You can even keep everything beyond the top-level directory in an temporary on-disk file, so your RAM usage could be less than 1MB with nearly instant performance.
The efficient implementation is tricky, error-prone, and involves handling a solid amount of corner cases, which is beyond the amount of hassle a typical hobbyist developer is willing to go through.
I think this is mostly likely the culprit. There almost definitely isn't that much contiguous memory available for a large file like this, so there are a lot of wasted pages (maybe 2-3x) which is causing that footprint to balloon.
Vim's implementation also feels somewhat line-oriented in that if you load a large JSON file that has everything on 1 line and you try to edit that line, uh... it will be sluggish, but if you do, like %!jq '.' and format it, you can them move around the file a little easier.
Edit - What the heck while I'm at it:
Emacs - Historically used a Gap-Buffer which optimized more for locality of edits than loading large files. It doesn't suffer though as much from heap fragmentation though as the linked-list of strings type approach.
Monaco - The editor in VS Code recently went from a linked-list of strings to a PieceTree buffer. Which basically loads a the whole file into an immutable buffer and then uses the PieceTree to manage the edits.
Others - Lately people keep talking about Ropes, which is another tree type deal with extra-smarts specifically for text editing. I don't know of a game-changing editor like the others mentioned that uses it tho.
An example for an editor that uses them to manage large text files is Xi-Editor: https://github.com/xi-editor/xi-editor (edit: written in Rust).
One downside, relative to this conversation, is it doesn't make working with pathologically large files a lot better. Without further optimization a gap buffer could push your box into swap if you open a large enough file.
I could think of a reasonable implementation for Linux that would use mmap and fallocate, but it wouldn't "just" be an mmap, as the swap file would still need to represent a rope or a gap buffer or something else of the sort for efficient editing.
It's frustrating but users don't understand how to read memory usage at all. There's even guides out there quite wrongly telling users to look at virtual memory usage for each app.
Mmap is a brilliant way to have the os handle which parts of the file actually reside in memory but you'll seriously have to deal with users that know only enough to be dangerous complaining that your app uses gb of ram when it's merely mapping a file into virtual address space to allow the os to page as it sees fit.
In fact I wouldn't at all be surprised if the memory measurements in this thread were judging virtual allocation rather than physical memory used.
It is a lot more performant than VS Code, but VS Code seems to have won on the sheer quantity of “good enough” plugins. Its pair programming features are a big deal as well.
For the most part I use VIM for quick things and VS Code when I need a plugin or to work with others. SBT kinda sits between those two extremes.
The pair programming feature from VSC is the only one I’d really miss from that side if I were to go back to SBT. Not sure about from the VIM side because I’m sure the SBT VIM plugins have improved over the past 8 years.
To know the offset of each line, you'll need to load the entire file from disk. So you're talking about loading it, processing it, and throwing it out again. You should qualify calling that "efficient" since you will need to re-load portions of the file, as needed.
Also, an OS's virtual memory systems may work well for you with minor tuning, so it's work looking into that before you spend a lot of time essentially writing your own.
I think unless you have a good reason not to, your best bet for a text editor on a modern OS is to read the entire file into memory and leave it there. (But don't make multiple copies of it or use two bytes to store each character as the current version of NotepadNext is apparently doing.)
No. You will need to read the file, not load it all at once in memory. And for a sequential scan, you need very little memory at any one time.
And if you are doing that, you can start displaying the file as soon as you have scanned enough to fill the view. The rest could happen in the background.
The problem with that for a text editor is that probably the first thing you want to do is process the line endings. So you’re just going to access the entire file, right away, as fast as possible anyway.
So mapping doesn’t really give an advantage in this case.
Except that this would most likely result in unpredictable corruption if another process modified the file while you had it open, which is contrary to the way basically every text editor works.
It would be nice if there was a way to mmap a file in such a way that the OS would give you a snapshot view, but AFAIK that doesn't exist, at least on Linux in a filesystem-agnostic way (not sure about other operating systems).
On Windows in particular you also have a ton of other facilities that could help with this (opportunistic locks, transactions, etc.) but notification is usually good enough.
My text editor supports files up to 256 GB in size, unless you have that much memory, you’ll have to wait a little bit to go from the first to the last line.
I think I'd be a bit unhappy if I was regularly loading megabyte blobs of text into editors, and start looking at other representations or tools.
The only files that big that I know of are log files, and very occasionally a CSV or two as some kind of data dump, and there's usually dedicated tools for those cases.
We wouldn't use large quantities of RAM today if there were any techniques that could make disk acces "nearly instant".
It's worth keeping in mind that typical CPUs from the past few years to today have memory bandwidths of several GB/s. Disks have sequential read speeds of ~100MB/s; SSDs can easily do 1GB/s and above.
Is there any open source text editor doing this?
Looks more line Qt's internal UTF-16 encoding in this case.
Sure, I didn't mean to imply that the whole document would be one giant QString. But those data structures might still use a QString as backing memory in the new implementation. I tried to have a look around, but didn't have too much time on hand to dive deep. I could see some QString usage but couldn't confirm if the document itself uses it for storage.
> Wouldn't Notepad++ use UTF-16 internally too, with its Windows heritage
Not necessarily, as the memory consumption of 586MB for a 500MB file shows.
Definitely my goto editor for non-code files. I've wondered if emacs can do the same, but when I've asked an emacs user responded with "Why would you want to open 800 files?" I take that as a "no"?
Emacs users aren't accustomed to getting questions they don't know the answer to, IMHO.
Emacs handles it fine. Two years ago¹, I wrote this:
“Just last week I opened, edited, and saved, a 4 Gigabyte SQL dump file, in Emacs. It was a bit slower than usual, but still well usable, and certainly no crashes.”