llama-fs: A self-organizing file system with llama 3
github.com
github.com
My first question was why would this be a file system rather than an app or a script, but I see it's actually an Electron app and python scripts, which I think is the right approach.
I think that something that would have a UI for automatically tagging, renaming, and moving files on request but not constantly would be very handy. Also, if you could somehow steer files into binning directories ("if X, put it in bin ~/X, if Y, put the file in bin ~/Y", "if it's an invoice or deals with payment, put it in ~/Documents/Finance"), that would be cool. Finally, Windows support would be amazing.
Key point here -- I might be old, but I like my files to stay where/how I left them.
Usually I just want the LLM to extract the document title and unique information (document numbers, companies mentioned on the document, people mentioned in the document, dates, and similar).
The entire workflow is this:
1. Prepare a batch of documents with Patch-T codes and load them into an automatic feed scanner.
2. Scan documents with NAPS2 and OCR, producing PDF files as they are scanned.
3. My python script monitors the output directory, picks up batches of files at a time, and processes them sequentially. The first page of each document is loaded as text, the text is fed into llama2 (the last one I used) through Ollama with a few-shot prompt, roughly in-context trained for my sort of documents, it outputs a file name.
4. The file is renamed and moved to a processed directory.
5. Optionally, I will also run a second step to generate tags.
I do not organize my every day files this way, it is specifically for digitizing and making redundant archival copies of old documents, which I collect.
This is a bespoke pipeline that fits very well for my workflow. The renaming is very accurate after I have perfected the context that I pass to Ollama, and the final files are easily searchable on slow-read media. I have not yet added any kind of computer vision because I scan so few photographs that I can label them myself. I agree with you that naming is personal, but I think that we can train the LLMs to name things the way we expect them to be named. And at some point, the human cost of labelling large amounts of data may make certain workloads simply impossible.
Interestingly, my pipeline is very slow. It runs on an old GPU and takes about 30 seconds to produce just the few tokens to name a file, and about the same for the tags. I will probably move it to an even slower pipeline on a Mac mini M1 when I have the time, just to save on electricity. Because ultimately, it finishes in a few hours and it doesn't take up any of my time. What would be a full-time labelling job for a human is now a full-time job for a machine, which makes the archival hobby feasible and cheap.
Could also have integrity checks that total number of files and their attributes didn't change after the commit
Cool project OP
I think one thing to improve the readme or landing page for this project would be a before & after for a sample ~/Downloads directory, maybe in `tree` format.
I am a grumpy AI hater. But Llama is not the security/data risk here. I don't think anyone should use this unless they are interested in contributing.
Most stuff like that doesn’t need an LLM and would probably decrease utility.
In my case, I built a local OpenAI-emulated proxy API to run against a local LLM, and used a modified OpenAI library to connect with it. This was the solution a year ago. Now, it's easier to deploy a local LLM.
Haven’t played with its tab organization, though.
We built this for the Llama3 Cerebral Valley Hackathon. The idea was this: My ~/Downloads folder is extremely messy, and I wanted an agent to fix it for me. So we built one.
LlamaFS reads file contents and metadata to Ollama with LlamaIndex, Moondream, and Whisper to understand what it’s about. It renames the files according to a specified pattern and organizes similar files into directories. Also, LlamaFS doesn’t overwrite anything until you explicitly ask it to (no risk of deleting anything precious)
One cool thing we benchmarked with AgentOps was the speed. It goes pretty fast, ~500ms per file.
In watch mode, LlamaFS starts a daemon that watches your directory. It intercepts all filesystem operations, updates i and uses your most recent edits in context to proactively learn and how, so you don't learns predict how you rename file. e.g. if you create a folder for 2023 tax documents, and start moving 1-3 file in it, LlamaFS will automatically creates, and move the right!”
Come again? Was AI used to generate the documentation?
Arc Browser has an AI rename feature (for downloaded files). I tried it out but I had to turn it off. I love Arc Browser BTW and their AI hover summary is useful. I found poorly naming of files to be disruptive and it's a lot better if I am more involved in renaming the file- that will help me remember it.
it's not "self organising" in this sense, but it's an easy way (imo) to organise files across your desktop.
As a data archivist, I would definitely recommend a setting to turn off rename of files as that can often be a database id, timestamp, etc.
Would be cool if I can build CI pipelines for daily stuff, just describe everything in YAML* and not have to do repetitive tasks all the time.
* or, hopefully something better that isn't a pain to write
With virtual systems, I think it could be really interesting, you could have a few different types, from conceptual, to project, to research area. That would be amazingly cool.
I like this idea.