Personal computing paves the way for personal library science
bramadams.dev
bramadams.dev
It has been extremely useful, especially because it's text-searchable and the really important papers are properly categorized.
A local LLM will make it 100x more useful. Also, it might not even need be "local." If I make it available via the web, I can probably sell access to other scientists and engineers in my field.
Recent advances really benefit data hoarders out there.
I'd add that these days it totally makes sense to download libgen's entire archive, because (1) storage has never been cheaper, and (2) you can use it to train local LLMs.
Out of curiosity. Does this statement come from complete ignorance of or complete disregard to copyright of the author?
While I rather like the idea of having to provide access to derivates of publicly funded works, I fear that people would rather not use it than invest money into innovative approaches of using it. Of course if the public pays for the training and development costs, then by all means it should be available.
And the library itself and the computing resources to operate it cost money that someone needs to pay. Publishers didn't pay for the research and yet they can profit from it - why this guy shouldn't?
A question that could be appended to many LLM discussions!
Of course, the idea of selling access to your private digital collection (or a derivative model) is relatively more absurd than the idea of monopolizing the original published work... Even so, this is as good a time as any to reconsider the practicality of copyright.
Hardly. Data hoarding comes with most downsides of hoarding physical objects. It's just smaller, cheaper & easier to process.
There's people that get rid of any physical object they haven't used in the past year (or 2, or 5, whatever). This makes sense. Imho, every object you own falls in 3 categories:
1) Things you use on a (semi?) regular basis. They make your life easier/nicer.
2) Things that are valuable. For flexible metrics of what constitutes value (sentimental, nostalgia, monetary, insurance against 'disaster', ...)
Not having used something in a long time is a good hint it's not valuable.
3) Luggage. Whose value is negative. It doesn't provide anything, just takes up space, drains mental energy (and possibly other resources), and in doing so gets in the way of other pursuits.
Data is no different. For any single piece of it, you either use it from time to time, somehow derive value from it, or it is useless luggage that you drag around at a cost.
Apply good judgement in what to hoard.
the new external drive is so much larger than the oldest ones i put the new stuff on it and use the rest of the space to backup old drives.
friends are constantly purging stuff but they seem unaware how much time and effort it takes.
the significant other wanted to clean up her old photos rather than upgrade the icloud. its an insane amount of work?
So the same downsides, except ... much better?
> There's people that get rid of any physical object they haven't used in the past year (or 2, or 5, whatever)
> Data is no different. ... you either use it from time to time
I suppose the real question then is what timeframe do you consider "using it from time to time"? It seems to me that this depends very much on the person, and likely on the objects themselves - probably a large appliance you haven't used in a year is likelier to be thrown away than a small item you haven't used in several. Considering data storage is smaller, cheaper, and easier to manage, I suppose a reasonable timeframe for keeping it, just in case, would be a lot longer than physical items.
This of course says nothing of the societal and cultural value such archivists safeguard.
That's not legal. Just because you own some files does not mean that you own the IP of the content within.
Well said. All anyone can do is to do the lonely work till you can't anymore or you find friends to not be lonely at that work anymore.
There is a rush in public to condense and summarize many authoritative publications to find patterns, or to replace a human expert with automated results.. yet that is fundamentally different than taking multiple incomplete perspectives to add to a human library-owners knowledge and investigations.
It is subtle to speak it but not subtle in its implications.. taking "data as facts" and condensing them or reordering them or rewriting an output based on them, using automation, is different than a human mind taking in many inputs for human mind knowledge and enabling new outputs from a human author.
You nailed it! Thanks for noticing the divergence!
> The focus of research at BCL was systems theory and specifically the area of self-organizing systems, bionics, and bio-inspired computing; that is, analyzing, formalizing, and implementing biological processes using computers. BCL was inspired by the ideas of Warren McCulloch and the Macy Conferences, as well as many other thinkers in the field of cybernetics.
On cybernetics, https://www.pangaro.com/definition-cybernetics.html
> Artificial Intelligence (AI) grew from a desire to make computers smart, whether smart like humans or just smart in some other way. Cybernetics grew from a desire to understand and build systems that can achieve goals.. it connects control (actions taken in hope of achieving goals) with communication (connection and information flow between the actor and the environment).. Later, Gordon Pask offered conversation as the core interaction of systems that have goals.
Happy to answer any questions
In this respect local LLM's are simply the tip of the iceberg, pointing out the vast amount of personal information processing that is available in principle but does not actually happen.
Thanks to Linux being used at scale in Android and WSL, it's now maintained and capable on the desktop, as a hypothetical foundation for personal computing innovation. But even there, native GUI toolkits took a backseat to web and CLI. Remember Chandler? http://www.osafoundation.org/
Investors poured small fortunes into cauldrons of smart devices, wearables and AR/VR, with little to show as nascent ecosystems failed to achieve escape velocity, due to closed hardware and software that forestalled the experimentation which birthed personal computing.
Apple Silicon has reinvigorated walled laptops. Hopefully next month's derivative Qualcomm SoC from PC OEMs can offer good price/performance/watt for Apple-competitive-yet-open Arm laptops and tablets that can run any Linux distro, with retail SSDs and RAM, plus AI silicon roadmap.
A modular Framework Arm laptop would be a good start to rebooting PC innovation.
If Arm SystemReady laptops with good performance/watt have an open security foundation (declarative, immutable OS at EL2) to support multiple competing "app store" equivalents on Linux, the resulting revenue and competitive market can reward innovative desktop software - open, closed or hybrid. Without an Apple tax on storage and memory, funds can be redirected to a competitive market of smaller ISVs.
Outside of Steam, not a single software distributor (including Canonical) has been able to do this. Linux succeeds in spite of everything you mentioned and none of it would particularly enable the sort of experience you're describing.
So it's easy to make a device powered by one when you control how linux is booted on it, but for whatever reason things like EFI NVRAM interface on windows qualcomm powered laptops was done non-standard and the only reason windows works is because there are drivers shipped which work around it - and I seriously doubt its intended by Microsoft, because Microsoft actually benefits from devices following their official, documented, hardware-interface specs - it makes for easy upgrades, reinstalls, etc. etc.
Would an LLM-driven "Personal Library" require manually annotated textual interpretation of each curated item, or could it derive personal interpretations from user history and the uniqueness of curated items/sets?
For those who have been using local, offline LLMs with a manually curated text/image corpus, what have been the most valuable or surprising use cases?
Author demo video (2023), https://youtube.com/watch?v=7TgqMRz2r3M & tooling comment (2024), https://news.ycombinator.com/item?id=39789712
> Inspired by the commonplace book format, I take highlights from Kindle and embed them in a DB. From there I build (multiple) downstream apps but the central one, Commonplace Bot is a bot that serves as a retrieval and transformer for said highlights.
No. In something like this you’d probably have the LLM annotate and curate your personal library for you.
Potentially by creating and assigning tags or topics based on the content of your library.
The PKM is the stored info to write to and query against (both for LLMs and humans). The data ingests are just a pipeline of digital inputs to the system, like chat logs, maybe (transcribed) webcam feeds, files i'm currently editing on desktop, browsing history, etc. The current world view is the interpretation of what i'm doing - to tie all the ingests together and give them context. Eg in isolation browsing some Rust crates might not be that useful. But if i'm also editing Project X on my computer then it's reasonable to assume the searching is related to X. However if it's been 8 hours since any Project X activity, it's less likely related. Same goes for context-less chat logs (as happens frequently in my house) where they are extensions of a voice conversation, etc.
All of this stuff is of course insanely privacy invading, so i'd only implement this locally. I also wouldn't even store most of it for fear of data invasion, but using it to fuel a PKM automatically seems pretty sexy. Like browser history, but for your life.
This is all just wishful thinking though, LLMs have been moving too fast for me to even bother toying with this. I should note though that i did not intend for LLMs to be "smart". Rather, in a RAG-like fashion (i think is the term), i want to just let LLMs do what they're good at - summarization & autocomplete, and let the world view / PKM store the real data.
I’ve personally found that tagging is less robust than LLM embeddings (mainly due to dimensionality), but human appended thoughts about a source — also embedded — serve even better as tags.
Example: “this is a quote about dinosaurs…” (Old way of doing things) Tags: dinosaurs, jurassic, history Query: “dinosaurs” > results = 1…
(New way of doing things) Embedded Quote: [0.182…] User Added Thought: “this dinosaur reminds me of a time i went to six flags with my cousins and…” Embedded User Added Thought: [0.284…]
Query: “dinosaurs” > results = 2 (indexes = sources, thoughts)
The "thoughts" index can do a second layer cosine similarity search and serve as a tag on its own to fetch similar concepts. Basically a tree search created by similarity from user input/feedback loops.
My final answer was to use the Library of Congress catalog system. They need to add some sub-categories for how-to explanations.
Then have a field for media type (video vs. PDF vs. image)
Then note the style of presentation (academic vs. folksy vs. a manual vs. a dad showing you how to do this)
Then note the language
You can see it on the Wayback Machine:
https://web.archive.org/web/20211127090321/https://wiki.shap...
In retrospect, I should have put some of that effort into:
https://en.wikibooks.org/wiki/Hobbyist_CNC_Machining
although since then, a machine owner worked up:
https://shapeokoenthusiasts.gitbook.io/shapeoko-cnc-a-to-z
I still regret a bunch of stuff I didn't keep copies of, esp. the scans of Barry Hughart's notes for his novels.
The irony is that one can see a bit of the result of discussion of this sort of thing at the top of one's browser window --- the URL bar, where URL == "Uniform Resource Locator" --- the originally proposed term was "Universal Resource Locator", but the argument against that was that people were not librarians, and that unlike Ted Nelson's Xanadu, there wouldn't an over-arching data structure and organization, so a given document wouldn't have a single canonical location.
Anyone interested in this sort of thing who hasn't read it, should read Tim Berner-Lee's book:
https://tug.org/TUGboat/Articles/tb24-2/tb77adams.pdf
(basically used copies of _The Bible_ and The Works of Shakespeare to determine if a given set of letters appeared in the English language or no)