Memex is already here, it’s just not evenly distributed (2020)
filiph.net
filiph.net
Annotation - The ability to mark up hypertext... yeah... HTML claims to do it, but you can't change a document once it's read only. You should be able to layer annotations on top of a document. HTML simply is not a Language for the Markup of Hypertext Documents.
[Edit/Extend] - Annotation is anything you could do with a printout, at minimum. Circle something, highlight a section, black something out, attach links, etc. It's not just tagging a file as a black box.
Trails - the ability to store a trail of annotations, linking up parts of two documents together, or more. Like a stack of bookmarks, except bookmarks can't anchor to part of a document.
Publication - One of the crucial features of the Memex was the ability to output a collection to allow sharing with others. It could spit out a copy of everything linked to a trail, all the files, all in one dump. Copyright holders would have a huge problem with Memex if it were implemented as originally shown.
I think it should be possible to build a proxy that records ALL local http(s) traffic, and thus record all the sources to documents you've seen... and then only keep those that you've tagged or linked to after a few months, purging the rest. A private Memex, but it couldn't even be an open source project, as copyright holders would work to squash it.
You can annotate any upload (arbitrary data, files, or collections of files) with tags and since the data is permanent, your links don't break.
Publication is possible too, and with the Universal Data License, copyright holders can define the terms for their part of the publication right at upload time.
Apple has lots of software that only exists to sell hardware. This seems compatible with that business model to me.
There was a Microsoft Research project in the early 00’s called “Stuff I’ve Seen” that did basically this. I don’t know why it never really went anywhere - I’ve wanted one ever since I saw the talk.
Edit: Link - https://www.microsoft.com/en-us/research/publication/stuff-i...
It's weird — we, the builders of HTTP-based APIs, very well understand the concept of "REpresentational State Transfer" ...but I've never seen a web-app server that actually accepts PUT or PATCH or POST of HTML Content-Type against the baked-out server-rendered HTML views, such that if you were to push a modified version of said representational state (the post-edit HTML), it would get "decompiled" on the backend into a delta of the internal state (e.g. Markdown in a CMS), and that change to the internal state would get persisted, reflecting back as a change in the representational state when the page is fetched.
If anyone would just actually implement that on a web-app backend, then you'd be able to "annotate, edit, and extend" HTML-presented content using freakin' NSCA Mosaic from 1993. Or the "edit in Frontpage" feature of IE4.
(No modern equivalents, though — the features that would generate a "native" browser PUT or PATCH against an HTML page have since shrivelled up and died, after it became clear that no server would ever bother to allow them.)
To dodge the chicken-and-egg problem of browser support, you could do this today, with a Javscript SPA serving the browser role: the frontend app would ensure that it's reflecting all internal state changes in its in-memory client-side internal state into changes to the HTML (even if those changes don't cause any presentation changes on the client side); and then, where a regular SPA would push those changes by calling an API endpoint that accepts a version of its in-memory client-side internal state, this SPA would just push the updated HTML, as a PUT or PATCH to the very same URL that was loaded to deliver the SPA.
I think what the world has needed basically since then is browser integration for something like annotations, so that web sites and commentary about them can live together. And with less of a direct line to advertising dollars. Perhaps something that could live in the Fediverse. Sometimes I want HN’s opinion and sometimes I want someone else’s.
Using a more general architecture, it was possible to add many kinds of structures on top of existing Web pages collaboratively. Navigational or spatial hypermedia, guided tours, etc. Link bases were stored outside of the linked resources, so that you could have different sets of structures applied to a corpus of documents.
There were also quite a few commercial attempts in this space. However, all, commercial and research alike, ultimately failed, and I believe that there were several reasons behind this: Web site owners disliked that others could annotate their pages; shared link bases often became polluted; bookmarks were ‘good enough’ for most users; maintaining link anchor consistency in (increasingly dynamic) Web pages became more and more difficult (despite our best efforts); and while I remain enthused about the notion of the MEMEX scholar carefully annotating their corpus of literature, I don’t think the appeal is all that widespread.
Still, there are remnants in some systems – e.g., I enjoy to see that others have highlighted the same passage as I in a Kindle book.
The author and the old guy in the video he linked to behave almost cult-like, especially the old guy: He literally claims that this IS the best method for working with documents ever, the www is a fork of his idea based on a "dumbed-down in the 70s at brown university", he does not understand why it has not already taken off and he thinks its the most important feature for the human race. Really?
If people really see potential in this and work in spaces like journalism and academic research there would be already big programs out there.
Yes it is good to have passion about something and yes it is good if someone has a real need for this and his delivered with a solution, but this will never go mainstream. And in my opinion not even in the segment of technical skilled people like engineers.
This is the typical invention that fits the "I know it is the best thing, I love it and almost pressure people to use it, but it has not taken off the lightest for decades" case.
There have been "big programs" but when the web came, fundamental hypertext research and development on other systems came to a grinding halt. Ted Nelson, and many other researchers, predicted many of the problems that we now face with the Web, notably broken links, copyright and payment as well as usability/user interface issues.
I don't know what an average user is, but what a user typically does or wants to do with a computer is somewhat (pre)determined by its design. Computer systems have, for better or worse, strong influence on what we consider as practical, what we think we need and even what we consider as possible. (Programming languages have a similar effect).
One of the key points of Ted Nelson's research is that much of the writing process is re-arranging, or recombining, individual pieces (text, images, ...) into a bigger whole. In some sense, hypertext provides support for fine-grained modularized writing. It provides mechanisms and structures for combination and recombination. But this requires a "common" hypertext structure that can be easily and conveniently viewed, manipulated and "shared" between applications. Because this form of editing is so fundamental, it should be part of an operating system and an easily accessible "affordance".
The Web is not designed for fine-grained editing and rearranging/recombining content and has started as a compromise to get work done at CERN. For example, following a link is very easy and almost instantaneous, but creating a link is a whole different story, let alone making a collection of related web pages tied to specific inquiries, or, even making a shorter version of a page with some details left out or augmented. Hypertext goes far deeper than this.
Although a bit dated, I recommend reading Ted Nelson's seminal ACM publication in which he touches many issues concerning writing, how we can manage different versions and combinations of a text body (or a series of documents), what the problems are and how they can be technically addressed.
[1] "Complex information processing: a file structure for the complex, the changing and the indeterminate" https://dl.acm.org/doi/pdf/10.1145/800197.806036
Here's where I'm stuck:
Hypertext - whether on the web or just on a local machine - can't solve the UX problem of this on its own, though. People can re-arrange contents in a hypertext doc, recombine pieces of it... but mostly through the same cut-and-paste way they'd do it in Microsoft Word 95.
The web adds an abstraction of "cut and paste just the link or tag that points to an external resource to embed instead making a fresh copy of the whole thing" but all that does is add in those new problems of stale links, etc.
So compared to a single-player Word doc, or even a "always copy by value" shared-Google-doc world that reduces the problems of dead external embeds, what does hypertext give me as a way of making rearranging things easier? Collapsible tags? But in a GUI editor the ability to select and move individual nodes can be implemented regardless of the backend file format anyway.
TLDR: I haven't seen an compelling-to-me-in-2023 demo of how this system should work, doing things that Google docs today can't that avoids link-rot problems and such, to think that the issue is on the document format instead of user tools interface side.
I have to catch some sleep, but I will address your questions as good as I can later. In the meanwhile, you might want to take a look at how Xanadu addresses the problems of stale links, and maybe some of your other questions will be answered.
[1] https://xanadu.com.au/ted/XUsurvey/xuDation.html
Also, I highly recommend reading Nelson's 1965 ACM paper I mentioned to better understand the problems hypertext tries to solve and the limitations of classical word processing (which also expands to Google Docs).
The intro of the Nelson paper/Memex discussion is similarly alien to me. I don't think it's human-shaped, at least not for me. The upkeep to use it properly seems like more work than I would get back in value out. It's a little too artifact-focused and not process/behavior focused, IMO?
Anyway, I have limited computer access at the moment, but maybe you find the following response I wrote in the meanwhile useful. Ill get back to you.
---
Some remarks that hopefully answer your question:
The Memex was specifically designed as a supplement to memory. As Bush explains in lengthy detail, it is designed to support associative thinking. I think it's best to compare the Memex not to a writing device, but more to a loom. A user would use a Memex to weave connections between records, recall findings by interactively following trails and present them to others. Others understand our conclusions by retracing the our thought patterns. In some sense, the Memex is a bit like a memory pantograph.
Mathematically, what Bush did is to construct a homomorphism to our brain. I think it is important to realize that, when we construct machines like the Memex or try to understand many research efforts behind hypertext. Somewhere in Linda Barnett's historical review, she mentions that hypertext is an attempt to capture the structure of thought.
What differentiates most word processing from a hypertext processor are the underlying structures and operations and ultimately, how they are reflected by the user-interface. The user interfaces and experiences may of course vary greatly in detail, but by large the augmented (and missing!) core capabilities will be the same. For example, one can use a word processor to emulate hypertext activities via cut and paste, but the supporting structure for "loom-like" workflows are missing (meaning bad user-experience), and there will be a loss of information, because there are no connections, no explicitely recorded structure, that a user can interactively and instantly retrace (speed and effort matter greatly here!), since everything has been collapsed to a single body of text. The same goes for annotations, or side trails that have no structure to be hanged onto and have to be put into margins, footnotes or various other implicit structural encodings.
Hypertext links, at least how Ted Nelson conceptualizes them, are applicative. They are not embedded markup elements like in HTML. In Xanadu, a document (also called version) is a collection of pointers and nothing more. The actual content (text, images) and the documents, i.e., the pointer collections, are stored in a database, called the docuverse. Each content is atomic and has a globally unique address. The solution to broken links is a bit radical: Nothing may be deleted from the global database. In other Hypertext systems, such as Hyper-G (later called Hyperwave), content may be deleted and link integrity is ensured by a centralized link server. (If I am not mistaken, the centralized nature of Hyper-G and the resulting scalability problem was its downfall when the Web came). Today, we have plenty of reliable database technologies and the tools for building reliable distributed systems are much, much better, so I think that a distributed, scalable version of Ted Nelson's docuverse scheme can be done, if there is enough interest.
How a document is composed and presented, is entirely up to the viewer. A document is only a collection of pointers and does not contain any instructions for layouting or presenting, though one can address this problem by linking to a layouting document. However, the important point is that processing of documents should be uniform. File formats such as PDF (and HTML!) are the exact opposite of this approach. I don't think that different formats for multimedia content can be entirely avoided, but when it comes to the composition of content there should be only one "format" (or at least very very simple ones).
I hope this answers some of your questions.
I think I see what you mean. Garbling, as you mention it, is actually what Xanadu is supposed to prevent. The problem is that it is not explicitely mentioned that a document/version (a collection of pointers to immutable content) should also be an addressable entity (an part of the "grand address space") and must not change, once it has been published to the database. In particular, if a link, e.g., a comment, is made to a text appearing in a document/version, the link must, in addition to the content addresses, also contain the address of that document/version (In Fig. 13. [1] that is clearly not the case and I think that's a serious flaw).
This way, everything that document C refers to - and is presented - is at the time it was when the composition was made. How revisions are managed is an orthogonal (and important) problem, but with the scheme above we lose no information about the provenance of a composition and can use that information for more diligent curation.
I don't see any reason why the folders in question should necessarily be shared, though. I get the impression that as defined by Bush originally, the memex is first and foremost a memory expander for your private memory, and any sharing would take place infrequently, or on a narrow subset of data contained within.
What do you have available to you after a hard drive failure? That's the point (or at least one of them) of having stuff in the cloud.
How do you balance the two? Local, personal storage backed up in the cloud?
I wrote more about offline documentation here: <https://news.ycombinator.com/item?id=27838105>
Everything, because I pretty much started shunning everything I can't reduce to a bunch of files I can use anywhere, and am using the same folder structure on all machines. If an application has a portable version (and doesn't have configuration that may vary between machines, e.g. hardware sensors displayed in the tray), I always use the portable version, install it on one machine, and then sync that to the rest. I love it so much. I can open a Sublime Text project file anywhere and it just magically works, because the paths are the same.
The only reason I sync some things manually instead of using Syncthing for everything is that I don't want to needlessly cause more SSD wear. E.g. Firefox/Thunderbird constantly write and delete stuff in the profile, so if I want to use a profile on another machine, I close it, sync, and open it on the other machine.
My source of truth is completely in those files, including my web stuff. Which, admittedly, is rather simple: just PHP and JavaScript and MariaDB really, or sqlite where that's enough. I thought I liked golang, but after spending an hour yesterday trying to compile something I haven't touched in 5 years, I decided to just rewrite it in PHP - because something that is slow but doesn't require constant hand holding beats anything else for me. I want to be able to pick up where I left, regardless of whether 10 years passed. PHP and JavaScript are unsurpassed in my personal experience when it comes to that. Even the shittiest JS still works fine 20 years later, and when you use PHP with warnings turned on, upgrading to new versions tends to throw very few errors, and they're actually helpful, so making the changes is trivial.
I miss out on a lot of fancy things that way, but I don't miss any of them. I love thinking of my stuff as where it is in my "home cloud" folder structure, and that though they're not actually thin clients, the particular machines really don't matter. Of course I also have backups, but the plan is to never need them because I always have more than one "live" working environment, which reads from and writes to that shared set of files.
Certainly not everything at this stage, but in a properly run Memex system, IMO, that really should be everything. Assuming enough local storage, there is no reason not to save every page you view while browsing the web in full.
> What do you have available to you after a hard drive failure? That's the point (or at least one of them) of having stuff in the cloud.
The standard answer to that, is to have good backups. Bite the bullet and spend some time and money on buying the proper equipment (a commercial NAS, a Raspberry Pi NAS with some USB-connected drives, a small form-factor PC with a couple of SSDs, anything you want), and integrating it into your other machines as a backup destination. Cloud makes it much easier, but also less private, with an ongoing monthly cost, and adds an additional point of failure that's outside of your control.
As for how I balance it personally? Local storage all the way. Even if I picked an end-to-end encrypted backup service, several terabytes of data is just too expensive to store in the cloud.
To me, the biggest problem of this kind of tool is that if you use it to its fullest potential, it necessarily becomes a crucial part of your daily life, something that will cause very serious problems if you lose it.
And that means anything proprietary or closed SaaS is an absolute no go. My Mind's promise to respect privacy and stay independent of outside influence is nice, but worthless. They can change their mind at any time - or get bought up, or go bankrupt. Being a paid service might reduce the likelihood, but does not eliminate it, especially not over a timespan of decades.
So at an absolute minimum, such a tool needs to be able and willing to export all of my data in an easily processible format. This is far more important than shiny features.
My current solution is a self hosted TiddlyWiki. It may not have all the shiny features, but it is mine, and always will be, even 50 years from now. The worst case scenario is that it stops working on new browsers and nobody maintains it anymore - but I still have the data. Easily processible text files are actually its internal storage format.
this appears to be source available not open source
I also think the interoperable argument breaks down when you want to do anything off the beaten path. The author mentions how you might have a folder with a presentation or document. But attaching any kind of nonstandard data for the format is impossible, and requires the format parsing clients to know about it. You're limited by what each format supports and the total number of available, mature formats with mature parsers/editors.
I think it's very important to have the ability to work directly with files, for draconian environments you describe, but just because it's the simplest, stable, most portable, least-time to _something_ abstraction. It just has to be balanced with more complex functionality files alone can't handle.
My "memex" tries to use well-known abstractions so that I can do certain tasks in other programs until my "memex" can do it. That and because for quite a few operations, I'd like to simply mount the window of a separate program into my "memex" as a sort of window manager (but deeply integrated with the other stuff the "memex" does, like data and tab management) instead of trying to remake such monolithic, decently-built programs
That said, I've certainly heard that different people prefer different ways of learning/consuming knowledge. I myself will choose video for some things, illustrations/graphs/images for other things, and text/tables for yet others. I do have a fondness for well done images, though.
Even the style of text can have an impact. Bullet points for providing instructions, instead of a long paragraph of sprawling text.
Another thing I picked up some 15 years ago is borrowed from newspapers: the concept of "above the fold"; get the main points out there near the top, and keep it concise.
A well-done graph can communicate data and the desired message quickly. At the opposite end of the spectrum, we have poorly done or deceptive graphs.
An illustrated magazine will communicate a story differently, compared to a text-only novel. There's nuance to each approach that the other can't quite capture.
A photograph or painting can elicit a different level (breadth, depth) of mood/emotion that words might merely hint at.
Of course, what's useful in one context could be pointless in another. An introduction to photography would do well to include some illustrations, whereas a book on meditation could probably do fine with few, if any, images. Another example is Gray's Anatomy, which would probably be more challenging to consume if there were no illustrations.
Some subjects such as learning a martial art can be helped by using pictures/illustrations, but realistically when used as the only source of knowledge fall short of adequately communicating proficiency in said martial art, where only in-person instruction would really suffice.
I do think spatial organization can help for human consumption of not easily orderable datasets >1000 items, but I think that's a niche case for everything ive interacted with so far.
What did you have in mind?
For revisualizing, it might be comfortable to plot it on a 2D map or to simply graph an interesting scalar.
For revisualizing in a 3D spatial setting... I haven't seen any actual good 3D spatial layout of data. Disregarding "experiences" (like art galleries, which lay out data in a way to promote interest in a topic and fractal-like amounts of subtopics, but is human curated, like Saganworks[1]), the only one I guess is like the high level "portal" to categories of data. Like "Documents", "Downloads", etc but think 500+ categories of more granular entity types. Everything is non-comparable apples and oranges (except for perhaps shared common data like timestamp) and there's no way to quickly scan it (5 or so seconds to find alphabetically perhaps), so spatial might help to quickly locate your entity type with muscle memory. I haven't seen another use case, but would love to explore what others are doing in this space!
You have a shared folder backed by version control.
You have projects which can serve as memex projects, and issues which you can edit via markdown and link to the files in the repo.
Plus you got CI-CD that allows you to have access to a cloud environment to perform work.
As he touches on, the missing, or a least somewhat disjointed, element is local data and connecting to that from other web destinations. My hope, and belief, is that we are breaking down those local <-> server barriers with newer JS APIs for better file like local storage.
It's just a pity in a way that the web browser and web server were not combined into one element from the start. It's still too hard to make content available on the network from your personal device without a server.
His conclusion about the need to cognitively organise such trails/collections though could well change in the coming years - and potentially sooner. It's clear the LLMs are good at organising and summarising text and data. I'm expecting us to be using tools that will ingest all our work activity (documents, meeting, email, calls, browsing) and automatically curate and file it with a chat based ui to extract that data. We just desperately need that to be local and not server based.
I'm using Notion. Its ok, but I'd like to have a way to snapshot documents to incorporate into my notes. And lately I feel like I'm paying for all their new AI stuff that I just don't use.
So help me HN: whats a good open source, self-hostable, replacement with decent desktop/mobile tooling and a practicable migration path from a Notion export/backup?
In my humble opinion these tools aren't quite there yet, which is not to say they aren't polished, but that they don't disappear from view like good tools should.
There's still a noticeable impedance mismatch due to having to use their editor, for example, which has its quirks and is different enough to a standard text editor to cause a context switch.
But it's open source and it's pretty good.
I was using Typora for a while, before Obsidian got better. It was just a tree of Markdown files and a _really_ nice Notion-like editor.
I had to write https://github.com/Cobertos/notion_export_enhancer to get my data out of notion in an actual nice tree-like structure.
I tried Obsidian again recently and their editor caught up to Typora's (and it was faster than the last time I used it). I'm trying to switch to it now, as the tree is more "file-y" than Typora (which removes blind spots for me, as I save PDFs and other stuff in my note tree).
I've resorted to writing my own though now. Notion, Typora, and Obsidian are great for one specific type of management (documents, markdown-like), but I think there's a higher level abstractions here.
Think about a case where you are trying to figure out a problem based on a bunch of error messages. You search the web for solution to the first message, then the second message and so on, until you arrive at the root cause and the solution to that. You then save this as a trail (or the memex does this for you automatically, using some fancy ML method).
Now, imagine you encounter a similar or identical issue 6 months down the line. You remember that 6 months ago you searched for a solution, and that you found one. Maybe you forgot to bookmark it. Not a problem, you find the original trail using some simple heuristics (how long ago it was, the keywords from the error message), and then you follow it to the end. Solution found, reasonably quickly.
That was my point with the original article. I'm with you that it would be fantastic to have memex that's described in 'As we may think' (only better), but until we do, I'm doing it the hard/manual way.
Nothing wrong with doing this manually, but I do think that having the Memex do some of the thinking re: how the data is organise would save a lot of time.
I’m personally at a stage of disillusionment with the modern web such that I think returning to the early ideas without improvement is an improvement over the status quo.
It may be true that what the author describes is rudimentary, but to me, that’s kind of the point.
This isn’t to say that better tools aren’t needed. But I’m fully on board with what I see as a movement/mindset shift back to the fundamentals of the early web.
If memex becomes a concept of interest in the 2020s, perhaps interested people will invest in building better tools.
Umm, no. That is the problem really. An "edge" is something defined in societal context. Rightly or wrongly society is largely unaware of memex and its gifts.
Memex, hypertext, wikis, semantic web etc. These are just early visions of plausible ways we can capture, organise and share knowledge at scale, using digital technologies. Alas it is a catalog of paths that the digitally interconnected society did not opt for.
Why this (on the face of it) irrational behavior? The very short answer is that "an edge" is only loosely coupled to superior knowledge management. In the vast majority it is obtained by proximity to existing power structures. The proverbial rolodex rather than the memex.
I somewhat disagree. A memex is effectively the plain old notebook/wiki, with the potential capability raised to the second or third power.
Does a personal diary, calendar, home budget, or a small library help you get into places of power? For average people, no. But it does make organising life, and figuring out where you are, and charting how to get to where you want to be, much easier than it would be for your peers without such tools.
So yes, it gives you an edge. Whether that edge has anything to do with power is a completely orthogonal concern.
There is an empirical fact though, that begs for some explanation: These tools are not used as much as one would think is commensurate to their (potential) added value. There is no virtuous cycle that would fund more development, create sharper "edges" etc.
The connection with "power" might appear remote but consider e.g. why powerpoint (no pun) - a manifestly inferior piece of software - became so succesful.
[1] http://scripting.com/2023/08/06/131842.html?title=miriamIsGr...
I'd argue that this sort of interlinking flexibility (alongside easy-to-use plugins) is what separates the current era of note-taking apps over alternatives that people seem to (ab)use: a folder of Markdown notes edited via VSCode, a Google Docs folder, or even a static site generator.
Shameless reference to Whatboard.app where we are working hard to address this issue. Admittedly, we are over-focused on sales, and less intra-document linked. Much work ahead for us, and several enhancements in the pipeline.
www.whatboard.app if you want check us out. Reach out if you have interest in getting involved with us.
https://en.wikipedia.org/wiki/Lifestreaming#Before_lifestrea...
https://en.wikipedia.org/wiki/David_Gelernter#Politics
>Politics
>He is a former national fellow at the American Enterprise Institute and senior fellow in Jewish thought at the Shalem Center. In 2003, he became a member of the National Council on the Arts.[21] Time magazine profiled Gelernter in 2016, describing him as a "stubbornly independent thinker. A conservative among mostly liberal Ivy League professors, a religious believer among the often disbelieving ranks of computer scientists..."[22]
>Endorsing Donald Trump for president, in October 2016, Gelernter wrote an op-ed in The Wall Street Journal calling Hillary Clinton "as phony as a three-dollar bill", and saying that Barack Obama "has governed like a third-rate tyrant".[23][24] In his capacity as a member of the Trump transition team, Peter Thiel nominated Gelernter for the Science Advisor to the President position; Gelernter did meet with Trump in January 2017 but did not get the job.[25]
>In 2018, he said that the idea that Trump is a racist "is absurd."[26] In October 2020 he joined in signing a letter stating: "Given his astonishing success in his first term, we believe that Donald Trump is the candidate most likely to foster the promise and prosperity of America."[27]
>Gelernter has spoken out against women in the workforce, saying working mothers were harming their children and should stay at home.[13] Gelernter has also argued for the U.S. voting age to be raised, on the basis that 18-year-olds are not sufficiently mature.[28]
https://arstechnica.com/science/2017/02/trumps-science-advis...
https://yaledailynews.com/blog/2017/01/25/gelernter-denies-m...
https://slate.com/news-and-politics/2018/10/david-gelernter-...
https://www.politico.com/story/2017/02/donald-trumps-shadow-...