Tabbed browsing: a lousy band-aid over poor browser state management (2014)
old.reddit.com
old.reddit.com
https://github.com/plateaukao/browser
It's preferable to any other GUI browser on the device (versions of Firefox and Chrome), and even beats out apps such as Pocket which should make for improved readability of Web content on a tablet but which fails to achieve what EinkBro does, most especially in ease and consistency of navigation.
(ProTip: touch regions beat gestures by lightyears.)
EinkBro does not have a particularly outstanding tab-management interface (though it's markedly better than, say, Google Chrome/Android). It does have a killer feature I've mentioned previously on HN: the ability to save multiple pages to a single ePub document, including updating said document at later points in time.
This means the ability to compile a single document with, say, items of current interest, a day's (or week's or month's) reading dump, research on a specific topic, etc.
And: when you're done with that particular interest, you can grab and nuke the entire collection in one swell foop. A major failing of tab-based browsers is clean-up, often an all-or-nothing proposition.
Or copy the compilation to (an)other device(s), email it to family, friends, or cow-orkers, post to a website itself (mindful of course of the greatest impediment to information interchange ever invented, copyright law).
So, Daniel Kau is single-handedly delivering one of the biggest transformations to the browser world I've seen in over two decades, and possibly three. Not as much progress as I'd like to see (it never is of course), but far more than trillion-dollar monopolists seem interested in or capable of executing.
A memex would save EVERYTHING by default as it came in, putting it all into a long (month+) buffer... any links to "interesting" material would then cause it to be moved to the local archive. At any point, you'd be able to share your "associative trail" with others, along with ALL of the content referenced. You'd never be forced to deal with a broken link for static content, ever. (Which is why it can't be done... copyright law)
The other worst timeline thing we got was HTML - it's not markup, and it's not hypertext. It's embedded formatting with links off the page. You can include images and other things in an HTML "document", why not portions of other documents? (Security, that's why)
I want my Memex... it sounds like you're closer to one that I'll ever get, good for you! 8)
I've done some digging into the history of the Web browser and how and when it came to have the set of features it does today. A large portion of that seems to have chrystalised with, or at least by the time of, the ViolaWWW browser, an early though advanced application, first released in April 1992. This inludes not only graphical webpage rendering and features such as bookmarks, but frames, stylesheets, and scripting.
See the features list: http://viola.org/viola/vwFeatures.html
What Viola, and most other browsers didn't have was the notion of automatically and by default downloading and retaining copies of remote content. There was download-on-demand and caching, but neither are epecially reliable or human-usable. A large part of the reason was simply that disk storage was too expensive and small --- by the end of the 1990s my desktop system still had only 6 GB of storage (over three disks), and that was generous for the time. For text storage, the real threshold is somewhere between 64 and 128 GB, at which time a collection of thousands of books worth of content becomes viable.
I'd argue that saving absolutely everything likely isn't useful, at least for specified domains, content, or other criteria, saving in something akin to the Internet Archive's URL-and-time-indexed format could be quite useful.
Document embedding or transclusions (of Xanadu fame) seem ... more trouble than they're worth. Though the idea of links that can indicate the specific version and/or signature of their reference could be quite useful.
Copyright stands between much of what really should and could exist, and what is permitted. This is most unfortunate.
I strongly agree, but spinning disk space is very cheap, and you never know what you're going to want later... so at least a full month of time there before auto-purge seems quite doable for everything 8)
I'd build a system such that if you refer to a given file/page in that timeframe, then the reference count goes up, and it gets moved to permanent storage. 8)
Among questions are:
- Curation of archive: what is selected, how, and why?
- Cataloguing of archive: how do you find what's within it? What purposes does the catalogue serve?
- Weeding of archive: what materials are removed, when, and why?
- Classification of archive: how are materials organised (especially a concern where physical storage is necessary) and grouped?
- Characteristics of archive: what is included, what's been added or removed, what's been circulated, etc., etc.
From the LoC sources there's a lot of discussion of how one such system evolved, and what the rationales driving that evolution were. There's also discussion of what usage patterns and concerns were, at least for a physical-book oriented limited-access (members of Congress and other US Government officials principally, members of the public on a limited basis). As of the early 20th century, it seems that very roughly 1% of the collection actually circulated. I should compare that with more recent data....
Which is prelude to some observations for a next-gen Web / online document personal archive curation:
- Web-surfing activity is itself a form of curation and selection. One that is, it should be noted, highly influenced by exogenous factors: search-engine biases, site-aggregator preferences (such as HN), and social media trends and fads, which have their own group and algorithmic foundations.... That said, it's a first cut.
- Very little Web content is in ideal long-term storage format as published. In general, I'd strongly prefer tools which stripped page metastructure (headers, sidebars, etc.) which are unrelated to the core article text, ensure that metadata (title, author, publication date, publisher, references/links, and referrers) are preserved and standardised, and that interstitial intrusions (recommendations and advertising) can at least be suppressed, if not eliminated entirely. Hotlinked references should be dereferenced and localised, to prevent future breakage or abuse. Yes, capturing a snapshot view of all the original sins of an article is probably useful for archival, but it's not the preferred form of most content. Guiding principle: the author/publisher/aggregator is an ass. Also known as FYWD, an acronym ambiguously defined as "fine young western dinosuars" or "fuck your web design". Either is equally correct.
- Indexing and cataloguing should be separate functions and roles than authoring or publishing. We see this in libraries already, and for a good reason. The interests of the archivist are not the same as that of the author or publisher, and tend to align more strongly with those of the reader, though with added domain-specific knowledge and expertise. Tools such as full-text search (a form of indexing) work to an extent ... until publishers discover that they can game this for benefit, at which point the arms race begins.
- Rather than a fixed duration storage, a fixed size cache might be used, with items evicted from it according to some basis. Given the sparsity of indicia, random eviction or a randomised weighted eviction might even be reasonable. Otherwise there are a few characteristics which might be used:
1. Specifically-listed domains or URL patterns. E.g., everything from a specific blog or news site. Everything matching a specific pattern, say, your own commonly-used username(s).
2. Keywords of interest.
3. Weighting by which search patterns have matched content. (Here: local searches against archive.)
4. Revisit counts. How often has a particular item been revisited and re-read? This will all but certainly follow a Zipf function, and probably be low. I'd be surprised if > 1% of archived content is revisited on an annual basis. This might be generalised to revisits against domain, keywords, topics, etc.
5. Indicators when initially visited. E.g., a "save this forever" option. The key difference between present "save page" features is that the auto-archival would have its own standardised structure and metadata. I'm thinking that borrowing heavily from wget's WARC functionality would be a good general basis for this and the archival in general.
6. References between items. E.g., if you've saved a page A and pages B, C, and D reference it, this upweights retention for A.
7. Project / process references. If you'e actively.
8. Specified site mirroring. One might designate specific websites to be mirrored / replicated on an ongoing basis. Typically blogs or news sites, though others might also be included.
I don't know if this is overly complex or useful. I think if sufficiently implemented as a coded system it would be useful.
On random eviction: this is often used in other caching contexts. It tends to work if the material can be restored from a canonical data store. The Web has proved a not-very-reliable such store. There is the option of incorporating archival to services such as the Internet Archive (this can be automated) or Archive.Today (there are manual steps, though the process can be streamlined, and I've done so for collections of thousands of documents for which I've had a specific interest in preservation).