The History of Microsoft Encarta
abortretry.fail
abortretry.fail
I inherited an encyclopedia from the 1930s and it is a constant source of surprise and wonder to me - both in what is different from today and in what I expect to be different but is not.
I think Encarta is the last of this generation.
Sure, I can download a snapshot from Wikipedia and I did and others did and do as well. It's still not the same, because all our snapshots are different. Anything surprising in there could just be fluke, an editing error, the temporary state of an edit war.
A Wikipedia snapshot proves nothing. My 30s encyclopedia on the other hand is a stable reference. If I doubted anything in my copy I could easily find another antiquated copy and compare. Same with Encarta, there are so many copies out there that this particular snapshot of human knowledge will never die.
If the study of history is a long-term, iterative process, where is this stability important (beyond the filtering of short-term noise, vandalism, political influence, etc)?
The 11th, 12th, and 13th editions of the Encyclopedia Britannica were from 1911, 1922, and 1926.
So if, as a researcher, you were interested in the state of knowledge in 1916, you're out of luck (if you're limiting yourself to a particular encyclopedia, for example).
I don't know what value you see in "stability", when the world's knowledge changes on a second-to-second basis. It's equally arbitrary whether you freeze knowledge at 1911 or in any given point in time since Wikipedia began.
When you say "A Wikipedia snapshot proves nothing", I have no idea what that means. If you're worried about vandalism or quick edits, it's trivial enough to compare with earlier and later versions of the page to ensure you're looking at relatively stable text. Heck, most of that could even be automated if you wanted.
And while Wikipedia articles are full of errors, so too were the articles in every edition of Britannica. The difference is that Wikipedia errors tend to get fixed a lot quicker, while the Britannica errors just remained on the paper they were printed on.
I genuinely don't see what difference there is between the "stability" of the Britannica 1911 edition, and the "stability" of Wikipedia at some arbitrary timestamp. Both capture a similarly arbitrary moment in time -- Wikipedia just gives you so many more to choose from.
I've had the idea of building something like this. The concept being that you would select an article and a time interval, and be shown the "best/most stable" revision of the article within the given window. The tool could use any number of metrics for determining which revision is best, the most reasonable one I've managed to come up with is "highest number of views during a state where the article was not locked/available to edit".
I wasn't thinking as much as identifying any single best revision, but utilizing more of a diff-like tool to identify the text/changes that remained most stable over time -- where a brand-new edit doesn't count for much, but the longer it stays around as other edits are made, the more trustworthy it presumably is.
I think the biggest problem comes as articles get rearranged and expanded -- a section gets split into two or three, something gets moved from one section to a more appropriate one, and so forth. Or heck, sometimes entire articles get split into multiple ones, or vice-versa. I'm not aware of any diff-like tool/algorithm that handles these situations well, to accurately track how the same information gets moved when it's not just a simple case of insertion.
I suppose you can't really count on the same text/markup being shifted around as articles get split and modified in the ways you've described. Also I suppose there is no such thing as a cross-article edit in MediaWiki terms iirc. Use vector embeddings? Throw an LLM at the problem? Rate editors on their familiarity with a given topic area (and track how that evolves over time)?
The idea of using edit information in addition to the raw text written by editors seems like it's extracting additional bits of information from human interactions.
I might have read some idea in a HN comment of training AI not just on code, but on how that code is edited in a git repo, or maybe I am just imagining it.
A non-historian mathematician does a good job explaining the issue of census data errors, https://m.youtube.com/watch?si=ySApTldsYVf3jv0W&t=1630&v=GVh...
The knowledge was never stable, even back then. The encyclopedias could only afford a team of some fixed size, to publish an edition every few years.
Wikipedia just scales that up to many editors doing real time edits. Arguably it's a more reflective representation of how organic knowledge transfer actually happens.
As far as scaling up is concerned, Wikipedia has "power contributors" like any UGC platform. One guy alone has 3 million edits.
https://www.cbsnews.com/news/meet-the-man-behind-a-third-of-...
I don't think it's still a realistic argument to be having in 2023 :)
But seriously, these were the discussions we've had two decades ago. If a professor today had such an issue, they're very old fashioned, and I'd probably walk out of that class because who knows what else that professor is out of date on.
So many fields today are evolving so rapidly, I wouldn't trust any single expert on a topic. Better to have a living crowdsourced reference that collects many sources.
As great a free resource as Wikipedia is, each article's quality relies on a knowledgable contributor to really make it worthwhile, and those are few and far between. Wikipedia is dry and lacks what Encarta had
https://www.hbs.edu/faculty/Publication%20Files/Reference%20...
https://web.archive.org/web/20031026131928/http://www.howtok...
I felt this with encarta, and old printed volumes. It's written by someone, or told in a different way. While wikipedia feels more like a text book reference.
Wiki is open source, so it's limited by having access to CC-NA licensed assets only. It's also firmly an "Internet-first" encyclopedia so it's beauty comes from its SEO-friendly information architecture and internal link structure .
I'd bet it's easier for someone to find specific info on Wiki quickly, whereas Encarta would be better as a slow browsing experience.
Encarta was available on the Macintosh in the 90s
See many usage examples here: https://twitter.com/conzept__
Mine was 2 Girls 1 Cup o_O
(I never owned Windows, so I never played with Encarta; it's possible the same experience was achievable, but I like dedicated software to play with when there's no real advantage to integrating it into a larger ecosystem.)
It's interesting. E.g. until "Encarta 96", the application was "16 bit". I wonder if this means that 95 and before worked on a plain 8088 5150. This 95 article[1] claims otherwise (386SX).
At Blekko we, of course, crawled all of Wikipedia. One of the more interesting aspects was how many dead links it had to references (which presumably at one time were not dead). Sometimes you could find those references in the wayback machine and sometimes they were just "poof" gone. This is the nature of the web, information isn't persistent. Hence I think periodically pulling it into static storage would be a good thing for longevity.
This is largely a solved problem. There's a number of bots on Wikipedia and other WMF wikis which periodically trigger Internet Archive dumps for all external links which are used as references, and which can replace those references with links to the archive if a site goes offline.
https://www.britannica.com/topic/Encyclopaedia-Britannica-En...
Also, _World Book_ is still in print:
https://arstechnica.com/culture/2023/06/rejoice-its-2023-and...
Since then, lots of folks have edited and contributed content.
It was a pretty open secret that when Funk & Wagnalls (where Encarta got its content) was first starting that they would pay college students to write articles, the college students would crib from EB or some other similar source, and then submit that.
I think it was a good tool. Had to uninstall it a number of times when hard drive space was low.
I remember using it for homework like History. I was luckly it provided a sentence or two. Not enough to really learn something.
Perhaps Encarta '98 onwards were much, much better.
Today, we have Wikipedia, google search/maps, assitance like Alexa and, now ChatGPT.
Imagine the next 20 years.
The Software Toolworks Multimedia Encyclopedia was better IMO (at least in terms of UI and multimedia), and then from 1995 onwards, Encarta was generally better (again IMO)...
https://en.wikipedia.org/wiki/Phil_Spencer_(business_executi...