Encyclopaedia Britannica Seeking $1B Valuation in IPO
bloomberg.com
bloomberg.com
But it seems like many of those authors are long retired, maybe dead, and I wonder how old their contribution is (Britannica seems to update it). And many articles say they are written by the 'Editors of Encyclopaedia Britannica', though maybe that always was the case.
I have nothing against the editors, but an expert in that domain will have perspective, insight and current knowledge (i.e., not yet in textbooks) that is impossible to match, imho. I'd love to see them again hiring leading people to write.
In Britannica the way it is now, I've noticed that related topics are sometimes written by experts with different viewpoints, and where they overlap you can see the contours of the terrain, so to speak.
If you look at specialist academic encyclopaedias (e.g. the Stanford Encyclopaedia of Philosophy), very often the article on theory X is written by one of the major proponents of theory X – because only they feel sufficiently motivated to write it, and because the editors think it is charitable to grant the proponent of a theory the opportunity to produce their best defence of it. But, they'll make sure to include coverage of the criticisms of the theory, because that's what an academic is expected to do. And the editors will make sure the article overall, and the coverage of the criticisms in particular, is fair rather than overly slanted – e.g. by recruiting one of those critics as a reviewer.
Edit. Ah!
Stanford Encyclopaedia of Philosophy. I read back a couple of comments and it clicked
And if true, I wonder what it would look like in the absence of Encyclopaedia Britannica, which accounts for 99% of the instances that I've seen the archaic spelling.
I'm guessing it's the last two letters that are going to get them that valuation.
Encyclopaedia Britannica is positioning itself not only as a tech firm, but an AI firm.
What's going on?
Britannica has a contract with Reddit? I tried searching around online but didn't find anything.
I wonder what the oldest company to go public is. This isn't easy to Google, because search engines want to give you the public company which is the oldest, even though that's not the same thing. I see this article about Birkenstocks, which company is apparently almost 250 years old. Surprising! But it does not mention Birkenstock being the oldest company at the time of IPO, so there must be an older one.
https://www.npr.org/2023/09/13/1198324691/birkenstock-ipo-st...
Also, I'd look in a country with both large securities markets and a long history (i.e., not the US). Maybe Japan, Italy, or France?
It was founded in the 13th century as Stora Kopparberg (first record is a private stock sale in 1288). It went public on the Stockholm Stock Exchange in 1901.
Today it's known as Stora Enso Oyi, following a merger with the Finnish company Enso in 1998.
If they can also completely eliminate hallucinations, they could be onto something useful.
Sure it won't come from just having a good data set... they'll need to do other stuff as well. Hopefully not impossible. :)
Another approach that can help, is after you generate the response, you then submit the response to an LLM (even the same LLM), with a prompt to check it for errors (errors of reasoning, misrepresentation of citations, etc). Often, LLMs can do a decent job on catching their own errors, including hallucinations.
The more advanced an LLM is, the less likely it is to hallucinate, and the more likely it is to pick up on its own mistakes. Llama-7b hallucinates a lot more than GPT-4, and is far less likely to detect itself doing so.
You can't make an LLM foolproof, but you can't make humans foolproof either. Humans "hallucinate" too – e.g. students write essays with erroneous claims, cited to sources which never actually make those claims.
Heh, we need "post-LLM" AI's then that don't hallucinate at all. ;)
Note that these are really very closely related, almost to the point of being slight variations on a single technique: both are “do a search on a database and include the results in a prompt to the LLM”; they differ in one uses something outside of the LLM to calculate the search query (typically, just an embedding of the user prompt used to search a vector DB), the other prompts the LLM so that it produces an appropriate search query.
They can produce rather different results – the LLM can do a better job than just an embedding at picking out what's truly important in the prompt, which can increase the relevance of the returned documents. I've seen before where a vector search returns documents which contain some similar language to the prompt but which are actually talking about something completely unrelated. LLMs are less likely to do that sort of thing, because they have a much better understanding of what's important in the prompt and what's just fluff/waffle
No. What a bizarre take. Why would anyone think that?
Having a training data set that's not based on fiction sounds like it could be a useful thing.
If they can somehow get rid of the hallucinations too it could be good.
Sorry (not sorry) for assuming that the second sentence in your comment was related to the first.
Hopefully it's clearer now. ;)
In both cases, you should not treat the information as canonically or authoritatively accurate or factual. Biases, gaps, outright lies and fabrication exist in any large collection of human writing.
I'm not doubting this, but do you have citations?
If you know it's heavily biased, maybe find something else?
I understand that you can't read every link in a comment, but you certainly don't have to post a comment like this if you haven't at least looked into what I posted.
Britannica's advantage is that it's written by experts and edited by professionals. If I want knowledge - for business, for my health, for legal, accounting, car repair, IT - I ask experts.
There are downsides - it's not nearly as large as Wikipedia, and it doesn't engage the crowd. But I know someone with expertise verified the facts, and the completeness, correctness, and consistency (the three C's of accuracy) of each article.
At the end of the day it's just a single centralised source
Probably wouldn’t be any less prone to confabulation than any other LLM, and given how limited in coverage the sum of any factual encyclopedia is compared to typical LLM training sets, would have even a spottier base of factual information that most LLM’s have to work from.
> If they can completely eliminate hallucinations
Yes, but that’s orthogonal to the training set.
Of course. I didn't say otherwise.
> An LLM that ...
Note that I also wrote "AI" rather than "LLM". ;)
Hopefully there will be better alternatives (to LLMs) at some point that don't hallucinate.
My point is that having one of those (something that doesn't hallucinate) trained on factual data sounds like it'd be useful. :)
There are AIs that don’t “hallucinate” (a bad metaphor), e.g., most of the technologies developed in the previous rounds of AI – like expert systems.
I doubt that we will find more advanced, equally-or-more general AI techniques than LLMs that don’t “hallucinate”, what we’ll probably do is find better ways to build systems around LLMs (or more future AI architectures) that allow them to resolve “hallucination” better before presenting results.
It’s not as if natural intelligences don’t both literally hallucinate and, more relevant to what is described as “hallucination” in AIs, confabulate when pressed to answer without access to facts, and we know with them that grounding in physical reality via sensory data and having access to factual references doesn’t mitigate those effects (and we’ve also been able to demonstrate the latter with LLMs.)
The AI’s already accidentally do that for topics people repeat a lot on the Internet, esp political. ChatGPT used to fight with me over some of those. That it would not relent or stop redirecting me showed that they can be trained to follow authoritative sources, too.