STORM: Get a Wikipedia-like report on your topic
storm.genie.stanford.edu
storm.genie.stanford.edu
This whole trend is going to get much worse before it gets better.
I've noticed that the hallucination rate of newer more SOTA models is much lower.
3.5 sonnet hallucinates less than gpt 4 which hallucinates less than gpt 3.5 which hallucinates less than llama 70b which hallucinates less than gpt 3.
Synthesis of Topic Outlines through Retrieval and Multi-perspective Question Asking.
From the paper it seems like this is only marginally better than the benchmark approach they used to compare against:
>Outline-driven RAG (oRAG), which is identical to RAG in outline creation, but
>further searches additional information with section titles to generate the article section by section
It seems like the key ingredients are:
- generating questions
- addressing the topic from multiple perspectives
- querying similar wikipedia articles (A high quality RAG source for facts)
- breaking the problem down by first writing an outline.
Which we can all do at home and swap out the wikipedia articles with our own data sets.
PROMPT: create 3 diverse personas who would know about the user prompt generate 5 questions that each persona would ask or clarify use the questions to create a document outline, write the document with $your_role as the intended audience.
We're long overdue for better sources of online revenue. I understand that AI costs money to train (I don't believe that it costs substantial money to run - that's a scam) but if we thought that walled gardens were bad, we ain't seen nothin yet. We're entering an exclusive era where the haves enjoy vastly more money than the have nots, so basically the bottom half of the population will be ignored as customers. The good apps will be exclusive clubs that the plebeians gaze at from afar, like a reverse zoo.
I just want something where I can pay 1 cent to $1 to skip login. Ideally from a virtual account that's free to use but guilts me into feeding it money. So maybe after 100 logins I pay it a few dollars. And then a reward system where wealthy users can pay it forward so others can browse for free.
I would make it in my spare time, but of course there is no such thing in the 21st century climate of boom-bust cycles and mass layoffs.
> Cuil worked on an automated encyclopedia called Cpedia, built by algorithmically summarizing and clustering ideas on the web to create encyclopedia-like reports. Instead of displaying search results, Cuil would show Cpedia articles matching the searched terms.
https://hachyderm.io/@inthehands/112006855076082650
> You might be surprised to learn that I actually think LLMs have the potential to be not only fun but genuinely useful. “Show me some bullshit that would be typical in this context” can be a genuinely helpful question to have answered, in code and in natural language — for brainstorming, for seeing common conventions in an unfamiliar context, for having something crappy to react to.
> Alas, that does not remotely resemble how people are pitching this technology.
One recent catastrophic failure I found: Ask an LLM to generate 10 pieces of data. Then in a second input, ask it to select (say) only numbers 1, 3, and 5 from the list. The LLM will probably return results numbered 1, 3, and 5, but chances are at least one of them will actually copy the data from a different number.
Unless of course we rephrase it as "when I roll 2d6, why do I sometimes get snake eyes?"
LLMs are looking at typical constructions of text, not an understanding of what it means. If you ask it what color the sky is, it'll find what text usually follows a sentence like that, and tries to construct a response from it.
If you ask it the answer to a math question, the only way it could reliably figure it out is if it has in its database an exact copy of that math question. Asking it to choose things from a list is kinda like that, but one could imagine that the designers would try to supplement that manually with a different technique from pure LLM.
The topic of knowledge synthesis is fascinating, especially in big organisations.
Moving away from fragmented documents into a set of facts from which LLM synthetize documents from, tailored for the reader.
There are few tricks that would be interesting to have working.
For instance the agent keep evaluating itself against a set of questions. Or user adding questions to see if the agent is able to understand the nuances of the topic and so if it can be trusted.
(Not dissimilar to what would be regression testing in classical software engineering)
Then the "homework" sections, when we ask human experts to evaluate that the facts stored by the agents are still relevant and up to date.
All these can then be enhanced with actions usable by the agent.
Think about fetching the PoC for a particular piece of software. It is the employer Foo.
If we write this down in a document, it will definitely get outdated when Foo move, or get promoted.
If we put this inside a knowledge synthesis system, the system itself may keep asking every 6 months to Foo if it is still the PoC for the software project.
Or it could daily talk with the LDPA system and ask the same question as soon as it notices that Foo has changed its position or reporting structure.
This can be expanded for processes to follow. Report to create, etc...
But I like the direction of the research. I'd like to be able to specify the output reduction prompts and to tweak the evaluation agents.
This is "just" multi-agent summarization and synthesis. Most summarizers are already doing this.
Nice thing that is this is open source, https://github.com/stanford-oval/storm
[1] https://www.phind.com/search?cache=z3qe9c0z6yb0x1hqbq64mrci
[2] https://www.perplexity.ai/search/please-summarze-and-explain...
(I built it because ChatGPT couldn't search the web yet. When Phind launched a few weeks later, my project was basically obsolete!)
It seems the main improvement this paper has over that naive approach is the quality of the inputs, i.e. using "trusted sources" rather than random web results. (They appear to get their sources from Wikipedia itself?)
I'm not sure how much value all the other steps in the process add though.
>Sorry, STORM cannot follow arbitrary instruction. Please input a topic you want to learn about. (Our input filtering uses OpenAI GPT-3.5, which may result in false positives. We apologize for any inconvenience.)
A final point, the notice states that "The risks associated with this study are minimal. Study data will be stored securely, in compliance with Stanford University standards, minimizing the risk of confidentiality breach." When I use STORM, I can see other people's request. Are they supposed to be confidential?
It's far from being perfect yet, sometimes too shallow and lacking a guiding thread, but after few iterations we believe it should offer all the information a visitor might need when planning a visit.
Also would love a way to try without google authentication
`Panther Moderns,' he said to the Hosaka, removing the trodes. `Five minute precis.' `Ready,' the computer said.I did discover a dearth of published information about just how much energy a 4x4 LUT requires per cycle, and it's idle power.
Lots of text, lots of headers, but extremly shallow.
They have a "Discover" page with previously generated articles, but I think that they have some sort of manual review process to enable public access and it's not updated frequently. The newest articles there were from July. I tried copying the link for a previously generated article of mine and opening it from a private browser window but I just get sent to the main site.
I would think sharing by URL should work, but has some bugs with it currently.
Everything is so siloed & biased now, it's hard to find any presentation of a topic from a source that has no agenda. AI to help surface, aggregate, and summarize real data, expert opinion, and analysis like this would be really powerful and much needed. Essentially on-demand wikipedia articles held to the same editorial standards. Wikipedia isn't perfect by far, but their model has been surprisingly successful considering the challenge.