Show HN: Goopt – Search Engine for a Procedural Simulation of the Web with GPT-3
github.com
github.com
1. https://www.theatlantic.com/technology/archive/2021/08/dead-...
There is already a concern in corpora creation in ML/AI projects. Researchers would like very much to have human-only generated content when training models on internet-sourced text. People posting GPT-created output has the potential to taint these corpora and create all sorts of strange loopy feedback.
In my fantastical imagination of that implementation, I would simulate "sim agents" as they live a life, collecting unique experiences that are the result of interactions with other sims. They then build up personalities, which influence their expressions in other mediums that real humans would consume. It's psychotic and possibly evil and I thought of it first, so no one better take my idea.
You might be able to see what I mean just in this comment itself. Could this have been written by an AI?
Has there been any successors to web of trust ?
OTOH, is proof of personhood really what we want long term, or is that just a proxy? Hypothetically, if an AI is good & trustworthy enough, why not allow a higher "source rating" for that than low quality human content?
[1]: Or maybe CRAAP https://en.wikipedia.org/wiki/CRAAP_test
[2]: They quickly realized the trick is make high quality low quality content. 100 pages of gibberish is way less effective than a convincing essay that happens to include a few key falsehoods.
It’s a project I have on the back burner to analyze Reddit history to check what ratio of comments are actually original, and I’d like to build a link aggregator that sorts by novelty.
Under such a system any company that begins producing spam could be removed, and we could go back to the lovely days of something simple like page rank being used to provide relevant results.
It would also be cool if I could upload my own crawling modules, so I could index more than just websites.
Startup entrepreneurs in the mood for a Hail Mary play take note. How do you have a web search engine in a world where there no longer exists any algorithm for telling spam apart from real content? "Go back to the original Yahoo" is a decent start but certainly nowhere near a complete answer in 2022!
My guess is that it may not even take the form of what we have today, with an arbitrary text box. Maybe you have to go down to a specific category at least. Who knows. I sure don't. All I can say is that it sure looks to me like the spammers are only a year or two from effective total victory in the current paradigm.
The problem is the space between "AI generated" and "human generated" is fundamentally getting smaller. That's a real problem. We don't seem to be that many steps away from that space going to zero for generalized writing.
I'm not expecting this to hold for very long (openai has partnered with microsoft; and there are at least 2 groups currently working on replicating it open-source), and expect strong detoriation of overall web content soon afterwards.
Whether it’s convincing depends on your background. Those in society who can distinguish reality from fiction will still be rewarded, for obvious reasons. So that’s a difference from our current world in degree, not in kind.
With some minor modifications you could port it to goose.ai and it isn’t against their TOS.
EDIT: Forking it here https://github.com/zitterbewegung/Goopt to add the functionality above.
I get this \"content\": [{\"name\": \"Keto Diet\", \n \"typeClass\": {}, \n \"description\":[],\n \"contentImageUri\": null }]}
GPT-J doesn't seem to compose similar broken JSON so formatResults returns empty.
Another option is to improve formatResults to fix more cases of invalid JSON responses. If you get any improvement a pull request is welcome.
I'll come up with a pull request.
Even so, it is a good idea to adjust, I'll keep an eye on your fork. Thanks!
If you can make any adjustments that respect the terms, contact me to see it together.
I think most of the current form of pregenerated web with search actually becomes completely unnecessary, and it'll basically stop existing.
> The procedural web will be the future of the web. It will offer us infinite content
Yup it'll be infinite "garbage"
The current state of adtech and near total surveillance isn't sustainable as more people wake up to the downsides, and as fake crap begins to accumulate.
Decentralization of social media, advertising, e-commerce and other web 2.0 staples will be a natural evolution of technology. The story goes "under Google's model of the walled garden web, SEO, spam, and bots achieved parity in all content metrics except actual meaningfulness to the user." Despite having all the compute and talent you could possibly bring together, Google is failing to uphold its core technology. They incentivized bad faith behavior, and are reaping the consequences of that. The acceleration of seo hacking and artificial worthless content is asymmetrical to the acceleration of the capabilities and market model Google has created.
A search engine can navigate self selected communities, human curated lists, and creatively bundle lists of lists to achieve high quality results based on actual humans self selecting and acting in their own interests. You can do things with higher quality classification and even provide regex over crawled data without huge technical barriers. Search agents will come about, whether locally or cloud hosted, and will eventually replace centralized engines like Google.
There are non doomed visions of the future. Maybe we won't suffer a digital trashocalypse.
This will be better personalized garbage.
As the saying goes, "one person's trash is another person's gold."
Increased personalization of content at effectively infinite scales is a future I think most people really aren't wrapping their heads around.
There will be both good and bad applications of it, but when it finally crosses the threshold, it's going to hit like a tsunami.
The Internet built the infrastructure for 1:1 content, but there's simply never been the capacity for building out the creative.
When you are watching a TV show that's being written and rendered live specifically for you, incorporating your social media data to develop relevant topics and story arcs, and pulse rate and eye tracking to gauge and adapt for interest/emotional connection -- traditional content simply isn't going to hold a candle.
To me, that show is almost certainly going to be trash. But to you, it will be gold.
The components for that future are arriving faster than I ever thought they would, and while it is still a ways off, it's increasingly an inevitable result.
Your comment reminds me of people pointing out the uncanny valley years ago in claiming that computers would never create realistic looking humans. Just last week was research that not only can't people tell AI generated people apart from real ones, but they find the AI generated ones more trustworthy.
There's a huge difference between the beginning of implementation of a technology advance and its half-life.
Isn't the "procedural web" built of mountains of (hopefully) human written content? How will the system get content about new subjects without the humans writing it? Isn't a system like GPT-3 currently limited to reflecting the ground truth data it has seen?
Note the above is a statement on some of the risks to a procedural web. Not a real market opportunity.
This limitation went away recently. A variant called RETRO (Retrieval-Enhanced Transformer) can use a search engine to take in the exact information up to date [1], assuming you can curate your own text corpus. It's also 25x smaller.
[1] https://deepmind.com/research/publications/2021/improving-la...
How do I ask this system about the recipe, or the history of the cocktail? Someone has to write an article about it, right? How do they get paid if it gets scraped once and people go to the scraping model for the answer instead of visiting the original article's page?
These first large language models are naive, unoptimized implementations of data structures we're learning to inspect and optimize. Something like retro that runs locally with a "just clever enough" service agent is so close to workable. I can't wait to see what happens in ML over the next two years, and who knows what kind of radical evolution the next big algorithm is going to bring.
I think it's a similar problem we see today with ad-supported news being indexed by search engines, but taken to another magnitude when those articles need to be scanned by a model only once to have near perfect recall of the details.
http://cuiltheory.wikidot.com/ Cuil Theory
After reviewing the materials I see that Cuil Theory has come a bit further since I last read. I believe that Goopt would be somewhere around -2‽ from Cuil theory itself, negative because it's literal reality, but distant because it's an abstract embodiment.
Slightly off-topic. During my cursory reading I see that imaginary Cuil got fleshed out. I'd like a second opinion. The way it reads to me is that 'i‽' is almost the literal definition of solipsism.
Can you mix procedural and static content? How can you verify accuracy of information? What if you could refine a web page’s content just-in-time? Modifying the query and context etc.
Through a lens of Roam/Notion: what if everything were a block that could be individually linked? what if every block could be edited by anybody? what if anyone could add links and annotations across pages? a blend of web and wiki?
* How can you verify accuracy of information? I think this is one of the main difficulties, as the AI would have to understand contexts and have a notion of truth, I think this would already start to touch the capacity of "consciousness".
* What if you could refine a web page’s content just-in-time? You will be able to do this for every part that you don't like enough and want something better, or just to see something different.
However I'm concerned these personalization systems will be too accommodating and just tell people what they want to hear. One way people use search is for motivated reasoning: I believe something, I search for it, I find confirming evidence. A procedurally generated system I imagine is especially prone to this kind of massaging of queries to get desired outputs - tweak a few input parameters in the form of a query, and out pops the answer you seek. It's a hard problem to test for.
Funny thing is when people search to find what they want to hear, they're often driven to content farms, and those are increasingly ML-generated. Seems this is just cutting out the middlemen!
But if your browser is compromised so that encryption doesn't work, I think you have bigger problems.
Also note that NeoX-20B is pretty good, but it's not GPT-3 quality.
there's also the flip side around SEO spam, which is partly why founded Breeze, a newish topic search engine that leverages curation to hedge against the dark side of human / bot spam, etc.
bottom line, love this, having worked with GPT-3 in past and the direct impact on day job, all things search
I started working on this because I think the idea of the procedural web is interesting, I think this technology will come in the future maybe not too far away, so we have to start thinking about the possibilities, problems, dilemmas, paradoxes. Well, I think this is a big change, which has a direct impact on the information we consume and how we do it, the online experience, it could even shape our behavior as it is our information, reference, entertainment system, etc. We must be ready for a change of this size, so I invite you to give life to this topic and put it up for discussion.
What is Procedural Web?
This is a term that has not yet been treated as such, so I have tried to give my vision on the matter. The procedural web will be the future of the web. It will offer us infinite content, since it will not be necessary that someone has written or created it before. All the content will be synthetic and generated at the moment, with infinite possibilities. From informative text, articles, images, videos, to games, applications and services with interfaces and functionality that are automatically generated. All this adapted to our queries and needs, and increasingly personalized to our preferences.
Web 4.0 could be the propitious evolution for the procedural web. The automation that this new version of the web poses about applications, services, interfaces, APIs, devices and others, could be exploited by connecting them with the procedural web. These reality data interfaces would help empower and enhance their generative capabilities, as well as connect their functionality to the real world, having services and devices on which to execute actions. This makes the procedural web even more interesting, because it endows it with cybernetic capabilities.
The procedural web is based on natural language processing (NLP) and procedural content generation (PCG). Advances in these fields, as well as in the field of computational creativity, will allow us to generate increasingly better synthetic multimedia, thus nurturing the procedural web with more and better content. It will be interesting to see if this content comes to satisfy us more than the traditional web and human creation. Maybe one day we won't ever be able to distinguish.
Demo video and usage guide in the GitHub repository: