And has anyone figured out a good way to minimize that when sourcing training data from e.g. Reddit?
And has anyone figured out a good way to minimize that when sourcing training data from e.g. Reddit?
It seems like provenance/traceability of data sources will become important, both for the builders and users of LLMs.
Now, I agree that having scalable generative AI and high-throughput-high-virality mass media (social networks) brings new unspeakeable horrors like feedback loops to the mix. Interesting times indeed.
from https://twitter.com/jathansadowski/status/162524580321127219...
This was pretty clever:
>CaligulAI <-- @Rogntuudju
I believe it's spelled "Caligulae" because it's first declension.
/Latin-joke
You know what would be funny ;
."Latin Joke Explainer" -- As explained in Latin Man-Splain-Terms."*
EDIT:
(You mastered my joke, and I appreciate it.)
EDIT AGAIN:
UI think we actually just coined a term ; 'Caligul::AI' -- Malicious AI for its own pleasure.
And it seems we're emulating that same cycle. Humans are trained not only on our own output, but also take changes from our environment. So in this sense LLM's environment would be human input.
Until that occurs, it's really just a miasma of bullshit.
At least according to this paper (I think, FYI I am not a expert) https://arxiv.org/abs/2305.17493
https://venturebeat.com/ai/the-ai-feedback-loop-researchers-...
Combined with pre-LLM data sets.
"Hey Steve, come on up, I'm spittin on bugs"