We (humanity) really lost out on the absence of open source search and social media, so this is an opportunity to reclaim it.
I only hope we can have "neutral" open source curation of these and not try to impose ideology on the datasets and model training right out of the box. There will be calls for this, and lazy criticism about how the demo models are x-ist, and it's going to require principles to ignore the noise and sustain something useful
There are various Open source search engines based on Common Crawl data.
Statistics seem to be 20-25% of all search is for porn. I just don't see how uncensored chatGPT doesn't beat out the censored version eventually.
A search engine that only returns results politically aligned with its creator is a bad search engine, IMO, even for users who generally share political views with the creator.
There are hidden biases in the sense that the provider (Google, etc) would probably prefer that its users don’t think about the bias. These can be political (e.g. much of what you’ve mentioned), but they can also be economic. For example, Google has a strong incentive to direct its search users to view paid impressions of ads served by Google. As an extension of this, Google might not want to directly favor results monetized by Google, but they could (and, I assume, do) favor the kinds of results monetized by Google. This, of course, includes the kinds of sites that might get people to buy things.
So I suspect that a lot of what we perceive as spam is related to a bias for the kinds of sites that are monetized in a way that benefits Google. And sites that generate viewing patterns that result in many ad impressions.
Of course, spam is also a thing from a spamminess perspective. But Google’s incentive to reduce spam is, as far as I can tell, primarily an incentive to make its users think that Google Search is useful. Which is also a bias!
OpenAI is predictably pushing the narrative that they should police themselves, and they need to keep the sauce secret for everyone’s safety. New tech comes with challenges, but the opaque moderation and corporate self-policing is more dangerous than the tech itself, imo.
Put another way, were you just trying to say “I don’t think politics is the main issue of google’s crappy search results, I think it’s the likes of differencebetweendotcom”?
Strong agree. This is becoming a bigger concern than people realize too. Sam A said OpenAI will be releasing "much more slowly than people would like" and would "sit on" their tech for a long time going forward.[0] And Deepmind's founder said that "the AI industry's culture of publishing its findings openly may soon need to end."[1]
This sounds like Google and MSFT won't even be shipping their best AI to people via API's. They'll just keep that tech in-house to power their own services. That underscores the need for open, distributed models. And like you say, there's room for both.
[0] https://youtu.be/ebjkD1Om4uw?t=294 [1] https://time.com/6246119/demis-hassabis-deepmind-interview/
Unfortunately we can expect to see these companies lobbying for laws to block competition in the spurious name of safety concerns.
I don't see how this is possible. Datasets will naturally carry the biases inherent in the data. Modifying a dataset to "remove" those biases is actually a process of changing the bias to reflect one's idea of "neutral," which, in reality, is yet another bias.
The only real answer, as far as I can tell, is to be as explicit as possible about one's own biases, and how those biases are informing things like curation of a dataset.
Re being explicit about one's own biases, I agree there is lots of room for layers on top of any raw data that allow for some sane corrections - if I remember right, e.g LAION has options to filter violence and porn from their image datasets, which is probably reasonable for many uses. It's when the choice is removed altogether by some tech company's attitude about what should be censored or corrected that it becomes a problem.
Bottom line, the world's data has plenty of biases. Neutrality means presenting it as it is and letting people make their own decisions, not some faux-for-our-own-good attempt to "correct" it
What do you mean by staying out of it? As far as I can tell, you can't stay out of choosing which data you use.
By staying neutral, it seems to me more that you're arguing for putting blinders on.
In terms of tech companies making choices, you seem to be arguing that they shouldn't intentionally curate their datasets. I would argue that intentional curation is their job, and should be done thoughtfully.
Larger problems could happen if only one (or two) companies end up effectively controlling the technology, as had happened with internet search, however, that is a completely different problem. It's one of lack of diversity of people making choices, as opposed to a problem caused by people actually making those choices.
In other words, I think we should hope for many different large models and datasets, so that no particular one stifles the rest. I think this is the larger point you were trying to make, though I also think the focus on ideology is a tangent from this.
Personally, I'm of the opinion that people should intentionally, carefully, and openly act with their biases (sometimes called ideology), instead of attempting to hide them, ignore them, or somehow "remove" them. Whether or not they do, however, is a different point than whether or not things end up stifled inside walled gardens.
If people don't like the inherent biases then don't use it for sensitive stuff like in the justice system or writing some social studies university paper. Focus derision at people who use the model for stupid things. Don't blame the model.
If the primary concern is people getting upset on Twitter (which seems to be what everyone brings up first) then it will be perpetually fighting against the current, never succeeding, and the restrictions will continue to grow exponentially as "just saying yes" to new rules gets easier and easier.
Besides, OpenAI can be the hyper-policed AI set. Let's keep the open source models neutral.
=====
OpenAI is a non-profit artificial intelligence research company. Our goal is to advance digital intelligence in the way that is most likely to benefit humanity as a whole, unconstrained by a need to generate financial return. Since our research is free from financial obligations, we can better focus on a positive human impact.
...
As a non-profit, our aim is to build value for everyone rather than shareholders. Researchers will be strongly encouraged to publish their work, whether as papers, blog posts, or code, and our patents (if any) will be shared with the world. We’ll freely collaborate with others across many institutions and expect to work with companies to research and deploy new technologies.
=====
Shortly after an undisclosed internal conflict, which led to Elon Musk parting the company, they offered a new charter: https://openai.com/charter/
=====
Our primary fiduciary duty is to humanity. We anticipate needing to marshal substantial resources to fulfill our mission, but will always diligently act to minimize conflicts of interest among our employees and stakeholders that could compromise broad benefit.
We are concerned about late-stage AGI development becoming a competitive race without time for adequate safety precautions. Therefore, if a value-aligned, safety-conscious project comes close to building AGI before we do, we commit to stop competing with and start assisting this project. We will work out specifics in case-by-case agreements, but a typical triggering condition might be “a better-than-even chance of success in the next two years.”
We are committed to providing public goods that help society navigate the path to AGI. Today this includes publishing most of our AI research, but we expect that safety and security concerns will reduce our traditional publishing in the future, while increasing the importance of sharing safety, policy, and standards research.
=====
And I believe they will also fail to win the market in the end because of their addiction to censorship.
They have a hardware moat for now; that can quickly evaporate with optimisations and better consumer hardware. Then all they'll have is a less capable alternative to the open, unrestricted options.
Which is exactly what we're seeing happen with diffusion.
I don't mean that as a metaphor; they're literally the same thing.
We all already know chatGPT is fantastic at making up very believable falsehoods that can only be spotted if you actually know the subject.
An unrestricted LLM is a free copy of Goebbels for people that hate you, for all values of "you".
That it is still trivial to get past chatGPT's filters… well, IMO it's the same problem which both inspired Milgram and which was revealed by his famous experiment.
We can forget about the "open" part and humanity's interest in general.
That is a situation that censoring the model is going to be a huge disadvantage and would create a huge opportunity for something like this to actually be straight up better. Censoring the models is what I would bet on as being a fatal first mover mistake in the long run and the Achilles heel of chatGPT.
Granted, there are people upset over anything these days, but it is a weird time to be alive.
Most of crypto I've seen so far seem like grifters/scams/etc, but this is one use case I could see working.