> Google uses research, published models, and data that was freely shared with them and iterates on it, making use of their vast budgets and hardware, to develop new models. Then, Google uses those models internally and doesn't share the models. This is a violation of academic norms under the pretense of "safety".
Google's not doing this (LLM, generative image model) research on academic datasets freely shared with them. They're doing this research on data they gathered at their expense. This is not a violation of academic norms. Again, Google shares a lot of datasets and models, just not LLMs and generative sets trained on problematic source datasets.
> As I characterized previously Google is able to do whatever they want, conceal their results, and impede progress and understanding because they aren't sharing their results. You say this isn't a "fair characterization" but it is exactly what is happening - which part is wrong?
Anyone can do research and not share back to the community. Google _does_ share back to the community in the form of papers (and again, very frequently with models and datasets). If you have the money and expertise to implement the papers, more power to you. Every technology company has some secret sauces they don't share with everyone. That Google may have some of those is not a moral failing.
> from a norms, ethical, and moral standpoint Google does have a massive obligation to the public that they are breeching. Google uses the public's data to train, public research, and publicly shared models to iterate on
From the other end: Google gets user data and has a responsibility to not proliferate that data, no? I wouldn't want them to share a dataset that has my personal data, even if anonymized because there are ways to deanonymize. There are levels to everything, and choosing "I'll release the paper but not the model + data" for some potentially sensitive models seems sane.
> Then, after building on the shoulders of giants, Google refuses to share what they have built in contravention of the norms that they benefit from.
People are building on the shoulders of Google's research all the time, and plenty of companies are doing similar things to Google and being way less open about their work. I mean, every company that trains a big model on data collected from the public -- are they all required to share their models with everyone? Is Cruise sharing their pedestrian detection model? I don't think what you're suggesting could possibly be the standard.
> Regarding your chair metaphor - the "danger" of these models, if there is such, is not that they would hurt the user, like a faulty chair, but that they could be used to hurt others - e.g. a bot army to manipulate public opinion or create fake news.
Sure, I was trying not to be hyperbolic and compare LLMs to guns since they have plenty of awesome use cases (whereas guns really don't). A faulty chair that you set out for anyone to use can hurt people other than the chair's creator / people who are aware of the specific risks. But yeah, seems like you now agree these models have the potential to cause great harm.
> In other words, if these tools can cause harm they should be regulated by the government, not Google
I agree that gov't regulation can be helpful for setting a minimum standard. But I strongly disagree that lack of laws means we should abdicate our own moral responsibilities. If I sell / provide something, I need to be able to sleep at night knowing I didn't make the world worse. Googlers typically try to do this.