Note that you can get the model weights on HuggingFace here: https://huggingface.co/adept/fuyu-8b
92 karma · joined March 19, 2017
website: https://www.augustusodena.com/
Note that you can get the model weights on HuggingFace here: https://huggingface.co/adept/fuyu-8b
We are spending a lot of time thinking about reliability and it's true that existing models fall a little flat here. I think ultimately the key to making this work really well is some combination of
a) collecting and training on human feedback and b) doing intelligent things to samples from the model after the fact
And we are not building a chatbot, we're building something collaborative that you can work with to accomplish the stuff you want to do!
1. There's a spectrum (sort of) between using full on RL techniques and just doing sequence modeling. We're trying to pick a reasonable place on that spectrum that lets us model whether things have gone well without doing too much fiddling.
3. It really depends on how closely related the domains are. I think it's safe to say that you should expect more transfer of abstract/high-level capabilities than nitty-gritty things related to the specific domain - that's part of why we're excited about training one big model to use all software tools.
I do personally, but there is some disagreement about this in the field. In fact, I would go further and say that (in addition to using large pre-trained models) we will need methods of training that are pretty substantially different in order to elicit robust reasoning behavior.
Even supposing I'm wrong about this, if you go and look at the scaling plots in figure 3 and try to figure out how big your model would need to be in order to be solving most of these problems, you'd get a really big number. Even if you had such a big model, it would still require post-processing of the samples to actually get the right answers. From the perspective of applications, that's fine, but it's a little unsatisfying from the perspective of studying intelligence. Even with those caveats (!) these problems aren't that hard compared to general software engineering tasks...
> What do you think about leveraging unsupervised training to improve program synthesis? Could synthesized programs be executed on generated input in a way that supports contrastive learning [5]?
I think this is an interesting idea and someone should try it! I do think that, even restricting our attention to just getting neural networks to execute programs, that we will need to do something a little more drastic to robustly get the results we want.
There are people who fall more on the side of bitter-lesson/scaling-law-maximalism, and I think it's probably healthy and valuable that there are people in the research community placing both types of bet.
More generally, whether these models are well-calibrated (that is, they know what they don't know) is an important area of research. I don't have references offhand, but I think it's true broadly speaking that these larger pre-trained models do tend to be better calibrated.
We are working both on synthesizing programs from scratch (see https://arxiv.org/abs/2002.09030 for example) and on understanding computer programs using machine learning (see e.g. https://arxiv.org/abs/1911.01205).
I'm always happy to correspond with people about these topics.
Maybe the work is derivative of a particular training image if we can easily predict the presence or absence of that training image given only the trained model?
IIUC, that's something that's been looked at in the ML Fairness literature. See this paper for example: http://www.aies-conference.com/wp-content/papers/main/AIES_2...
> Question: How does copyright work for GAN output? If I input 300,000 copyright protected photos of celebrities and generate images of new celebrities that do not exist, are the generated images public domain or would there be copyright issues?
AFAIK, this is not a settled issue, but I'd be really interested hear what an actual lawyer thinks about this?
I'll answer what I think you're asking and you can tell me if I got it wrong:
There's been a lot of effort spent on coming up with different loss functions for GANs, but https://arxiv.org/abs/1711.10337 shows that, according to the metrics in problem 5, they don't really improve results compared to the original GAN loss function.
There's something called a Wasserstein GAN https://arxiv.org/abs/1701.07875 that you may have heard about, but IMO the useful thing to take away from that paper is their 'gradient penalty' technique and not the new loss function.