HNHacker News
TopNewBestAskShowJobs

tasdfqwer0897

92 karma · joined March 19, 2017

Augustus Odena

website: https://www.augustusodena.com/

submissionscomments
tasdfqwer0897··on Fuyu-8B: A multimodal architecture for AI agents
Hey I work at Adept and helped make this! Happy to answer questions. The thing I think is especially neat/notable is how simple you can make the model architecture while still getting good performance. I expect we'll continue to see bits of these models get deleted in the next few years

Note that you can get the model weights on HuggingFace here: https://huggingface.co/adept/fuyu-8b

tasdfqwer0897··on Act-1: Transformer for Actions
Yeah this is a good point!

We are spending a lot of time thinking about reliability and it's true that existing models fall a little flat here. I think ultimately the key to making this work really well is some combination of

a) collecting and training on human feedback and b) doing intelligent things to samples from the model after the fact

tasdfqwer0897··on Act-1: Transformer for Actions
Yeah, we did have to custom-build our own benchmarks.

And we are not building a chatbot, we're building something collaborative that you can work with to accomplish the stuff you want to do!

tasdfqwer0897··on Act-1: Transformer for Actions
Thanks - glad you like it! I probably won't get to all of these but let me try a couple:

1. There's a spectrum (sort of) between using full on RL techniques and just doing sequence modeling. We're trying to pick a reasonable place on that spectrum that lets us model whether things have gone well without doing too much fiddling.

3. It really depends on how closely related the domains are. I think it's safe to say that you should expect more transfer of abstract/high-level capabilities than nitty-gritty things related to the specific domain - that's part of why we're excited about training one big model to use all software tools.

tasdfqwer0897··on Act-1: Transformer for Actions
We used a combination of human demonstrations and feedback data! You need custom software both to record the demonstrations and to represent the state of the Tool in a model-consumable way.
tasdfqwer0897··on Act-1: Transformer for Actions
Yes! We plan on putting out a more detailed technical post soon.
tasdfqwer0897··on Act-1: Transformer for Actions
Hey, I helped make this! Happy to answer any questions.
tasdfqwer0897··on Program Synthesis with Large Language Models
> do you see improvements in Transformer or attention based architectures as essential...

I do personally, but there is some disagreement about this in the field. In fact, I would go further and say that (in addition to using large pre-trained models) we will need methods of training that are pretty substantially different in order to elicit robust reasoning behavior.

Even supposing I'm wrong about this, if you go and look at the scaling plots in figure 3 and try to figure out how big your model would need to be in order to be solving most of these problems, you'd get a really big number. Even if you had such a big model, it would still require post-processing of the samples to actually get the right answers. From the perspective of applications, that's fine, but it's a little unsatisfying from the perspective of studying intelligence. Even with those caveats (!) these problems aren't that hard compared to general software engineering tasks...

> What do you think about leveraging unsupervised training to improve program synthesis? Could synthesized programs be executed on generated input in a way that supports contrastive learning [5]?

I think this is an interesting idea and someone should try it! I do think that, even restricting our attention to just getting neural networks to execute programs, that we will need to do something a little more drastic to robustly get the results we want.

tasdfqwer0897··on Program Synthesis with Large Language Models
I personally agree that this experiment is evidence that there are certain problems that cannot be solved simply by making the models bigger, and one of the main research questions I'm interested in is what we need to do to elicit more reasoning-like capabilities from them.

There are people who fall more on the side of bitter-lesson/scaling-law-maximalism, and I think it's probably healthy and valuable that there are people in the research community placing both types of bet.

tasdfqwer0897··on Program Synthesis with Large Language Models
Unfortunately not, but we do release both the programming dataset and the math questions dataset, so in principle you could try those out with one of the open-source models from e.g. huggingFace.
tasdfqwer0897··on Program Synthesis with Large Language Models
I think it might be a mistake to think that the model is not confident because its response is something a human might say if they were not confident. The model is 'just' completing the prefix text with something that has high likelihood from its perspective, so it may just be used to, for instance, seeing people hedge in similar conversations it has read in its training data.

More generally, whether these models are well-calibrated (that is, they know what they don't know) is an important area of research. I don't have references offhand, but I think it's true broadly speaking that these larger pre-trained models do tend to be better calibrated.

tasdfqwer0897··on Program Synthesis with Large Language Models
Hey, I am one of the lead authors of this paper. Happy to answer questions. This is a twitter thread going over the main results:

https://twitter.com/gstsdn/status/1427794393373626368

tasdfqwer0897··on Symbolic mathematics finally yields to neural networks
We have been working on this recently on the Google Brain team.

We are working both on synthesizing programs from scratch (see https://arxiv.org/abs/2002.09030 for example) and on understanding computer programs using machine learning (see e.g. https://arxiv.org/abs/1911.01205).

I'm always happy to correspond with people about these topics.

tasdfqwer0897··on Open Questions about Generative Adversarial Networks
This actually might have interesting connections to ideas from differential privacy.

Maybe the work is derivative of a particular training image if we can easily predict the presence or absence of that training image given only the trained model?

tasdfqwer0897··on Open Questions about Generative Adversarial Networks
So if you wave your hands enough, it seems like maybe there's an argument to be made that the weights of a trained GAN somehow correspond to a 'compilation' of the training data as it's defined in this doc: https://www.copyright.gov/circs/circ14.pdf
tasdfqwer0897··on Open Questions about Generative Adversarial Networks
So you are worried that your existing attribution method is too focused on 'obvious' attributes and you want to see if you can make it focus on less obvious things?

IIUC, that's something that's been looked at in the ML Fairness literature. See this paper for example: http://www.aies-conference.com/wp-content/papers/main/AIES_2...

tasdfqwer0897··on Open Questions about Generative Adversarial Networks
Someone on the machine learning reddit asked me this:

> Question: How does copyright work for GAN output? If I input 300,000 copyright protected photos of celebrities and generate images of new celebrities that do not exist, are the generated images public domain or would there be copyright issues?

AFAIK, this is not a settled issue, but I'd be really interested hear what an actual lawyer thinks about this?

tasdfqwer0897··on Open Questions about Generative Adversarial Networks
Hmm, I'm not sure what you mean by applicable loss functions?

I'll answer what I think you're asking and you can tell me if I got it wrong:

There's been a lot of effort spent on coming up with different loss functions for GANs, but https://arxiv.org/abs/1711.10337 shows that, according to the metrics in problem 5, they don't really improve results compared to the original GAN loss function.

There's something called a Wasserstein GAN https://arxiv.org/abs/1701.07875 that you may have heard about, but IMO the useful thing to take away from that paper is their 'gradient penalty' technique and not the new loss function.

tasdfqwer0897··on Open Questions about Generative Adversarial Networks
Hey, I wrote this! Happy to answer questions.