> On both datasets, we find that synthesis performance scales log-linearly with model size
This "scaling law" observation is starting to become a major trend in NLP [1] and other other modalities such as speech recognition [2] and protein structure/function prediction [3]. Do you have any insight or commentary to offer regarding the next step in improving program synthesis? For example, will techniques to scale up model size continue to be a primary focus (eg as in [4]), or do you see improvements in Transformer and attention based architectures as essential for pushing the limits of what has been achieved by you and your colleagues?
> We find that even our best models are generally unable to predict the output of a program given a specific input.
What do you think about leveraging unsupervised training to improve program synthesis? Could synthesized programs be executed on generated input in a way that supports contrastive learning [5]?
Thanks in advance for your time and comments here.
[1] Scaling Laws for Neural Language Models https://arxiv.org/abs/2001.08361
[2] Pushing the Limits of Semi-Supervised Learning for Automatic Speech Recognition https://arxiv.org/abs/2010.10504
[3] ProtTrans: Towards Cracking the Language of Life’s Code Through Self-Supervised Learning https://arxiv.org/abs/2007.06225
[4] Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity https://arxiv.org/abs/2101.03961
[5] A Simple Framework for Contrastive Learning of Visual Representations https://arxiv.org/pdf/2108.00587.pdf