yeah, the tricky thing about the experimental loop is that:
1. its very difficult to do it in a reproducible manner (the same experiment done twice often gives different results due to small undocumented changes)
2. its expensive to do at scale.
Both of these properties make it hard to hill climb on experiment. What's worked for us so far is precisely what you said - having human experts review and provide feedback. we distil their reviews into rubrics, and have LLMs act as proxy experts using these rubrics. We expect the models will hill climb using this approach, and will reach (close to) human expert level by doing this.
thanks for sharing! We think the Experiments-as-code path is not the right approach. the beauty of LLMs is their ability to ingest and reason over unstructured data - they remove the need for formalizing experiments. We tried using declarative templates to document our experiments, but realized that most of the interesting insights (for example, how viscous a liquid feels) is easier described by ranting about the experiment to a LLM, than formalizing it via constructs/code.
interesting! we haven't really considered these yet. However, we may soon have to - the recurring feedback we hear from industry is that crystallinity is a pipe dream, and that most materials are going to be amorphous (or perhaps quasicrystalline). MLIPs have lowered the computation cost for a large number of atoms/odd cell size, but it may be a while before they're accurate enough to simulate these scenarios
that's true, we've seen examples of many promising startups that are a few years in and stuck because the industry is so risk averse. will explore the other domains you've mentioned!
good points. one of the reasons we picked the semiconductor industry is that its less price sensitive than others-companies are willing to pay if the performance is there. Effort is a different story though, and definitely a tradeoff to keep in mind.
We're doing experiments ourselves now at university partner labs (UC Berkeley and Stanford), which helps us get moving quickly. At some point, we'll need a partner though - the equipment and testing process quickly get very expensive.
yeah we were surprised by how much it does it. Our approach has been retroactive - we monitor the thinking trace, spot reward hacking behavior and then fix things.
We haven't faced this issue with Sol though - its been much more well behaved
We're still figuring this out. We'll need some synthesis equipment (think CVD, PVD etc) and characterization (XRD, Raman spectroscopy) tools in-house to validate that we're making the right materials. We're considering developing these tools in-house - the models sometimes come up with clever modifications to them so that they can deposit new materials. We think equipment is as central to new material discovery as the material itself, and will probably need to be rethought to allow for high-speed AI based experimentation
There’s a variety of computational techniques that help us establish some confidence on the materials. Atomistic simulations can estimate stability and bulk properties of a new material, and we have synthesis experts (min qualification: PhD in thin film deposition) come up with rubrics on how to judge if a material/synthesis recipe is worth trying. All these approaches have known limitations, and improving them is the bulk of our work as a company!
There’s also a lot of work to be done in figuring out the minimal set of experiments required to know if a research direction/material set is worth pursuing
fully agree with this. Plus, if you're worried about "promotion competitiveness" souring friendships with people in your team, you can always make friends with people in other departments and meet them during lunch. This doesnt have to be a hard binary rule
I think the point of the comparison with chess is to show what happens when AI becomes vastly better than humans at something. GPT4 is not vastly better than humans at coding, but the author is using chess as an analogy to visualize what the world might look like if these GPTs continue to get better
agree with this. in their new values, they literally use the same words that were in the old ones ("unpretentious", "collaboration"). I also dont think this is a change "at the drop of a hat" as the article suggests - OpenAI's been through a major inflection point with GPT3.5/4, and the company's not the same one it was a few years ago. It makes sense to make updates to the company's core values
imo they dont have batching because they pack sequences before passing through the model. so a single sequence in a batch on OpenAI might have requests from multiple customers in it
yeah I think there are a lot of use cases for an assistant trained on your chat history. Given how privacy sensitive this use case is, I think maybe Apple is the best suited to build something like this? Hope they come out with something cool
Very cool. I like the introspection bit, I've realised quite a bit about my texting style from talking to Llama too. I think Im also very "type first and think later" on WhatsApp
Yeah, this is definitely a dicey ethical question. Would be interested to know what guardrails you're considering for these digital avatars, and how you'll ensure that people use them in a healthy manner and dont get dependent on them.
Yeah, this can be extended to create a "simulation game" of us and our friends. This paper (Interactive Simulacra of Human Behaviour https://arxiv.org/abs/2304.03442 ) has a setup on how we could create a Sims game with us as the characters
Yeah I wondered if few shot prompting would yield better results than finetuning. For the amount of finetuning I've done (1 epoch, 7B model with 4 bit quantization), I think it might be comparable. But if we scale this to a bigger model and longer training times, I think finetuning should produce much better results. Hoping someone with access to compute will try it out and update us!