I reviewed the doc, and it is just curated list of links on various projects (ML frameworks, proof assistant, some tutorials), and does not outline plan you described here. Or I missed something?..
*Some people's terminology would classify such algorithms as "AI". I wouldn't, but they are nonetheless on-topic because you might invoke them to assist, focus, refine, or accelerate other techniques.
But there are activities where an inherently high-temperature assistant (e.g. some Instruct-inspired tune) are a great fit, and those are almost definitionally at the boundary of “objectively faithful” and “a bit stochastic”.
I’ve never done any work in novel mathematics myself, but I’m a big fan of those who do, and any study of those who do paints a picture of roughly that boundary: those who are proficient can easily spot a falsehood but only with difficulty generate inspiration for a new idea.
Of all the highly optimistic applications for 2024 LLMs, mathematics sounds pretty plausible?
Even the best funded boosters are making no such claims AFAIK, I think a much more reasonable and sober assertion might be like “they are useful to doing real mathematics”.
The word “if” in both the parent and GP is doing a lot of lifting here though, if they can do amazing feat X, then they will likewise do amazing feat Y. And I think in every instance you’ve made a credible argument.
But I don’t think we have the if yet. If I could just slip GPT-4 an executive summary of the code I meant to write and the emails I meant to send on a given day I would be doing it.
But it can’t: it’s catastrophically wrong with extreme confidence routinely. These things are easy to cherry-pick in citation but still fuck up on HellaSwag in instances a child would not.
Even though I am deeply skeptical whether AIs (as they are commonly understood today) will be helpful for hard scientific problems (for quite different reasons), I don't think this argument necessarily holds: couldn't a (hypothetical!) AI propose much better experiments and experimental designs?
Of all the highly optimistic applications for 2024 LLMs, mathematics sounds pretty plausible?
You probably already know this, but I think it bears pointing out explicitly: in this context we're not talking (only) about LLM's. There's a LOT more to AI than just Large Language Models, and the list of resources linked reflects that.
This is already being used by a few hundred scientists in a lab. We are aiming to extend to thousands next year with a focus on CS, climate & bio.
Can you share any details?
Meanwhile, this is the blog of our earliest work from January that will be published in ICML24: https://blog.allenai.org/data-driven-discovery-with-large-ge...
Also if you are interested - happy to correspond more by email. Lot more to share privately.
MIT, the Steel Factory, Cal, UDub and what, 37 authors pushed a preprint and no one heard about it?
@dang I can’t be the only one who has had it on this GPT-5 test flight shit. I don’t care how many clean Azure IPv4 blocks Altman has, throw this motherfucker out.