4,593 karma · joined April 21, 2010
http://mikeknoop.com https://x.com/mikeknoop
The new and surprising thing about test-time training (TTT) is how effective it is an approach to deal with novel abstract reasoning problems like ARC-AGI.
TTT was pioneered by Jack Cole last year and popularized this year by several teams, including this winning paper: https://ekinakyurek.github.io/papers/ttt.pdf
This is fair critique. ARC Prize's 2024 messaging was sharp to break through the noise floor -- ARC has been around since 2019 but most only learned about it this summer. Now that it has garnered awareness, it is no longer useful, and in same cases hurting progress like you point out. The messaging needs to evolve and mature next year to be more neutral/academic.
One big update since June is that progress is no longer stalled. Coming into 2024, the public consensus vibe was that pure deep learning / LLMs would continue scaling to AGI. The fundamental architecture of these systems hasn't changed since ~2019.
But this flipped late summer. AlphaProof and o1 are evidence of this new reality. All frontier AI systems are now incorporating components beyond pure deep learning like program synthesis and program search.
I believe ARC Prize played a role here too. All the winners this year are leveraging new AGI reasoning approaches like deep-learning guided program synthesis, and test-time training/fine-tuning. We'll be seeing a lot more of these in frontier AI systems in coming years.
And I'm proud to say that all the code and papers from this year's winners are now open source!
We're going to keep running this thing annually until its defeated. And we've got ARC-AGI-2 in the works to improve on several of the v1 flaws (more here: https://arcprize.org/blog/arc-prize-2024-winners-technical-r...)
The ARC-AGI community keeps surprising me. From initial launch, through o1 testing, to the final 48 hours when the winning team jumped 10% and both winning papers dropped out of nowhere. I'm incredibly grateful to everyone and we will do our best to steward this attention towards AGI.
We'll be back in 2025!
We still have a long way to go for the grand prize -- we'll be back next year. Also got some new stuff in the works for 2025.
Watch for the official ARC Prize 2024 paper coming Dec 6. We're going to be overviewing all the new AI reasoning code and approaches open sourced via the competition [3].
[1] https://deepmind.google/discover/blog/ai-solves-imo-problems...
Curiosity is the first step towards new ideas.
ARC Prize's whole goal is to inspire curiosity like this and to encourage more AI researchers to explore and openly share new approaches towards AGI.
One core idea we've been advocating with ARC is that pure LLM scaling (parameters...) is insufficient to achieve AGI. Something new is needed. And OPs approach using a novel outer loop is one cool demonstration of this.
Ryan's work is legitimately interesting and novel "LLM reasoning" research! The core idea:
> get GPT-4o to generate around 8,000 python programs which attempt to implement the transformation, select a program which is right on all the examples (usually there are 3 examples), and then submit the output this function produces when applied to the additional test input(s)
Roughly, he's implemented an outer loop and using 4o to sample reasoning traces/programs from training data and test. Hybrid DL + program synthesis approaches are solutions we'd love to see more of.
A couple important notes:
1. this result is on the public eval set vs private set (ARC Prize $).
2. the current private set SOTA ~35% solution also performed ~50% on the public set. so this new result might be SOTA but hasn't been validated or scrutinized yet.
All said, I do expect verified public set results to flow down to the private set over time. We'll be publishing all the SOTA scores and open source reproductions here once available: https://arcprize.org/leaderboard
EDIT: also, congrats and kudos to Ryan for achieving this and putting the effort in to document and share his approach. we hope to inspire more frontier AI research sharing like this
> We found that humans were able to infer the underlying program and generate the correct test output for a novel test input example, with an average of 84% of tasks solved per participant
We are also trialing a secondary leaderboard called ARC-AGI-Pub that imposes no limits or constraints. Not part of the prize today but could be in the future: https://arcprize.org/leaderboard
Relevant to OP, all platform auth creds are secured and not exposed after initial setup.
Separately, we just launched custom actions which allows anyone to extend existing public apps on Zapier with new endpoints (https://help.zapier.com/hc/en-us/articles/16277139110157-Cre...)
"We support standard oauth butttt..."
Super excited for this. Tool use for LLMs goes way beyond just search. Zapier is a launch partner here -- you can access any of the 5k+ apps / 20k+ actions on Zapier directly from within ChatGPT. We are eager to see how folks leverage this composability.
Some new example capabilities are retrieving data from any app, draft and send messages/emails, and complex multi step reasoning like look up data or create if doesn't exist. Some demos here: https://twitter.com/mikeknoop/status/1638949805862047744
(Also our plugin uses the same free public API we announced yesterday, so devs can add this same capability into your own products: https://news.ycombinator.com/item?id=35263542)
Random memory: using infinite gold hacks and the suicide dagger (which would the drop gold for other players) over and over hour hours inside the throne room or the castle west of Burger Refuge.
One thing I’m particularly proud of was figuring out how to build an Assembly program workaround for TI’s “testguard” (a tool to force clear someone else’s calculator memory via link cable) in middle school: https://mikeknoop.com/upload/SafeGuard.zip