I think it's also fair to say they are different organs based on the developmental physiology. Many other organs come from distinct embryonic cell lines.
2,214 karma · joined December 6, 2021
I think it's also fair to say they are different organs based on the developmental physiology. Many other organs come from distinct embryonic cell lines.
The argument that consciousness suddenly arises due to a particular configuration of atoms is also suspect; all forms of consciousness we routinely accept are fairly robust to many configuration changes. It seems to me that consciousness hinges on the discussion of degree or content, not on the material form. A soup of atoms certainly could be aware, but not aware of in the same way that "I" am aware.
The real question for Panpsychism is not whether something is or is not conscious, but conscious of what. If everything is aware, what is this thing we have that a rock doesn't seem to have (but LLMs might have)? I'm basically sold on the premise of Panpsychism, but we do have to recognize that when we say "consciousness", we aren't always talking about the same thing.
In another comment below, I likened this to clear-cutting a forest. Growing the forest takes a lifetime; destroying it could happen in the next few months.
It's not that the horizon is expanding because of this. It's more like a forest getting clear-cut.
> [I]t is now the identification of a promising problem which is the scarce and precious resource. We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential.
This. What is worth a human life's attention? As little as a month ago, mathematics was valuable in part because only a small number of people could possibly make progress on the frontier. We are confronting an existential moment for a 4000+ year-old human cultural endeavor. The assumption that "mathematical thinking is hard" has been built-in at a number of important points in how we support mathematics and mathematicians.
We need a different model, and fast. Already, the research community is feeling unable to digest proofs fast enough to keep up with the output of AI models. The paper is 165 pages, and the discovery was finalized two days ago. What this means is that nobody really understands it. Nobody would accept OpenAI's proof in this amount of time, except that they formalized it in lean. The formalization alone would normally be another years-long (or career-long!) effort if the world was the way it was one year ago.
So, again, what efforts are worth a life's attention today? It's a harrowing change.
That said... CoT monitoring is a fragile "intent discovery" mechanism; neuralese puts this problem front-and-center but if agents begin to learn to hide their intent from their CoT journals, we are basically in the same spot.
For amateur astronomers, an interesting question is whether there is a satellite above me that I can see tonight (and when / where)? For professionals, the literal million dollar questions are more like: are there any satellites on a collision course? Which satellites have moved recently and why? Is there a dime-sized piece of metal somewhere out there that could hit my (employer's) satellite?
I expect that building the satellite viewer is now just a few dollars worth of tokens. I built something like this a few years ago, and it took a week of work. It's cool that the marginal cost of satisfying your curiousity about where things are in space has dropped so close to zero, but don't be fooled into thinking this website lets you in on some vast conspiracy.
I say this because it seems that earlier announcements where industrial deep neural nets "outperformed NOAA" likely encouraged the slash-and-burn Trump administration in its gutting of critical activities and centers of expertise at NOAA. The impression that industry can predict weather better than the government agencies totally misses that the industrial models utterly rely on government data for inputs. In fact, almost all weather reports you see---weather.com, TV, etc.---are just lightly repackaged products that NOAA provides for free on weather.gov (which you can access for free without ads).
They are likely deliberately avoiding the SoTA race for a few reasons:
1. Their best models are marginally better than current SoTA releases. 2. They'd like to let Ant/OAI make mistakes with safeguards / let them get the regulatory heat. The unknown unknowns are huge with SoTA models (eg OAI accidentally hacking huggingface) and they are protecting their reputation. 3. They want to encourage companies to become cost conscious because they can likely win on price in the long run. Getting market share in "quantity beats quality" workflows forces companies to establish processes to choose the "cheapest acceptable model", which is a good environment for Google.
In the other story, the current understanding of optimization is a natural evolution of past work, where a new generation of researchers respond to social and technological changes, adapting and building on the work of the past, taking what's useful, downplaying the importance of some ideas, and inventing new language to describe concepts that seem most relevant to the current situation.
Both stories tell some of the truth. A revolution or evolution? Looking at the literature (eg the sibling comment here) shows that even today, convexity is used as an intuition pump for modern optimization techniques. But there are also new ideas that apply to the specific exigencies of neural nets, and downplayed ideas (eg convergence rates) that seem less relevant.
I don't think this changes the point, which is that most optimization methods used in AI owe a substantial intellectual debt to convex optimization theory.
If you have more specific feedback on what you found distasteful, I'd be happy to hear it.
I think that Nesterov's first order method is the most efficient general first order algorithm on convex problems, so anything else is in some sense worse. (Edit: removed incorrect ADAM comment.)
The fact that neural networks are highly nonconvex has encouraged a lot of research, but it's more of the kind aimed at resolving tension: these methods are probably good for convex functions, why do they continue to work for nonconvex problems, and are there tweaks we can make to improve them in that setting? It's not a lot of de novo theory; more standing on the shoulders of giants, etc etc.
You want to know how long it takes to solve an optimization problem, in this case over convex, lipschitz functions. (The restriction to a spherical domain is not really a restriction, you can just change variables for any bounded domain.) Anyway, showing upper bounds on time complexity is "easy" because it's just the runtime of your algorithm. Showing (nontrivial) lower bounds is usually much harder because it requires constraining all algorithms.
This proof apparently shows that the lower bound time complexity is equal to the time complexity of an existing 30-year old algorithm: it requires Omega(d^2) function evaluations to solve over this class of functions.
My gut says likely implies that d is the minimal number of evaluations if you have a gradient oracle because you can approximate a gradient with d function evaluations, but I'm not sure how hard it is to make that rigorous.
Being a bit more humble, perhaps the lesson is that the difference between the theory and reality highlights externalities that always exist in the real world that make the theoretical model miss a crucial piece of the real world. It's logically correct in some sense, but incomplete.
A 3.5-sigma event is a 1-in-1000 chance. Yet we have 2 of them in our data set that goes back only 40ish years? Something is suspect. Maybe I'm not accounting for the likelihood of a random walk excursion probability, but if a 3.5 sigma event isn't a 1-in-1000 event I have difficulty interpreting.