Integer multiplication below n log n
github.com
github.com
"By how much?"
"2⁻¹⁸² off the exponent."
"Go back to bed."
We're in full vibe-code mode at work, so I understand both how powerful frontier models can be and how often they can over-confidently state subtly (or not so subtly) wrong things, even when you're taking great efforts to try to keep that from happening.
So without a Lean development or extensive human verification, I guess I'm a little bit skeptical, and even sort of hoping this is wrong - not just because of my not so positive feelings about AI, but by my disposition towards beauty in math. n log n is an awful lot nicer than what we have here.
But in regards to beauty, i feel like multiplication already has a lot of non beautiful exponents. Best known matrix multiply is O(n^2.371). For integer factorization, the inverse of this problem, general number field sieve is a crazy subexponential.
If factorization is just barely subexponential, is it really that surprising that multiplication is just barely sub n lg n ?
1. https://github.com/openai/math/blob/main/preprints/Matrix-Mu...
This also reaffirms my (wishful) thinking that if there’s a way to do FTL communication it’ll be something with an absurdly tiny factor like 2^-182 with a slight asymmetry in a probability somewhere.
Then you’re not violating FTL, just gaining a very slight chance that you might know something FTL – probably.
From that angle, beating light speed by some absurdly tiny factor would probably correspond to a means of predicting the future at some almost absurdly tiny factor better than random guessing.
Edit: Actually...it doesn't make sense to call this FTL communication, it's just predicting the future state of a system given some previous state. FTL comms would have to be predicting the future state of a system without information about the previous state.
Practically speaking predictive modeling would be a means of compensating for light speed comms, kind of like branch prediction in processors or speculative decoding in LLMs, but that wouldn't actually be FTL comms.
It’d likely involve exponentially more energy as well. It’d be a good sci-if plot point if FTL communications required machines the size of Jupyter to get a few milliseconds of prescience.
If you know what will happen in one minute, write down the message you see yourself writing down in one minute. In a minute, do the same thing. Now you can pass messages back two minutes.
Wouldn't faster than FTL mean that that's the speed of causality? If I send a message from a light year away telling you to "jump up and down" and you receive it in 6 months doesn't and you jump up and down, then doesn't that mean that the speed of causality in that cause was actually 2*C?
So, yes, we could've measured c wrong. We just would have no idea if we did.
Source: Veritasium did a very fascinating video explaining this problem.
We even use the Lorentz transform, developed for post-Michelson aether theories, precisely because you cannot tell whether you’re in a varying aether or a varying geometry.
We define distance and time relative to c, so we cannot measure it in the usual sense.
We know we've got it pretty well close to accurate, for the bulk of the observable universe.
If c changes, then chemistry changes, and stuff like hydrogen absorption lines shift. This is how scientists have looked for changes in c in the early universe.
Causality is conceptually out of time, so it's not traveling. Instead we except causality to operate everywhere anytime uniformly, and all physical dimensions to be bound by causality.
No way this is the correct upper bound, and I imagine it'll get refined fairly quickly. IIRC, the GapCVP results were released with a 1/n^400 complexity term, but people quickly got it down to 1/n^8 by more careful accounting.
sqrt(n) now https://github.com/Mira-acc/cvp
The fact that you can in principle go faster than n lg n, even if just by an almost imperceptible amount, is kind of surprising. It raises the question of, if n lg n isn't the limit, what is? How far down can we get the speed? If we can get it a little past n lg n, maybe we can go a lot further.
[or at least that is my understanding. not a theoretical computer scientist]
"We give a deterministic algorithm that computes the discrete Fourier transform at every length n in O ( n ( log n ) ** (1 − 10 ** −13)) operations. The model uses exact complex arithmetic, unrestricted coefficients, specified Fourier roots, and unit-cost logarithmic-size indexing; scalar preparation and array organization are included."
Multiplication can be done by FFT, which is an exceptionally efficient algorithm that has nlogn complexity, you'd be called crazy if you claim you have something more efficient than FFT.
i will NEVER care about proposed multiplication speedups unless they are truly generalized
It's most interesting when the lower bound can actually be proven. In lack of that, we have to guess what the best possible algorithm might yield (generalized or not). This tells us that need not be O(n log n) and we have the opportunity to still find better algorithms than we typically thought would be possible. This does the latter, which is interesting, but it just leaves us to hunger more for what the real limit must be :).
Your issue is that I am viewing this proof as what it really is in terms of progressing the field and not from an imaginative perspective. I think that it is important to ground our selves somewhat in reality when discussing research like this because at the end of the day open ai is not doing for fun either.
openai wants to show the world what their product can do and i am simply not impressed
OpenAI. An explicit power saving for the exact discrete Fourier transform.
Here's a random excerpt:
8.3 The middle transform and the final permutation The factor QFt in (35) can be computed from a cyclic convolution and two pointwise phase multiplications. The chirp identity below performs the frequency change in Q without applying Q as a separate permutation of the array. The second identity shows how the retained source permutation R cancels when computing a convolution. Here ∗ denotes cyclic convolution on the product of the coordinate groups and a dot denotes coordinatewise multiplication.
There are also a bunch of other stinkers, like building a Turing machine out of Navier-Stokes fluids -- except that it only works if you can encode literally infinite amounts of data in the relative positions of two particles. I.e. assuming physics is based on set-theoretic real numbers, something we've known is wildly false for over a century: https://en.wikipedia.org/wiki/Banach-Tarski_paradox
Just like vuln reporting, the AI industry has put zero effort into triage here, and the models are really good at making their findings sound more important than they really are.
The cynic in me suspects this is a smokescreen for the Navier-Stokes tokenstream plagarism fiasco.
It’s not meant to be an engineering improvement, it’s an important theoretical result because it casts doubt on things that were previously thought to be impossible.
So yeah, among the few physicists who understand enough set theory to comment on this, most of them are totally fine with admitting that our reality is not based on set-theoretic real numbers.
For mathematicians, sets are the base reality and everything else is constructed in terms of those. For physicists experimental validation is the base reality, and if it disagrees with set theory that's okay with them; it just means that the set-theoretic real numbers aren't the correct real numbers for doing physics. They're very pragmatic when it comes to mathematical foundations, and that's probably the right approach to take for doing physics.
There have been a few attempts to formalize real numbers that more closely match our observed reality; look into measurable sets. Most of these require a large cardinal, so they go beyond ZFC.
Sharing AI Progress in Mathematics
Math is incredibly rich, and even the simplest things have insanely complicated structure when you zoom in. However this all ends up, math is bigger than LLMs, and the people who claim it is getting "solved" and we are running out of open problems haven't stared into the abyss enough.
Also, it still seems that AI has a much different style from humans, with more brute force and using obscure literature results, and the future might still end up human/AI complementary. We aren't in an AlphaZero situation where the AI learns everything through self-play. (Yet? But we don't even seem to be moving that way much? Can anybody qualified help out?) Things are just moving really fast now and it's hard to process everything.
Is that true? Look at the average math paper on arxiv. Of course it's obscure to those not in the exact sub-field - math has become very specialized.
I'm not very familiar with this kind of low-level theoretical computer science, so please forgive my possibly stupid question: is it common to use a model of computation like this, with registers which can store exact complex numbers? Since no computers like that really exist, would that limit the practical application of this finding?