What P vs. NP is about
vasekrozhon.wordpress.com
vasekrozhon.wordpress.com
https://www.righto.com/2016/12/die-photos-and-analysis-of_24...
I have to say that I find these pictures extremely aesthetical. It's crazy how the brutally powerful CPUs we have today are successors to processors created not so long ago that also, though very intricate, fit into one picture...
Interesting post but not sure why the author calls it underrated. It seems to get a fare share of attention and love amongst those that care about such things.
Edit: presenting the problem with inverting functions is an unusual approach tho.
https://stackoverflow.com/questions/1857244/what-are-the-dif...
If you can solve the decision problem, you can solve the optimization problem by doing a binary search on the x.
That means that tsp and tso-opt are in the same ballpark of complexity. The reason tsp-opt isn't called np-complete isn't related to its complexity, but to the fact that only decision problems get to be called that.
As I said, there are problems where the decision problem and the optimization problem are completely unrelated in terms of complexity, but tsp isn't one of those.
Assuming you have an oracle for any question that you know how to verify, this is very easy to reduce:
Is there a solution in which vertices {1, 4, 6} are all positive?
First, let's establish that the question is legal: presented with a working solution, we can calculate the defining constraint of the problem, and we can check the polarity of vertices 1, 4, and 6. If this problem is in NP, verifying the constraint will take at most polynomial time. If not, not, but our question can be verified in however much time it takes to calculate the constraint.
The complexity of determining a solution by getting answers to questions of this form is interesting.
Case 1: The answer is "no" for all single-vertex sets. All vertices must be negative. This cannot actually happen, because all vertices being negative is the same thing as all vertices being positive.
Case 2: All vertices may be simultaneously positive. In this case, the answer to every question will be "yes", and we'll ask one question for every vertex in the input graph, solving the problem in sublinear time.
Case 3: We need some positive and some negative vertices. By asking about single-vertex sets, we can quickly identify a vertex that can be positive, vertex A.
We will include that vertex in every subsequent question. By the time we get there, we will have already asked about zero or more other vertices and been answered "no". We need never ask about those vertices again; they have to be negative and if we include them in any set, we'll get another "no".
So, let's ask about a never-before-examined vertex, vertex B. Can {A, B} be simultaneously positive?
If not, we can toss B into the "must be negative" pile, since we know A can be positive.
If so, we can include it in our developing solution; future questions will include vertices A and B. We still never need to reexamine any vertex that has ever been included in any of our questions.
So, unless I'm missing something, in this case we'll also produce a solution using one question per vertex in the graph.
> but for this one I don’t see the point making the distinction.
The distinction between NP-complete and NP-hard has nothing to do with the distinction between yes/no questions and open-ended questions. Those are unrelated concepts.
An NP-hard problem is at least as hard as any NP problem. It might be much harder.
An NP-complete problem is NP-hard, but it's also guaranteed to be an NP problem. That's the distinction.
My point being that transforming tsp-opt to tsp only adds a polynomial factor. So in the case of tsp-opt, it's called NP-hard rather than NP-complete because it's not a decision problem, not because it's not equivalent in terms of time complexity as a problem in NP.
People simply call it NP-hard because that term is better known than NPO or NP-equivalent.
Both NP and NP-hard are defined for the class of decision problems.
You may review the following to clarify the distinction between the various NP complexity classes:
https://en.wikipedia.org/wiki/NP-hardness#NP-naming_conventi...
>A decision problem H is NP-hard when for every problem L in NP, there is a polynomial-time many-one reduction from L to H
Under the assumption that the <= question can be answered in polynomial time, this entire algorithm will also complete in polynomial time.
What problem formulation do you have in mind?
If you want to represent the problem as a bunch of points in a Euclidean plane with free travel, that's a different (and easier) problem.
And even then, while it might take an infinite number of steps to specify the answer at infinite resolution, it will only take a finite number of steps to specify the answer at any level you're capable of writing down.
That’s basically the good old “everything is O(1) because int64/float64 has 2^64 possibilities” misconception. It’s not how complexity theory works. For any N, 2^N is still finite, we’re obviously not talking about undecidable problems.
Edit: I see that you may be responding to the “finitely enumerate” part of my comment. Sure, it wasn’t phrased well. Replace with “enumerate in P”.
What are you fixing? Suppose you want to know the answer to within one part in 10¹⁰⁰. That will take you 333 questions.
Suppose you don't actually need 100 decimal places of the answer, or more likely that even if you had them you'd be unable to use them, and you can only represent the answer to 20 decimal places. That will take you 67 questions.
You can easily enumerate this answer in P. The problem in your argument isn't that float64 only has 2^64 values. Use as many bits to hold the answer as you want. No matter how many that is, it will be a finite number, and you'll be able to specify them all in a polynomial amount of time. Each question takes polynomial time to answer and fills one bit of the solution.
I'm having some trouble with this. As far as I can see, finding a path of a given maximum length must be in NP, because it's very easy to deterministically verify that a given path has length no more than the maximum.
If determining the optimal path length is in NP, and identifying a path of that length is also in NP, how can determining an optimal path fail to be in NP?
> Hodge conjecture
I for one have no idea about topology in general. I had to look up what a nice shape was... I thought it was some mathematical term.
The only thing I remember about is topology is someone telling me how it would be easier for me to untangle cables if I knew topology but alas I never learned.
My own pet theory is that it's an evolution of speech towards "hardened" claims that you can't easily disagree with. You cannot disagree with "over/under-rated" because there's no official "rating" method. You cannot disagree with "vibes" because the source of a vibration is impossible to determine. They both allow a speaker to make a claim without evidence.
It was ever thus.
See, for example, https://blogs.illinois.edu/view/25/96439
Though I guess for a lot of young people it feels like that, with none of their friends talking about such bands.
There's another term, i'm sure, because politicians use it all the time; "I never said X" when it was very heavily implied, and the listener was led down the garden path to the conclusion, only to find out the conclusion is bitter and unpleasant. The speaker can say "oh, that's your own biases/misconceptions/dogma, i never said <something extremely specific>"
I'm in the weeds here, but a simple analogy would be: "I don't like sunny days without clouds or fog." and i say "why don't you like blue skies?" and they say "i never said that" when the salient points of a sunny day with no clouds or fog are 1) there's a sun, and 2) the sky is blue.
contradictory reply: "not at all, they were Grammy-nominated and had N number-one hits",
weasel reply: "that's not what I meant... lately they haven't been paid enough attention. way to miss the point, $NAME_CALLING".
It's a way to say "I like them" without having to refute anything other than supporting comments, aka a weasel-word.
(A) That's not going to stop anyone; (B) the claim can easily be obviously false. For example, Taylor Swift is underrated.
We have here an example of problem (B); P versus NP is possibly the single most famous problem in computer science. It isn't underrated.
(B) anything can be "obviously false" under contrived/extreme situations. try again with "Sting is underrated" and suddenly you can easily argue it either way. Try using an argument that doesn't rely on p-values being at the 3+ sigma levels, but of course you won't because you know that will invalidate your position.
"I think that the framing with inverting a function f is quite underrated."
Of course, using the title "What P vs NP is actually about" is a bit of a stretch -- "Interesting take on P vs NP" would be more honest.
I think it's just that the first time you use underrated, it seems to be referring to the problem itself, not to your specific take. With the video as context, or even just reading a few sentences in, the meaning is clearer.
1) With others, I work on high-effort Youtube videos, our channel's called Polylog. We recently made a video about an underrated take on P vs NP. [1]
2) After working on the high-effort video, I always write a low-effort stream-of-consiousness-style very-much-unedited blog post where I clarify some details or explain some more esoteric stuff that did not make it into the video. The blog post you are discussing is such a blog post. So you should watch the video first and read the blog post only if you really liked the video and want to know more. :)
3) If you have any questions, I'll be happy to answer them!
The BBC also had him on (the Today programme iirc but it was a while ago) the one time to comment when someone claimed to have proved it.
For example, let's say we invert a checker algorithm and "run it backward from YES". Does that find an arbitrary single solution? Does it find all solutions? What if there are an infinite number of solutions?
In other contexts, what is meant is g : B -> 2^A such that x is in g(f(x)) for all x and f(x') = f(x) for all x' in g(f(x)). Here, g is said to produce the pre-image under f, and is a generalization of the above definition of inverse for non-injective functions.
In other contexts still, what is meant is g : B -> A such that f(g(y)) = y for all y such that there exists some x with f(x) = y. This is a different generalization called a one-sided inverse. (People probably call it a left inverse or right inverse, but I can never remember which is which because it emerges from an arbitrary notational convention.)
The article uses inverse in this third sense. "Find some input to f which would yield a given output y, if one exists."
Disclosure: I was that kid. My program was a mess, but good enough for a proof of concept.
I incorrectly assumed it would have some basic linear algebra solution because of how simple the problem seemed.
> To me, the essence of deep learning has nothing to do with trying to mimick biological systems or something in that sense; it’s the observation that if your circuits are continuous, there’s a clear algorithmic way of inverting/optimizing them using backpropagation.
> From the perspective of someone used to algorithms like Dijkstra’s algorithm, quicksort, and so on, this declarative approach of thinking in terms of loss functions and architectures, rather than how the net is actually optimized, sounds very alien. But this is how the whole algorithmic world would look like if P equaled NP! In that world, we’d all program declaratively in Prolog and use some kind of .solve() function at the end that would internally run a fast SAT solver to solve the problem defined by our declarations.
Relatedly, probabilistic programming was originally imagined pretty much like your second quote: you define a model, get some data, run them both through the built-in inference engine, and you get the parameters of the model likely to have produced the data. In practice though, there's no universal inference engine that works for everything (some people disagree, but they're NUTS ;) I guess pretty much for the same reason P is probably not equal to NP.
Yep, in particular there are classes called #P and PP that are closely connected to NP that can capture the hardness of problems like computing partition functions, sampling from posterior distribution and so on.
Backpropogation is a way to optimise a neural network. You want to know how best to nudge the weights of the network to optimise some loss function, so what you do is compute the gradient (ie partial derivative) of that function with respect to each of the weights. This allows you to then tweak the weights of the function so your network gets better at whatever task you're trying to get it to learn. See [2] to understand how this works and and [3] to understand how this relates to the Jacobian, but generally if you're trying to go "downhill" in your loss function it's easy to see intuitively that knowing which way the function slopes (ie the effect of tweaking each of the weights) is important and that's what the Jacobian tells you.
The inverse of a matrix[4] and its transpose[5] are two different operations in linear algebra. Transpose turns rows into columns and columns into rows and the inverse of a matrix is a little harder to grasp maybe without background, but you could think of multiplying one matrix by the inverse of another as a little like division (since you can't actually divide matrices).[6]
[1] https://en.wikipedia.org/wiki/Jacobian_matrix_and_determinan...
[2] https://www.youtube.com/watch?v=Ilg3gGewQ5U
[3] https://www.youtube.com/watch?v=tIeHLnjs5U8
[4] https://math.libretexts.org/Workbench/1250_Draft_4/07%3A_Mat...
[5] https://math.libretexts.org/Bookshelves/Linear_Algebra/Funda...
[6] algebraists please don't shoot me for that.
This is also not true.
If you want to be picky, it's true that the direct analogue of continuous optimization would be discrete optimization (integer programming, TSP, etc) rather than decision problems like SAT. But there are straightforward reductions between the two so it's common to speak of optimization problems as being in P or NP even though that's not entirely accurate.
The section was written horribly -- while I was talking about backpropagation, I was thinking about "what can be done in polynomial time" and there's a mismatch as you explained. Thanks and shame on me, I rewrote it.
In any case, I would recommend watching the video first, the article is just accompanied stream of consciousness for those few who really liked the video.
wat?
P vs NP is one of the foundational open problems in computer science.
But, of course, P vs NP is clearly central to computer science and math for a bunch of other reasons, hopefully, our video provided some intuitions.
It was very similar to a previous unsolved problem in a "haha, history repeats itself" sort of way.
This comment might have a lot more comments someday.
If you have a genuine interest, and especially if you could provide feedback or warm introductions to someone who can, I'm happy to share.
You can email me at atmanthedog at Google's free email service domain. Put something HN related in the subject. *If you do, please share your mathematical background so I can better tailor a response.
Eagerly awaiting your published proof
[1] https://www.claymath.org/millennium/p-vs-np/ and especially https://www.claymath.org/wp-content/uploads/2022/06/pvsnp.pd...