Function spaces can’t always be viewed as vector spaces because there may be no concept of addition of functions or a concept of scalar multiplication on the functions that behaves well with whatever notion of addition the functions satisfy.
> A function f from a set X to a set Y, denoted f: X -> Y, is a relation from X, the domain, to Y, the co-domain, that satisfies two properties: (1) every element in X is related to some element in Y, and (2) no element in X is related to more than one element in Y.
Susanna S. Epp. 2010. Discrete Mathematics with Applications (4th. ed.). p.384
1. When we want to consider both positive and negative square roots, we can just say that instead of the function sqrt(x): ℝ ⟶ ℝ which always gives a nonnegative real number, we are using a function full_sqrt(x): ℝ ⟶ ℝ² which gives an ordered pair of real numbers.
2. We mostly define arcsine as yielding a single value between -pi/2 and pi/2, but it's just as easy to define the equivalence relation "x ~ y when sin(x) = sin(y)" and then say that the unique value given by the arcsine function is one of the equivalence classes which that relation induces over the real numbers -- or that it is the subset of the reals included in such an equivalence class.
3. If we are concerned about graphing a multi-valued function, we can do that too; for your parabola, we can define the three-dimensional function "z(x,y): ℝ² ⟶ {0, 1} = 1 when x = y² and 0 otherwise" and the graph of this function is the parabola you want.
So you're wrong about the definition of a function, except that you're right in spirit; the way we normally think about functions, as opposed to the way we define them, definitely allows them to yield multiple values for one input.
You're catastrophically wrong about bijective functions; those are an important concept and your definition is not related to the actual one. A bijective function is one for which all values in the target set are reached by exactly one, no more and no less, value in the source set.
(In fact, this definition is incomplete: it must also be the case that the function is defined over the entire set A; each value in A maps to a value in B.† This is a common assumption to make about functions, but since I violate it earlier by talking about sqrt(x): ℝ ⟶ ℝ, I really should make it explicit. The definition I gave is motivated by the synonymous term "invertible function".)
"All values in the target set have property X" and "Each value in the target set has property X" are exactly equivalent English-language statements.
† "f: A ⟶ B satisfies ∀b∈B ∃a∈A ∀x∈A (f(a) = b ∧ f(x) ∈ B ∧ (f(x) = b ⟶ x = a))", but while that technically works, it feels kludgy.
There is no high principle that says "collecting multiple values and giving them a collective name that implies they are a single value is metaphysically superior to collecting multiple values and admitting that they are multiple values".
Maybe some kind of "it's all of them at once" mental model is useful in math, but object oriented programming gives you a hierarchal mindset, everything is in a container of some sort, so multiple return values having a container like a set makes perfect sense and just a bunch of loose values isn't a very familiar concept.
But perhaps if you're actually doing math, things are different?
Over in the programming world, Common Lisp allows you to return multiple values without wrapping them in a container. If you want the primary value, you just treat the function as if it returned one value. If you want additional return values, you use multiple-value-bind.
(I believe the general idea is along these lines: your function does a certain amount of computation, producing a set of related values. Most of the time, the caller will be interested in just one of those, which you return as the primary value. But some of the time, the caller will be interested in more than just that one value, and you have to compute the secondary values whether the caller wants them or not, so you return those too.)
For a different example of formalism in math, it is conventional to say that there are two boolean logical operators, negation and implication. You can still write about conjunction and disjunction, but everyone understands that when you write "p and q are both true", what you really meant to write was "it isn't the case that the truth of p implies the falsity of q". The point of the formalism is that you can do your proofs by considering negation and implication and then ignoring everything else. (It isn't conventional to say that there's just one logical operator NAND. You might think that would be even better, but the effort saved by only considering how one operation works ends up being less than the extra effort involved in doing proofs about NAND.)
The situation with functions is more or less the same thing; at many points we want to rely on the assumption that when a = b, f(a) = f(b). So we define functions that way, and functions that give multiple values have to be treated as giving a single composite value instead. But in a context where you have some value a and what you want to know is "what is f(a)?", the fact that the answer may consist of multiple values won't bother you.
Now, I have never seen someone take the position that boolean conjunction and disjunction don't exist as concepts just because that is how logic is normally defined, but the analogous position seems to be more popular for multiple-valued functions.
The expressive power of single value functions is very powerful and the constraint is not necessarily restrictive but may even drive a stronger analysis. (Where does this function have multiple returns? Is it for the whole domain? Etc)
By contrast the expressive power of negation and implication are relatively low given very intuitive and well defined alternatives exist.
Second, there already exist good enough paradigms for dealing with multi-valued functions. Splitting the function up into multiple functions, mapping to an ordered pair, etc.
Defining a function the traditional way is more than just notational convenience. The single value constraint allows for many simplifying assumptions, enough that is worth to pay the cost when dealing with relations that you want to talk about in functional ways
The concept of a special kind of mapping between two sets, where to any element of the first set corresponds a unique element of the second set is very important and it needs a special name.
The choice of the names is arbitrary and one could use for instance the term "univocal function" to mean a mapping like described above and "function" for any mapping between two sets.
Nevertheless, by far the most widespread convention in mathematics is to use the term "relation" for any mapping between two sets and the term "function" only for those relations where to any element of the first set corresponds a unique element of the second set.
There exists no reason for not following this convention, from which it also results that an invertible function is a function where for any element of the second set corresponds a unique element of the first set, so this convention also provides a simple meaningful name for another important concept that needs a special name.
The functions in programming languages that return multiple values, unless they return partially or totally random values (in which case they are not functions of only the input arguments, but also of an internal state or of time), are just functions that return values which belong to the set that is the Cartesian product of the types of the individual values. So the name "function" is usually correct in the mathematical sense even for such functions. If they had not been functions, the programmer would not have known what values to return, when writing them.
Moreover, I disagree that in most contexts when you want to invert a function "it isn't actually important that functions yield a single value".
In the overwhelming majority of the cases that appear in engineering and science when you want to solve equations a.k.a. to invert functions, you want to obtain a unique solution that can be directly implemented in practice. Whenever you cannot obtain a unique solution, you need to add extra criteria that allow the selection of a unique solution that is usable. Those extra criteria are actually equivalent with transforming the original non-invertible function into a function that can be inverted.
Of course the function that gives you inverse images is nothing more than an inverse function that is allowed to give multiple values.
You can’t use the word function for something that isn’t a function. It makes no sense to say inverse function that is allowed to give multiple values. Hence the need for terms like pre-image of a set.
In all contexts it is important that a function yield a single value for a given input because that is the definition of the word.
Of course, these are still functions which take exactlt one element in their input set (the reals for sqrt, or pairs of reals for square) and return for each a single element in their output sets (the pairs of reals for sqrt, or the reals for square).
It would have cost you nothing except a minute or so of your time to have refuted yourself before posting: https://arxiv.org/abs/1912.08274 https://arxiv.org/abs/1703.01700
(See also: https://en.wikipedia.org/wiki/Multivalued_function)
EDIT: Complex analysis was developed well before the formal definition of functions, relations, set theory. As such there are legacy terms in use. Mathematics has a lot of abuse of notation. For instance x=2 can mean assignment or an equation to solve. It depends on context.
In each branch of a multi-function the mapping is a function and, as a convenience, a single term “multi-function” is used to encapsulate all this. It cuts out the verbosity. A multi-function is not a function though.
My pedantry in this comes from the blog post. It was poorly written and left out some, from a mathematical point of view, important details. For instance the author uses a Schauder basis and has an uncountable sum at one point whilst ignoring the need distinguish between a Hamel basis and Schauder basis or the need to use a linear ordering in the reals to properly make sense of the uncountable sum.
They don’t think of Log(z) as a set of values. At least I don’t. I assume one means the principal branch when writing Log(z) unless specifically stated otherwise. Maybe in their mind they do think of an infinite set when seeing Log(z). I’m skeptical of this possibility though.
The insight behind multi-valued functions seems to be that for certain types of functions that seem unambiguous in the real case (roots, logarithms, ...) there's no clear candidate for which one is "the" logarithm function etc. So theorems look more like "let l be a logarithm function".
If you actually want to do complex analysis, it would not be very useful to have a vector valued function (even if that's technically possible) for logarithm etc. because most of the theory is developed for (in particular, differentiable) functions C -> C. So you'd rather just pick a branch.
Isn’t the function definition meant to be interpreted as “maps consistently to exactly one element of the codomain”? So an ordered pair of R^2 is still one element of R^2
It seems the parent has just mixed up the domain and codomain, because under that assumption he would be right about both the definition and bijectivity
The formal definition guarantees that whenever a = b, f(a) = f(b). You use it when you need that guarantee.
An ordered pair drawn from ℝ² is in some sense a single value. In another sense, it is two values. Which way you want to think about it depends on what you're going to do with it; if you're thinking about square roots of real numbers, it will be more useful to think of it as two values.
> It seems the parent has just mixed up the domain and codomain, because under that assumption he would be right about both the definition and bijectivity
He still wouldn't be right about bijectivity; you also need the assumption that a function is defined over its entire domain.
What I tryed to say is that the sense in which “exactly one element” is used in the definition of function is inclusive of codomain being R^n, so it confused me why you would provide a function that has a codomain of R^2 as something that suggests deviation from the formalism. It just seemed misleading to phrase it that way
I thought consistently would convey the idea I had, nevermind if it doesn’t
My bad about bijectivity, I see it now, you’re right
Yes, nothing stops you from defining the square root function such that it returns an element of R^2. It's perfectly legal mathematics, functions can map arbitrary domains to arbitrary codomains after all.
But I've never seen it used anywhere. Maybe there is an area where such a definition makes sense... but regular real (or complex) analysis doesn't really normally do such a kind of thing.
The problem is that you want your typical functions like square root (and logarithms, which in the complex case also aren't unique) to be functions R -> R (or C -> C). Otherwise you can't compose them and then everything is kinda weird. In particular, you don't get the vector space structure that the article is all about.
When working in a structure, one often doesn't want to leave that structure and go to another one, but keep the same structure (or at least one that is related in an interesting way, such as a quotient structure). That just leads to more interesting mathematics.
Any claims of injective or surjective properties (and by extension bijection) are moot without the domain and co-domain being specified.
To map math ↦ dev semantics: functions need to be pure, deterministic and declared with strong types. The properties of square root in particular will vary wildly based on what those types are.
You could make it a function by defining sqrt : N -> Z×Z, so sqrt(4) = (-2,2), and we're back to the property that any value in the domain (the natural numbers) is mapped to a single value in the co-domain (a unique pair of integers). Of course, this is not the traditional sqrt function that people normally use, since it doesn't obey the property (sqrt(x))² = x (since obviously taking the square of a pair of integers will not return an integer, for usual definitions of squaring).
Note that a bijective function has both this property and the opposite, that for any value in the co-domain there is also a unique value in the domain that maps to it. Sqrt, both in the traditional definition where it only returns positive roots, and in this definition where it returns a pair of roots, is actually also bijective. Squaring is a non-bijective function, though.
sqrt: C -> (C×C)/~,
where C is the complex numbers and ~ is the equivalence relation generated by (a,b) ~ (b,a) for all a,b in C.
Of course, this is not a surjective function, but it is a function nevertheless.
As an aside, the notation R x R means ordered pairs of reals. So for your definition, you would instead want 2^R as the codomain (i.e. the set of subsets of R).