If you're in 2024 and you're in some programming class making a big deal about pass-by-value versus pass-by-reference, ask for your money back and find a course based in this century. Almost literally any topic is a better use of valuable class time than that. From what I've seen of the few unfortunate souls suffering through such a curriculum in recent times is that it literally anti-educates them.
$ python3
Python 3.12.3 (main, Jul 31 2024, 17:43:48) [GCC 13.2.0] on linux
Type "help", "copyright", "credits" or "license" for more information.
>>> def x():
... v = 1
... y(v)
... print(v)
...
>>> def y(val):
... val += 1
...
>>> x()
1
A pass-by-reference language would print 2.Everything in a modern language is passed by copy. Exactly what is copied varies and can easily be pointers/references. But there were languages once upon a time that didn't work that way. It's a dead distinction now, though, unless you go dig one of them up.
If you want a specific one to look at, look at Forth. Note how when you call a function ("invoke a word", closest equivalent concept), the function/word doesn't get a copy of anything. It directly gets the actual value. There is no new copy, no new memory location, it gets the actual same memory as the caller was using, and not as a "pointer"... directly. Nothing works like that any more.
Python passes primitive types by value, out rather "as if by value", because it copies them on write.
if you modify your experiment to pass around a dict or list and modify that in the 'y', you'll see y is happily modified.
so Python passes by reference, however it either blocks updates (tuple) or copies on write (int, str, float) or updates in place (dict, list, class)
https://stackoverflow.com/questions/373419/whats-the-differe...
No, you won't.
x = {'a' : 1}
foo(x)
print(x)
def foo(z):
z = {'b' : 2}
You'll see that this prints `{'a' : 1}`, not `{'b' : 2}`. Python always uses pass-by-value. It passes a copy of the pointer to a dict/list/etc in this case. Of course, if you modify the fields of the z variable, as in `z['b'] = 2`, you do modify the original object that is referenced by z. But this is not pass-by-reference.I would sooner believe the example is showing you shadowing the z argument to foo, than foo being able to modify the in-parameter sometimes even if it's pass by value.
Because it is passing a pointer by value under the hood.
This is the part that messes everyone up. Passing pointers by value is not what passing by reference used to mean.
And it matters, precisely because that is extremely realistic Python code that absolutely will mess you up if you don't understand exactly what is going on. You were passed a reference by value. If you go under the hood, you will find it is quite literally being copied and a ref count is being incremented. It's a new reference to the same stuff as the passed-in reference. But if you assign directly to the variable holding that reference, that variable will then be holding the new reference. This is base level, "I'd use it on an interview to see if you really know Python", level stuff.
Everything in a modern language involves passing things by value. Sometimes the language will gloss over it for you, but it's still a gloss. There were languages where things fundamentally, at the deepest level, were not passed by value. They're gone. Passing references by copy is not the same thing, and that Python code is precisely why it's not the same thing.
Passing a pointer-by-value similarly seems to be avoiding the point; if you tell me Python is always pass-by-value I'll expect an object I pass to a function to be a copy & not a reference, thus not be able to be mutated, and that's not the case.
That would be a misunderstanding. It would only make sense if you think Python variables are Python objects. They are not: Python variables are pointers to objects. The fact that assignment to a variable never modifies the object pointed to by that variable is a consequence of that, and doesn't apply just to passing that variable to a function.
The important point is that it's not "a reference to x" that gets passed, it's a copy of x's value. x's value, like the value of all Python variables, is a reference to some object. The same thing applies to setting variables in Python in general:
x = {1:2} # x is a new variable that references some dict
y = x # y is a new variable that references the same dict
y[1] = 7 # the dict referenced by x and y was modified
x = None # x no longer references the dict
print(y) # y still references the dict, so this will print {1:7}
y = None # now neither x nor y reference that dict; since y was the last reference to it, the dict's memory will be freedHonestly, "x's value, like the value of all Python variables, is a reference to some object" makes me think it's more accurate to call Python pass-by-reference only.
What those values represent and how they can be used is a completely different topic. Take the following code:
x = "/dirs/sub/file.txt"
with open(x, "w") as file:
file.write("abc")
foo(x)
with open(x, "r") as file:
print(file.read_all()) #prints "def"
def foo(z):
with open(z, "w") as file:
file.write("def")
Here x is in essence a "reference to a file". When you pass x to foo, it gets a copy of that reference in z. But both x and z refer to the same file, so when you modify the file, both see the changes. The calling convention is passing a copy of the value to the function. It doesn't care what that value represents. def foo(x):
x['a'] = 1
y = {'b': 2}
foo(y)
print(y)
foo can modify the object y points to, but it can't make y point to a different object? Is that what "This is impossible in Python" is referring to?so we disagree on terminology?
in my CS upbringing, sharing the memory location of a thing as parameter was tagged "call by reference". the hallmark was: you can in theory modify the referenced thing, and you just need to copy the address.
call by value, in contrast, would create an independent clone, such that the called function has no chance to modify the outside value.
now python does fancy things, as we both agree. the result of which is that primitives (int, flot, str) behave as if they were passed by value, while dict and list and its derivatives show call by reference semantics.
I get how that _technically_ sounds like call by value. and indeed there is no assignment dunder. you can't capture reassignment of a name.
but other than that a class parameter _behaves_ like call by reference.
however if you do z['b'] = 2 in foo, then you'll see the global dict bound to x has been modified, as you have stated.
well, that's _exactly_ pass by reference.
Pass-by-reference doesn't exist in Python. Here's what it looks like in C#, which does suport it:
auto x = Dictionary<string, int >() ;
x.Add("a", 1);
foo(ref x);
System.Println(x); //prints {b: 2}
void foo(ref Dictionary<string, int> z) {
auto k = new Dictionary<string, int>();
k.Add("b", 2);
z = k;
}
Here z is just a new name for x. Any change you make to z, including changing its value, applies directly to x itself, not just to the object referenced by x.The classic example of "pass by copy-reference is less expressive" is you can't have pass a reference to number and have the caller modify it. You have to explicitly box it. I understand you understand this, but it's worth considering when thinking about whether the distinction means absolutely nothing at all.
This is really not true. Depending on how your language implements pass-by-reference, you can pass a reference to an int without boxing in one of two ways: either pass a pointer to the stack location where the int is stored (more common today), or simply arrange the stack in such a way that the local int in the caller is at the location of the corresponding parameter in the callee (or in a register).
The second option basically means that the calling convention for reference parameters is different from the calling convention for non-reference parameters, which makes it complicated. It also doesn't work if you're passing a heap variable by reference, you need extra logic to implement that. But, for local variables, it's extremely efficient, no need to do an extra copy or store a pointer at all.
That is, in code like this:
ReferenceType a = {myField: 1}
foo(a)
print(a.myField)
void foo(ReferenceType a) {
a.myField = 9
a = null
}
Whether you translate this pseudocode to Python, Java, C# (with `class RefType`), C (`RefType = *StructType`), Go (same as C), C++ (same as C), Rust, Zig etc - the result is the same: the print will work and it will say 9.The only exceptions where the print would fail with a null pointer issue that I know of are C++'s references and C#' s ref parameters. Are there any others?
Depending on the language, that is very likely the whole picture of how function calls work. In a rare few modern languages, this is not true: in C# and C++, when you have a reference parameter, things get sonewhat more complicated. When you pass an expression to a reference parameter, instead of copying the value of evaluating that expression into the parameter of the function, the parameter is that value itself. It's probably easier to explain this as passing a pointer to the result of the expression + some extra syntax to auto-dereference the pointer.
You're talking about parameters of type int; I'm talking about structs that are strictly larger than pointers. Structs which may be nested; for which deep copies are necessary to avoid memory leaks / corruption. And here, the distinction between these "mental models" exhibits a massive gap in real performance.
Here's a deliberately pathological case in C++; I've seen this error countless times from programmers in languages that make a distinction between references/pointers and values:
bool vector_compare(vector<int> vec, size_t i, size_t j) {
return vec[i] < vec[j];
}
int vector_argmin(vector<int> vec) {
if (vec.size()) {
size_t arg = 0;
for(size_t i = 1; i < vec.size(); i++) {
if (vector_compare(vec, i, arg))
arg = i;
}
return arg;
} else return -1;
}
The vector_compare function makes a copy of the full vector before doing its thing; this ends up turning my linear-looking runtime into accidentally-quadratic. From the perspective of this solitary example, it would make sense to collapse reference/pointer into the same category and leave "value" on its own.But actually these are three distinct concepts, with nuance and overlap, that should be taught to anybody with more than a passing interest in languages and compilers. I'm not here to weigh in on what constitutes a modern language, but the notion that we should just throw this crucial distinction away because some half-rate programmers don't understand it is patently offensive.
Everything (except reference types) is pass-by-value, but of course values can have wildly different sizes.
Also, the problem of accidentally copying large structs is not limited to arguments, the same considerations are important for assignments. Another reason why "pass-by-pointer" shouldn't be presented as some special thing, it's just passing a pointer copy.
Your vector<int*> is a red herring. The distinction I'm making is between passing a (vector<int>)* and a vector<int>, because those two objects have radically different sizes, and the distinction can and does create severe performance issues. And yet, pointers are still different from references: with a reference, you don't even need your object to have a memory address.
But this doesn't change the fact that they are both passed-by-value when you call a function of that parameter type.
See my cousin post: https://news.ycombinator.com/item?id=41220384
This is a distinction so dead that people basically assume that this must be talking about whether things are passed by pointer, because in modern languages, what else would it be? But that is not what pass-by-reference versus pass-by-copy means. This is part of why it's a good idea to let the terminology just die.
The fact that is a sensible thing to say in a modern language is another sign the terminology is dead.
That is, the following holds true:
int a = 10;
foo(ref a) ;
Assert(a == 100);
void foo(ref int a)
{
a = 100;
}
There's also a good chance that in this very simple case that the compiler will inline foo, so that it will not ever pass the address of a even at the assembly level. The same would be true in C++ with `void foo (int& a)`.So if one invokes low-level implementation details, Forth is also a pass-by-pointer-value in the same way as C# "ref" and others, at least on x86.
However I don't think appealing to implementation details is useful, what matters is what is observed through the language, with the results you point out.
foo: proc options(main);
dcl sysprint print;
dcl a(3) fixed bin(31) init(1,2,3);
put skip data (a);
call bar(a);
put skip data (a);
bar: proc(x);
dcl x(3) fixed bin(31);
dcl b(3) fixed bin(31) init(3,2,1);
x = b;
b(1) = 42;
x(2) = 42;
put skip data (b);
put skip data (x);
end bar;
end foo;
outputs A(1)= 1 A(2)= 2 A(3)= 3 ;
B(1)= 42 B(2)= 2 B(3)= 1 ;
X(1)= 3 X(2)= 42 X(3)= 1 ;
A(1)= 3 A(2)= 42 A(3)= 1 ;
demonstrating that X refers to A in BAR and assigning B to X copies B into X (= A). #include <iostream>
#include <string>
void s_ref_modify(std::string& s)
{
s = std::string("Best");
return;
}
int main()
{
std::string str = "Test";
s_ref_modify(str);
std::cout << str << '\n';
}
The equivalent Rust requires passing a mutable pointer to the callee, where it is explicitly dereferenced: fn s_ref_modify(s: &mut String) {
*s = String::from("Best");
}
fn main() {
let mut str = String::from("Test");
s_ref_modify(&mut str);
println!("{}", str);
}
Swift has `inout` parameters, which superficially look similar to pass-by-reference, except that the semantics are actually copy-in copy-out[2]:> In-out parameters are passed as follows:
> 1. When the function is called, the value of the argument is copied.
> 2. In the body of the function, the copy is modified.
> 3. When the function returns, the copy’s value is assigned to the original argument.
> This behavior is known as _copy-in copy-out_ or _call by value result_. For example, when a computed property or a property with observers is passed as an in-out parameter, its getter is called as part of the function call and its setter is called as part of the function return.
[1]: https://doc.rust-lang.org/book/ch04-02-references-and-borrow...
[2]: https://docs.swift.org/swift-book/documentation/the-swift-pr....
The important lesson is that assignments are by value(copy).
There are reference types in Go even though this is also not a super popular term. They still follow the pass-by-value semantics, it's just that a pointer is copied. A map is effectively a pointer to hmap data structure.
In the early days of Go, there was an explicit pointer, but then it was changed.
Slices are a 3-word structure internally that includes a pointer to a backing array and this is why it's also a "reference type".
That said, everything is still passed by value and there are no references in Go.
That's like saying C++ doesn't have references since it's just a pointer being copied around
Go is always pass-by-value, even for maps [0]:
x := map[int]int{1: 2}
foo(x)
fmt.Printf("%+v", x) //prints map[1:2]
func foo(a map[int]int) {
a = map[int]int{3: 4}
}
In contrast, C++ references have different semantics [1]: std::map<int, int> x {{1, 2}};
foo(x);
std::print("{%d:%d}", x.begin()->first, x.begin()->second);
//prints {3:4}
void foo(std::map<int, int>& a) {
a = std::map<int, int> {{3, 4}};
}
[0] https://go.dev/play/p/6a6Mz9KdFUhIt's a quirk of C++ that reference args can't be replaced but pointer args can.
And the fact that C++ references can be used for assignment is essentially their whole point, not "a quirk".