I answered you points below, but then I gone meta and thought about discussion becoming a metaphysical conundrum and decided to go back and tried to come from another angle. And this what I came with:
The worst part of C abstract machine in Rust is lifetimes. Lifetimes becomes hard to think of, borrow checker becomes impossible to satisfy. You move value from x to y, therefore you move memory, now you do not need to track a lifetime or a scope of x, it is irrelevant, it have no memory, it doesn't exist anymore. But you need to track y, because it owns some value, which is the memory. Lifetimes of values are different from scopes of variables, but it is lifetimes that matter to a borrow checker.
And now a lot of words about how I manage to think this way. I'm not sure it is relevant to the discussion, but because I already had typed it and if you might be interested...
> Well that's literally how C works. There's an "abstract machine" that runs your code, and that's how you know what side effects it's obligated to produce. Though the C abstract machine isn't very concerned with memory locations, so it's not particularly relevant to this discussion.
Yes, thinking about it, I agree. With C it is natural to believe in a canonical representation. I thought exactly this way before I moved to rust. Possibly I have in my mind some representation of this even for rust, but I'm not sure. Oftentimes I do not think of functions as of sub-programs that must be called, because mostly they are inlined. Functions are chunks of code that are combined into one to compose a working code. I do not think of variables as of something existing in a machine code, they are compile-time abstractions allowing me to relate to ideas of how to compute their values or to the values itself (it feels like lazy-evaluation mode of thinking, and it is, but evaluations in Rust are not always lazy).
> It's often easy for the compiler to avoid allocating memory on the stack at all. But your claim was that you could have uninitialized variables with no memory allocated, then initialize them, and they would now have memory allocated.
Not necessarily. A value can live in the memory of compiler. It can be computed statically and then used without really forming a structure or something. When you move a value of variable into a Box, it can easily happen. Value was calculated, but have no representation in the machine code, but then when Box is initialized, machine code does a few "mov"s of immediate values into Box.
> Putting something in a Box isn't that. A Box has two parts, the small Box struct itself, and the memory it allocates for the contents. If you have a Box variable, it will be given space for the struct even before you run Box::new.
No, it will not get space before Box::new
struct Point {
x: i64,
y: i64,
}
impl Point {
fn new(x: i64, y: i64) -> Self {
Self {
x: x,
y: y,
}
}
}
#[inline(never)]
fn make_a_point() -> Box<Point> {
let p = Point::new(2, 3);
Box::new(p)
}
How does it translates into a machine code?
_ZN3tmp12make_a_point17h643905a77db82570E:
.cfi_startproc
pushq %rax
.cfi_def_cfa_offset 16
movl $16, %edi
movl $8, %esi
callq *__rust_alloc@GOTPCREL(%rip) ; <-- allocation, now %rax is a pointer to uninitialized Box<Point>
testq %rax, %rax
je .LBB10_1
movq $2, (%rax) ; <--- (%rax).x = 2;
movq $3, 8(%rax) ; <--- (%rax).y = 3;
popq %rcx
.cfi_def_cfa_offset 8
retq
See? There were not place to store Point { x: 2, y: 3 } before Box was created. And the only place in a program it exists as a Point (which can be addressed for example) it is a memory owned by a Box. So Point::new is evaluated
after Box::new, and then inlined Point::new initializes not the memory on the stack but memory of the Box directly. In other case compiler might pass %rax into a constructor as a kind of `this` pointer in C++, so it could initialize Point directly in the memory owned by the Box.
%rax holds "second part" of the Box, "the small Box struct itself", which is just the address of the memory owned by the Box. We can think of %rax as of variable holding the Box, though there is no corresponding variable in the source code.
> With the caveat that the compiler might try to shrink the scope where the variable exists, so you might need to do a little bit to prevent that.
Why may I want to do this? I mean, it is an important trait of Rust to avoid copying while moving. It wouldn't work if it was necessary to allocate structure on the stack before moving it into the box. Theoretically it is ok, but what about performance of the code? C++ deals with it using `this` pointer, and a special syntax for calling a constructor. Rust doesn't do it and constructor returns object "by value" and to get performance on par with C++ it is necessary for rustc to optimize it, so the machine code of the constructor was able to initialize the target memory directly.
Yeah, it is the optimization, we can probably ignore it. But... but I cannot ignore it, because it is important. I need to know that my code doesn't do easily avoidable copying. So I think of it as of part of the language itself.
> I'm aware that variables can be optimized out. But so can values! You can have a value in your program that takes up 0 memory even though it gets used to do something.
Above there is an example for this. The Point comes to an existence at some point, while variable p doesn't. So variables do not exist at run time, while values do.
> I think we should ignore optimizations like this.
Why? Just to protect a representation of an abstract machine? It mostly works with C, and so it makes sense with C. But it mostly do not work with Rust, it doesn't predict a generated machine code well. For example, in C you can rely that function will be called, one needs to explicitly inline it to avoid the call. In Rust you need to explicitly forbid inlining to get a call instruction reliably (I did it in the example above to separate the code of make_the_point from main function). In C you can rely on Point being created before a memory malloc'ed, and you can infer from that that Point must be stored somewhere, probably on the stack. In Rust as you can see it is not the case.
> It's often easy for the compiler to avoid allocating memory on the stack at all. But your claim was that you could have uninitialized variables with no memory allocated, then initialize them, and they would now have memory allocated.
It is not just that. If you "deinitialize" variable by moving value out of it, it loses its memory.
let p1 = Point::new(1, 2);
let p2 = p1;
There is just one Point allocated. Not two. In C we would think of it as of two Point-sized slots on the stack, and about value being copied from one to another. But here you should think of it as of one Point-sized slot which was named p1 before we moved it to p2, after which p1 is not associated (lexically bound) with the memory slot of the Point, p2 does.
> You have a much more narrow definition of "variable" than I do. It has a name, it can hold a value...
But it works differently. It pins field memory to the containing struct. One cannot just change association by moving:
let p = struc.point;
Rust will forbid it, we will need to use Option<Point> instead of Point as a type for the field, or to find some other way around it. Though I suspect it is harder now to learn because borrow checker becomes smarter and it allows to move out of a struct when it sees that lifetime of the struct is bound and it will not be used anymore.