Untapped potential in Rust's type system
jakobmeier.ch
jakobmeier.ch
Normally this would be a compile time error, but using typeid can turn this into a runtime error.
https://docs.rs/assert-type-eq/0.1.0/assert_type_eq/ allows you to assert that two types are the same, forcing a compile time error for this scenario.
I'd rather not write that myself.
My best guess, going by a comment in the source code, is that this macro will not be necessary in the future, and therefore makes less sense to put in the standard library:
> Until RFC 1977 (public dependencies) is accepted, the situation where multiple different versions of the same crate are present is possible.
[a]: Not so much conservative, but just not caring
Or if your language supports metaprogramming (something like Nim, Zig, or Jai), you could write compile-time code that maps each of your annotated types into deterministically-determined ids. (An example of this in Nim: https://gist.github.com/PhilipWitte/dd6c670fca3baf573490)
These should be defined in the API as simple newtype wrappers over some basic binary-level type (generally either unsigned binary word or u8 array) with conversions to and from safer higher-level types defined in Rust code, leaving it to the compiler to optimize these to no-ops whenever possible. This ensures that "binary compatible" types also keep the same structural identity, and conversely, that binary incompatible types are automatically detected as well.
I basically had similar goals and implemented it with a similar design as the one described in the last part. I had the same interrogations about what should should be captured by a universal type id. In my own code, I decide to punt the discussion and require the user code to pick a globally unique name. I planned initially on using Typetag [0], a library achieving the same result as `UniversalId` from the article. The reason why I couldn't use `Typetag` in the end is the lack of support for generics. Generics are a tough problem to deal with when deriving an id because it is not clear how they should influence it.
In my case I have a common pattern where structs are defined as:
Message<S: AsRef<str>> {
message: S,
}
This allows me define methods on both borrowed (`Message<&str>`) and owned (`Message<String>`) versions without issues. When remaining inside the process, the borrowed version is passed around and when sent across the process boundaries I can deserialize it to an owned version. This pattern prevents me from using Typetag and I am still not sure how it should be solved.A related problem is also how the registry is built on the receiving end. In my project the registry is built manually, similarly to the example in the article. There are also crates such as Inventory [1] and Linkme [2]. Which allow to mark types at their definition point and then collect them in order to register them.
[0]: https://github.com/dtolnay/typetag
struct Message<'a> {
message: Cow<'a, str>,
}The more general issue is that removing generic requires the lib to pick the implementation and remove choice from the consumer. Sometimes it's not that important, sometimes it matters more. Another example I have in my project is that I have a trait describing a `User` (with e.g. `get_name`) it may be implemented as a `LinuxUser` or `WindowsUser`, each providing extra fields. How do you generate a universal id for a struct generic over a `User`?
- Some types have dynamic memory allocation (Vec, HashMaps etc...), so those would have to be computed at runtime and can change at any moment.
- Some types have shared ownership (Rc/Arc), so it's unclear how you would measure memory usage then.
- Some types, especially in foreign interfaces, will effectively just hold a pointer to some black box data, you'd need a special API to figure out how much memory it hides. For instance what's the memory usage of a database handle or a JPEG compression library context?
- When you care about memory usage things like fragmentation are usually very important, and the amount of memory used by a given object can be misleading. If you have a string that takes up 12 bytes but it's the only object left in the middle of a 4KiB page, then just counting "12 bytes" for this object is misleading because you have a huge fragmentation overhead.
The only advantage of the borrow checker is that safe Rust forces you to make ownership relationships explicit but not all Rust code is safe and there are many escape hatches that muddy the water (like Rc/Arc mentioned above, but also threads and a few other things).
I thought that's exactly the question that was being asked? Something like a htop for a rust program. This could be helpful in improving code efficiency.
> - Some types have shared ownership (Rc/Arc), so it's unclear how you would measure memory usage then.
Same way we measure the memory usage of programs running on systems that support shared libraries and mmapped data. Seeking perfection here is the obstactal; the goal is to attribute memory usage to an object (and preferably, a function call) so that we can improve its characteristics.
> - Some types, especially in foreign interfaces, will effectively just hold a pointer to some black box data, you'd need a special API to figure out how much memory it hides. For instance what's the memory usage of a database handle or a JPEG compression library context?
It's true that if you integrate with external systems, you don't get the benefits of the system you chose to make your home. This doesn't mean you should try to get as much benefit as possible.
> - When you care about memory usage things like fragmentation are usually very important, and the amount of memory used by a given object can be misleading. If you have a string that takes up 12 bytes but it's the only object left in the middle of a 4KiB page, then just counting "12 bytes" for this object is misleading because you have a huge fragmentation overhead.
This is interesting to me, and I don't know much about it. It sounds like it's not an actual limitation to the utility of the tool, but it should certainly guide how it is built and how its results are interpreted.
> The only advantage of the borrow checker is that safe Rust forces you to make ownership relationships explicit but not all Rust code is safe and there are many escape hatches that muddy the water (like Rc/Arc mentioned above, but also threads and a few other things).
In conclusion, I don't think that renders the activity pointless. The fact that hazy information is hazy and requires careful interpretation doesn't make it useless. People find test cases and static type checks useful even though they don't answer the question "is my code correct". And we're all the time relying on half measures that answer parts of the questions here (I once developed a program and used the load averages from uptime to tell me if it was performant enough :/). The pursuit of perfection is the enemy of improvement. It might be that there is some fatal flaw, but I don't think you've mentioned any here.
`loupe` provides the `MemoryUsage` trait; It allows to know the size of a value in bytes, recursively. So it traverses most of the types, and its fields or variants as deep as possible. Hopefully, it tracks already visited values so that it doesn't enter an infinite loupe loop.
We are using it inside Wasmer, a WebAssembly runtime.
Not worth the effort but it can be done in not much code