Rust: 128 bit integers preparing to be released
github.com
github.com
The current release of stable Rust is 1.13, beta is 1.14, nightly is 1.15.
New features land on master, hence 1.15. So this means that, if it does land, it will become available on nightly soon. But it's still behind a feature flag. So it won't actually come out in Rust 1.15.
I am not on the relevant subteam here, so I am not 100% sure of the procedure, but usually, stuff has to sit in nightly for a full release cycle to be eligible for stabilization. So if this lands tonight, it'll be in nightly 1.15, 1.16 will be its full cycle, and it'll be elligible for release in 1.17, which would be March 16, 2017.
Does this cause any problems in the Rust view of the world? Will there be unanticipated issues for developers who have been getting along just fine so far unknowingly assuming that all operations are atomic, which they have been just by virtue of how the hardware works?
LLVM supports arbitrary length integer types, regardless of the platform. Rust already has u64 and i64 on 32bit architectures.
But even if something fits in one register, a variable isn't thread safe anyway.
Rust DOES NOT allow a mutable reference to exist simultaneously with non-mutable ones.
For thread safety, you always need special types, which Rust provides:
True, this won't come up in a run-of-the-mill program, but if Rust is going to supplant C for OS work, or DataBase implementations, this is an issue to be aware of.
Succinctly: If you're in unsafe code, using 128-bit integers, you may experience behavior you've never seen before.
The compiler would force you to use an atomic. Once you realize no AtomicUint128 atomic exists, you may try to explicitly use unsafe to make things work, but even then that requires a lot more explicitness than just an unsafe block. You will realize your mistake some point here.
Unsafe in Rust doesn't just turn off all the checks. It gives you the power to circumvent the checks, but these circumventions are still individually explicit.
Anyway, u128 doesn't worsen this situation. You could already make thread safety mistakes with wider value types if you really really wanted to, just like u128.
The good thing about Rust is that you actively have to circumvent the languages safety checks with unsafe code to run into those problems.
And if you are using unsafe code, you should be aware of such low level considerations anyway, and would refrain from using types when not appropriate.
The consistency problems you describe apply to any datastructure after all.
It's up to the synchronization primitives to ensure memory barriers are executed, not the language. So e.g. a Mutex<T> would execute a memory barrier when locking and unlocking. Rust then allows you to pass ownership in a trusted manner only, i.e. through those primitives.
Of course, you can also create an unsafe block and say: "Rust, trust me, I know what I'm doing, don't bother me, and these are the invariants I want you to enforce in safe code..." Mutexes are implemented this way.
But what I was asking was at runtime, when ownership transfers how does Rust ensure that writes to the value by the previous owner appear to the new owner before the transfer of ownership appears to the new owner?
You can statically guarantee that only one thread owns the object, but you can't statically guarantee the order in which the processor will apply the instructions your compiler generates, without barriers.
But the other person answered - you need to ensure that there is an explicit memory barrier yourself when you transfer the object.
To add to/clarify this: I don't think you can transfer an object to another thread in safe Rust code without a primitive that will handle the barriers for you. Static ownership tracing doesn't actually know what threads are, because it doesn't even need to.
The channels do this.
To be clear, you need to write the code yourself, but the compiler won't let you transfer ownership between threads without doing so.
There is no way to compile your code without proving to the compiler that you're data race free.
The language itself doesn't know about memory barriers. It gives you the tools for enforcing thread safety provided your synchronization abstractions have memory barriers in the right place.
So, if you wanted an atomic type, you'd find it here: https://doc.rust-lang.org/stable/std/sync/atomic/index.html
In other words, no, it shouldn't be a problem, at least in my understanding.
That's already the case: simply reading a 64-bit integer in a 32-bit processor is not atomic, and that's already possible today.
But that's not a problem in Rust. How would you access the same 128-bit integer from more than one thread?
- You pass the thread a copy of the value: the thread's copy can't be accessed by any other thread, so whether or not it's atomic doesn't matter.
- You pass the thread a mutable reference (&mut) to the value: one of Rust's rules is that there can be only one &mut to a memory location, and while that &mut exists, nothing else can read or modify the value. Therefore it's the same case: only one thread can access the value, so whether or not it's atomic doesn't matter.
- You pass the thread a non-mutable reference (&) to the value: another of Rust's rules is that while any non-mutable reference exists, nothing can modify the value, so the reads not being atomic don't matter.
- You wrap the value in something like a Mutex: the Mutex has a lock which prevents concurrent accesses in the middle of a write.
- You use an Atomic version of the 128-bit integer: this one doesn't exist, so can't be used.
The main reason for all the interest in Rust is that the compiler protects you from many kinds of mistakes. Accessing the same variable from more than one thread, without using a lock or an atomic, is one of the things the compiler protects you from.
The entire basis of the memory protection model of Rust is that it makes the above impossible in safe code. If you attempt to create such a situation, the compiler will fail to compile your code. Every single integer operation could stop being atomic and not a single line of safe rust would break.
The only language I've seen where this is possible is Julia (which uses LLVM too).
Compilation strategy and code hotswapping are independent of static vs dynamic types. Java is statically typed, however its implementations support interpreted execution and code hotswapping.
Types that have fixed size at compile time are more useful in Rust and can be optimized much better.
Huon Wilson worked on an implementation a year ago:
https://github.com/huonw/float
Related- there was a discussion a few years back in Reddit about future-proofing math/numbers in Rust:
https://www.reddit.com/r/rust/comments/1uy7rt/an_appeal_for_...
However, when it comes to speed, working with primitive types has gotta be faster if supported natively, so anything else anytime soon will play second fiddle.
[1] https://www.forth.com/starting-forth/5-fixed-point-arithmeti...
It also is a lot less of a concern on modern systems, where floating point operations aren't much slower than integer ones.
At least on x86-64, the 64x64->128 bit multiplication is a single instruction, like the 32x32->64 bit and the 16x16->32 bit multiplications. Doing the calculation with only three limbs is clearly faster than doing it with five limbs; to start with, using 3 limbs you need 9 multiplies, while with 5 limbs you have to do 25 multiplies. The carry propagation and reduction steps also take time proportional to the number of limbs.
ping6 42540577535212633203815888880477462122
Like you can do this: ping 2158835347
PING 2158835347 (128.173.54.147) 56(84) bytes of data.
64 bytes from 128.173.54.147: icmp_seq=1 ttl=52 time=20.4 ms
64 bytes from 128.173.54.147: icmp_seq=2 ttl=52 time=20.5 ms1. I guess having bit mask operations for IPv6 addresses could be useful.
This is the big annoyance in Golang with it's net.IP type - it's `type IP []byte`, which means you can write ip1 == ip2 and you can't pass by value easily, nor use it as a map key.
I've ended up inventing my own type for that a lot of the time as a struct wit static fields, since those you can copy around and do 1:1 comparisons.
Though to be honest I'd be super happy if there was a drop-in varint type, or you could trivially have the compiler calculate instructions for a arbitrary fixed size ints.
Most UUID libraries I've seen and written use [16]byte as the concrete UUID type.
Sorry for the rhetorical device.
type IP []byte
That is not the same as [16]byte. // Note that in this documentation, referring to an
// IP address as an IPv4 address or an IPv6 address
// is a semantic property of the address, not just the
// length of the byte slice: a 16-byte slice can still
// be an IPv4 address.
What I believe that comment is saying is that something that is [16]byte can still be an IPv4 address, but that doesn't mean that all IPv4 addresses are stored as [16]byte. At least that would be my interpretation based on the comment above it: // An IP is a single IP address, a slice of bytes.
// Functions in this package accept either 4-byte (IPv4)
// or 16-byte (IPv6) slices as input.(Maybe you want to foreach over an array every time you want to apply one. I'd rather not.)
In my world integer is a mathematical construct with no particular representation, making things like bitmasks and or shifts nonsensical.
If you really want to work with fixed length bitstrings why not just have a type for that? Operating on a string of 128-bits should be valid on all such bitstrings no matter wether those represent a number or a string of code points.
And equally operations on integers should not care about particular bitstrings representations of the number in question.
Your world doesn't map to the reality of silicon and registers, whereas Rust does. As it happens, you can be fixed much more easily than the whole of modern computing.
> If you really want to work with fixed length bitstrings why not just have a type for that?
I don't. I want to work with integers. An IPv6 address is not the hex format that you read--it is a 128-bit integer. You can go read RFC 2460 if you don't believe me, but it's true. It is an integer that I can add and subtract from; I don't add 1 to an octet of an IP address and then do a bunch of carries if I want the next IP address in my network, I add 1 to the IP address. I don't perform some magic operation to determine what a subnet looks like, I bitand the integer. They are inescapably based on the representation used both by my computer and by my network hardware. (As is the performance of both my network hardware and yours. There's a reason that your router doesn't use BCD or whatever.)
There are programming languages that do not represent the underlying system. They are, for the most part, bad at dealing with the kinds of problems Rust is tailored to effectively represent. You can use those. It's pretty presumptuous to suggest that languages designed for lower-level problems accommodate your peculiarity.
About the only mathematical operations I can think of which are ever done to them are bitwise-anding, bitwise-inclusive-oring, and testing for zero (and, as mentioned, equality).
If you have internalized an IP address as a dotted set of octets expressed in ASCII digits, that's a problem of comprehension. When you don't, it's pretty natural to just use this stuff like any other integer.
It's like saying a pointer to RAM isn't an integer. Of course it is, and you add to them every time you de-reference an array in C. That it has additional semantic meaning doesn't mean it stops being a integer. The pedantry you're peddling doesn't fly.
That is a cool link though, I'm gonna give it a deeper read.
if ip >= IP(10.34.12.3) && ip <= IP(10.34.12.9)
//is one of our database servers
> They are never added, subtracted, if abs(atk1.ip - atk2.ip) < 16
//same source likely, modify threshold
Netmask are useful, but sometimes doing regular math is a better fit for your problem. Adding an integer to an IP, subtracting IPs or comparing IPs all yield meaningful results.Added: I'm talking about implementing specific, optimized hash functions on 128-bit values (e.g. GUIDs, IPv6 addresses), and not generic hash functions that can take any length input (although many of them speed up linearly in the size of int that gets used internally).
Hashers provide conveniences for feeding in a u8, u16, u32, u64, etc., but most hashers just implement this as casting the value to an array of bytes and using the generic implementation. This is because most algorithms are defined in terms of bytes. Evidently there isn't any interesting optimization to do by statically knowing you have a u32 vs a u64 for these algorithms (appears to be the case for SipHash, XXHash, and Fnv).
And indeed, u128 continues this tradition in the PR: https://github.com/rust-lang/rust/pull/37900/files#diff-2327...
I know fortran and Cobol code that banks use and is used for science (tm) often defines numbers this large for certain operations.
One such i can think of is a map reduce of large datasets. Let's say you want to find as a result a huge sum. 2^128 is a bit bigger then 2^64 and that difference may be big enough to provide the computation needed for getting to Mars rather then getting to the moon on 2^8 machines.
I haven't used it, but Julia lets you declare any(?) fixed size bitset. Which would probably come in handy for Go.
It would be much easier for them to just build on top of i128.
Also, you can see that one of the implementations is copied from Cairo, so it seems graphics libraries have some use for 128-bit integers too.
But I don't think rust has support for the Emotion Engine in any case.
In other words, they weren't removed because we fundamentally didn't want them. They were removed because of maintainability, usability, and usefulness concerns.
The RFC with justifications for why this was added (and all the related discussion) is already linked elsewhere in this thread. The same could still happen with f128 in the future.
> The Duration type could be simplified with this: instead of using a u64 for seconds with a seperate u32 for nanoseconds, it could just be a single u128/i128 count of nanoseconds.
Of course, just a single u64 at ns precision would get you 500+ years of range. But ok.
Floating point errors are hard to reason about and often makes equality a very fuzzy concept. If your not starved for bandwidth or memory, large fixed precision numbers are incredibly useful.
You can work around the problem - time since program start, time since capture start, etc. - of course, this can lead to fun edge cases, where e.g. Windows 95/98 crashed due to a 32-bit milliseconds timer rollover after around 50 days. For comparison, pow(2,53) nanoseconds is only a little over 100 days. Of course, a floating point value won't roll over in quite the same way, but...
The problem is surmountable if you throw enough edge case handling at the problem. Force the devs to be vigilant about choosing the necessary precision, the proper time to measure relative to, for each possible application of anything time related, etc...
... or you could just throw more bits at the problem, and suddenly my accounting software can accurately calculate compound interest on both a 400 year old debt, and a 49 nanosecond debt clearing HFT trades, if that's your kind of thing - without as many edge cases to worry about.
EDIT: Use pow(x,y) formatting since HN collapses x double-star y to simply xy ...
"On September 20, 2013, NASA abandoned further attempts to contact the craft.[76] According to A'Hearn,[77] the most probable reason of software malfunction was a Y2K-like problem (at August 11, 2013, 00:38:49, it was pow(2,32) of one-tenth seconds from January 1, 2000)"
I do agree that a general timestamp needs to be domain-agnostic, and I'm certainly not saying Rust should use it, not least because Rust aims to preserve the semantics of underlying APIs.
I can see why (b) might smell a bit bad, but think of it this way: we share a backend with clang, and so if they think support is mature enough to ship, then that's a very positive sign.
fn is_token(c: u8) -> bool
http://kamalmarhubi.com/blog/2015/09/15/eliminating-branches...We probably won't see larger then 64bit bus/address access for some time as we don't need it currently.
For integer/fixed point math, the functionality is here since SSE2: https://en.wikipedia.org/wiki/SSE2
Unless you’re planning to do reverse engineering, or planning to work on a compiler, learning assembler is mostly pointless. In most cases, real-world algorithms contain both vector code in the inner loops, and scalar code everywhere else. When coding assembler you need to use it for both, and assembler ain’t exactly user-friendly. Using C or C++ language with SSE intrinsics is the way to go. All modern compilers support them.
Intel’s documentation is the best so far, but there’s no offline searchable version. I’ve created one: https://github.com/Const-me/IntelIntrinsics/releases
Memory layout is a king. You need to keep the input and output data SSE-friendly: aligned, dense, sequential access patterns are preferred. This could mean you need to [re]design some parts of your software specifically for SSE.
There’re multiple generations of hardware. When writing manually-vectorized code, the compiler won’t tell you what CPUs it’ll run on. SSE2 is the most compatible. Here’s some statistics about Windows users: http://store.steampowered.com/hwsurvey/ click on “Other settings”
glhf
lg 10 is 3.322
3 is ten percent accurate. 3.333 is one percent accurate.