Announcing Rust 1.24.1
blog.rust-lang.org
blog.rust-lang.org
"There are only Copy types on the rust stack frame being jumped over."
Why are developers (since this is an issue that shows up with cargo) running Windows 7 without security patches installed? Especially since the issue only shows up on Windows 7 installs that haven't received security patches since June 2016.
> libgit2 created a fix, using the WinHTTP API to request TLS 1.2. On master, we’ve updated to fix this, but for 1.24.1 stable, we’re issuing a warning, suggesting that they upgrade their Windows version.
That's really a neat and responsible way of handling this.
You now have to visit and manually download the updates from Microsoft's Update Catalog website. Oh, and that website only works on Internet Explorer.
Maybe you ran into this issue? https://www.myce.com/news/windows-7-dont-receive-security-up...
Enterprise. It’s only recently I’ve stopped seeing XP boxes around, but Windows 7 is everywhere.
The issue shows up on patched machines. The patch does not enable TLS 1.2 by default. You have to add reg keys for it.
https://github.com/rust-lang/cargo/issues/5065#issuecomment-...
https://doc.rust-lang.org/std/str/struct.SplitWhitespace.htm...
Why would you create a special data type to represent a string split by Whitespace? Lunacy
Basically the same as calling .split(‘ ‘) on a string in JavaScript
You can find a list of Rusts primitive types here: https://doc.rust-lang.org/std/#primitives
It's a named struct. That's pretty datatype-y.
I don't agree with the GP that this is an example of overengineering. If anything it's an example of current Rust being slightly underengineered. Presumably a lot of these temporary types can be killed off once `impl Trait` goes mainstream?
When "impl Trait" lands, it will be possible to hide details like this.
Because it's an iterator. A bespoke iterator type is also created behind the scenes in many other languages. How else would you do it?
There is a `splitWhitespace` iterator[1] in the Nim programming language as well and doesn't require a separate type. In Nim, the iterator is inlined.
1 - https://nim-lang.org/docs/strutils.html#splitWhitespace.i,st...
Note that it’s exactly the same in other major statically-typed languages like Java or C++: there are no existential types without indirection (which in the case of Java and similar managed languages is implicit and mandatory).
Rust is in the progress of adding existential trait bounds in the form of the `impl trait` feature. It’s already available in nightly.
It's not just statically typed languages. Python also returns loads of special types from iterator functions, they are just usually not documented as being types:
>>> type(itertools.chain([1, 2], [3, 4]))
<class 'itertools.chain'>Big chunks of Rust are build around iterators, and `split_whitespace` returns an iterator of type `SplitWhitespace`. Because this is a concrete type, the Rust compiler and LLVM will then work to together to completely inline it, and they will generate code that looks like a hand-rolled loop.
There are some downsides to this system—it usually takes me about 10 minutes to write custom iterators for a new data structure—but iterators are very nice to program with and they go fast.
There's a new 'impl Trait' feature scheduled for later this year which will eliminate the need to export a custom struct like this. And it will eliminate the 10 minutes I spend writing iterators. Of course, it adds a new language feature. Nothing's free.
> Rust gives me a headache. I want to like it, but it just seems so... overengineered
I admit, Rust does sometimes have a "heavy industry" feeling to it. But this has some nice benefits, too:
1. I can write cross-platform CLI tools that Just Work on Linux, MacOS and Windows, because Rust has good abstractions for paths, files, threads, etc.
2. I can write multi-threaded code that does things like, "Read an arbitrary stream of bytes in a background thread, compress it, break it into 5MB chunks, and upload each of those chunks to S3 in parallel, using no more than N worker threads and applying backpressure, and do all this in the background while I work on something else." And thanks to Rust's threading rules, all this will work on the first try, with no nightmarish threading bugs.
3. In general, if my Rust code actually compiles, there's about an 85% chance that it will work flawlessly on the first try.
Personally, these are benefits that I'm willing to pay for. And Rust does require some familiarity both with how processors work, and with functional programming. And of course, everybody has different tradeoffs. But for certain kinds of work, Rust really hits the sweet spot.
Do you use a crate like Rayon for this? I'm just starting with Rust, and my current understanding is that without using a library, one can only spawn threads and distribute load by hand (as opposed to automatic scheduling a-la OpenMP). Is this correct?
`rayon` is awesome, and it uses `crossbeam_dequeue` internally. But `rayon` is really best-suited to computational parallelism on many small pieces of data, and less suited to I/O parallelism. So it may be worth using `crossbeam` directly in that case.
All the threads in our S3 uploader run on a single machine, because multi-part uploads are mostly limited by how much memory we want to use for buffering S3 object parts, not by CPU.
(Of course, I'll probably overhaul a lot of this code later this year once `#[async]` and `futures` stabilize, so I can also stop paying for the memory used by thread stacks.)
But the cool part is that we can already do streaming input, compression, chunking and parallel uploads without ever creating a temp file. (And we can upload multiple data streams using the same worker pool.) This took about three days to build.
In particular, for this task, this requires to:
1. Return references to subranges of the original string, rather than copying them, so that no copy happens if you only need to examine the component instead of storing it
2. Not use reference counting to do so, but rather statically checked references with lifetimes, to avoid unnecessary instructions to update the reference count and lack of a static finalization point
3. Provide a way to get components one by one, so that if you only need e.g. the first two, time is not wasted to split the whole string
4. Provide that through a generic Iterator trait, so that it may be passed to generic methods (like one that collects the result into a vector)
5. Dispatch that generic trait statically rather than using an indirect call as that would destroy performance
6. Make the state manipulated by such an interface into a first-class object, and allow to put them in a data structure (like an array) while still doing static dispatch, so that you can, for instance, split multiple strings into components and interleave them without ever making an indirect call.
The combination of these essential requirements results in the creation of the SplitWhitespace<'a> data type, which represents the state of a parser splitting a string into 'a-lifetime references to whitespace-separated its components one by one, implementing the Iterator trait, and usable in a data structure.
Second, you're making it sound as if the position Rust (and those who program in it) is somehow undesirable. This state of things is the consequence of an explicit design goal, which was to accomplish automatic memory management without runtime cost.
And yes it's a consequence of a design goal, I said as much. The outcome however isn't beautiful enough to feel smug about the rest of programming languages.
It's automatic memory management without runtime accounting or a garbage collector, which means Rust doesn't need a runtime at all.
Call it whatever you want, but there are languages where you don't have to "pass ownership" manually for every frigging thing.
- You split a string, returning an iterator or slice where every element is a slice of the original string.
- You change the original string.
- Now the splitting may be invalid.
Such bugs are not possible in Rust, since; (1) the SplitWhitespace struct borrows the str immutably; (2) the borrows checker does not permit simultaneously borrowing data as mutable and immutable.
Another problem that is prevented by Rust is 'memory leaks' where someone splits a large string and uses only a smaller substring. If the substring is a slice of the original string, a GC cannot deallocate the larger string. Rust prevents this, because a string slice reference (&str) cannot outlive the underlying String. [1]
tl;dr: better of the half points are indirectly and indirectly caused by Rust's ownership model that prevents a lot of ownership bugs. That does not mean that you do not have to think about ownership problems in other languages.
[1] Such memory leaks were one of the reasons why Oracle Java switched to a much slower, copying implementation of substring:
http://java-performance.info/changes-to-string-java-1-7-0_06...
A large number of engineers do desire such a tool, and Rust provides that in a way that provides as little cost as possible over one of those more dynamic languages.
Your phrase about 'painting itself into a corner' implies you disagree with point (a), and that was the case I was trying to make.
For example:
Point 1 is known because the Item type is &str.
Point 2 is known because, well, that's the norm, but beyond that, SplitWhitespace is parameterized over a lifetime
Point 3 is known because it's an iterator; next() is its primary interface, which returns things one by one.
Point 4 is the "impl<'a> Iterator" bit
Point 5 is known because the return type of split_whitespace is this iterator, not Box<Iterator>
Point 6 is related.
If we repeated all of this stuff in every single bit of docs, it might make it more useful for some audiences, but also kinda destroy the docs for intermediate/advanced Rust users.
I'd love to make docs generally more accessible, but I'm not aware of any great solutions to this particular problem.
However, thinking about this more, I think I understand what you mean and I, sort of, agree.
What I believe is going on is that Rust exposes a lot of its engineering. This can be either good or bad depending on where you're coming from. If you're working at a low level then this is usually a good thing. If you're working at a high level then this can be quite annoying.
What I think you'll find is that this is largely a point in time thing. Rust is still relatively young and there is substantial work in flight under the banner of ergonomics. Even though this is more engineering, I'm pretty confident that the net effect will be to make the language feel less engineered. It probably won't completely get there this year but I don't think it'll be long.
Filter<Split<'a, IsWhitespace>, IsNotEmpty>
The split[1] method is generic, and for example, one can indeed use a regex for it. (I wouldn't use it for CSV though, since it would almost certainly be wrong.)[1] - https://doc.rust-lang.org/std/primitive.str.html#method.spli...
Yes, it's convenient, but using data types can help with many things, from implicit documentation, to optimizations to enforcing a code contract/interface.
And there's actually very little chance that you, as the developer, are ever going to explicitly specify that type - there's pretty good type inference. As a rust user, I don't care that it's another type because I don't need to care about it in order to write code.
So lots of benefits, very low cost/impact.