My Favorite Rust Function Signature
brandons.me
brandons.me
fn tokenize(code: &str) -> impl Iterator<Item=&str> {
but including them to be more explicit does make sense for the post.- Each elided lifetime in the parameters becomes a distinct lifetime parameter.
- If there is exactly one lifetime used in the parameters (elided or not), that lifetime is assigned to all elided output lifetimes.
So, Since the signature has one `&str` as a parameter, the lifetime is assigned to the output lifetime also per these rules.
[1]: https://doc.rust-lang.org/reference/lifetime-elision.html
(Edited to improve formatting)
There are a few a few example cases in the doc linked about that walk through the conditions it's necessary to add explicit lifetime annotation.
fn foo(bar: &x) -> &y
is inferred to be: fn foo<'a>(bar: &'a x) -> &'a yThis pretty much sums it up for me. You need to know so many rules... you need to know what those symbols mean... when people refer to Perl's syntax that they frown upon, they are referring to this: heavy use of symbols whose meanings must be memorized. I cannot just look at the code and understand it as I can with say, Ada. Pretty frustrating. I do not have to learn Ada to know what the code does. Funny that. :D At times Rust code looks very similar to Perl and Haskell. Especially Perl.
Forget the rules and don't write a lifetime where you needed one? You get this:
error[E0106]: missing lifetime specifier
--> src/lib.rs:1:29
|
1 | fn foo(x: &i32, y: &i32) -> &i32 {
| ---- ---- ^ expected named lifetime parameter
|
= help: this function's return type contains a borrowed value, but the signature does not say whether it is borrowed from `x` or `y`
help: consider introducing a named lifetime parameter
|
1 | fn foo<'a>(x: &'a i32, y: &'a i32) -> &'a i32 {
| ^^^^ ^^^^^^^ ^^^^^^^ ^^^
It even shows the syntax and where things changed.Now, it's not an AI; this may not be the signature you wanted, exactly, but the point is, you don't need to memorize the elision rules. You just don't write any lifetimes until the compiler makes you deal with it, and then you deal with your problem, and move on with life.
"There comes a time when words are more appropriate than symbols."
I find Ada very hard to read. ... come to think of it, doesn’t it also have unbalanced apostrophes? It’s been a while.
> unbalanced apostrophes?
What are you referring to exactly? Can you give me an example? Do you mean something like Integer'Last, which is the largest value of type Integer?
> The concept of attributes is pretty unique to Ada. Attributes allow you to get - and sometimes set - information about objects or other language entities such as types. A good example is the Size attribute. It describes the size of an object or a type in bits:
type Byte is range -128 .. 127; -- The range fits into 8 bits but the compiler is still free to choose.
for Byte'Size use 8; -- Forces the compiler to use 8 bits.
I find it particularly neat.Of course one of the best things I love about Ada is pre- and post-conditions or contract-based programming (that are actually checked, unless in languages like Kotlin that just bolted on some contracts that are just there for the reader, i.e. pretty useless, to be honest). Ranged types are awesome, too, as you tell the compiler the range you want and the compiler chooses an appropriate sized integer to suit the range, which is the important bit. For example, instead of asking for "16-bit signed integer", you say: "I want a type which can store values between -1 and 1000", and maybe even a hint for the compiler where you can ask the compiler to pick you a faster type or a smaller type, so you say:
type Index is range -1 .. 1000;
Types like these are much more readable, and you can use attributes on them such as 'Min, 'Max, 'Range, and so forth. It is such a good feature, and you can remove the range checks if you want, and you can even have them checked at compile-time. These kind of features are what I would like to see from a C replacement.See more at https://learn.adacore.com/courses/intro-to-ada/chapters/stro....
---
https://www.adaic.org/learn/materials/intro/part1/
https://www.adaic.org/resources/add_content/docs/craft/html/...
https://www.adaic.org/advantages/features-benefits/
https://www.radford.edu/nokie/classes/320/designGoals.html
https://hackaday.com/2019/09/10/why-ada-is-the-language-you-...
Those links might give you some idea about the language and its goals, if you really are interested. Of course there are much more, up-to-date sources, but I think these got most of it right.
index = num(-1->1000);
or something. :DI dunno how attributes would work, but man, they really are handy. It is very, very difficult to shoot yourself in the foot.
I see what you did there (but maybe you don't see it yourself :-)
PS: at least MY parentheses are balanced.
That seems unlikely, or anyone - even if they hadn't learned to program ever - would be able to read Ada. More likely that a) you've already learned enough to cover Ada by learning other programming languages or b) you assume you know what the Ada code does but are incorrect.
Let's charitably go with a) then what you're implicitly supporting is that "all languages must look and behave the same, and there's no possible benefit in a language being different enough that it requires me to learn something new".
And therefore that the lowest-common-denominator intersection of all popular languages combined is also the pinnacle of language design.
Something which seems unlikely.
Perhaps you could compare some Ada code (something from AdaCore (GNAT, GNATColl)) with the Rust equivalent.
If you don't know Ada then you don't /know/ that you understand it, you're assuming you do - and if you do that's because it looks like languages you already know. And if that's what you're saying is good about it, then what you're saying is "things I know are good". You're seeing "Ada looks familiar" and saying "Ada has good design" as if those are the same thing. You only need to look at some beginner programming discussions to see that block-structured, scoped, procedural, looped, branched, etc. features are not things people are born understanding - that you know them already and Ada has them is convenient for you.
> "What the code does is not hidden behind symbols, for example."
I looked at the Wikipedia page, it has equals, colon-equals, slash-equals, greater-than, less-than, apostrophe, parentheses, ampersand, double-quotes, equals-greater-than which seems to be an arrow not a comparison, double-dots, and of course "when" "is" "protected" etc. they're as much symbols as they are words.
> "Ada and Rust, Ada is way much easier to understand"
I'll quote Erik Naggum from an old comp.lang.lisp(?) comment: "When I work at this system up to 12 hours a day, I'm profoundly uninterested in what user interface a novice user would prefer.".
Here are the rules:
1. Every reference as a parameter gets its own lifetime parameter. There is only one in this case.
2. If there is exactly one, then every possible lifetime in the return type is assigned to it.
3. If there are multiple parameters, but it's a method, not a free function, then the lifetime assigned to self is assigned to all outputs.
This is written in a bit of an abstract way, but, #2 there describes this function: we have one input, so we assign the same lifetime to all outputs, and there's only one. This means that it's not 'static, and in fact, it would fail to compile if it assumed that.
The original RFC, if you're curious https://rust-lang.github.io/rfcs/0141-lifetime-elision.html
(And, I actually think there's been some tweaks over time; if you have no inputs, the outputs can turn into 'static... the core of it is these three things though. Oh yeah, that is in the reference, which is the up to date spec, https://doc.rust-lang.org/reference/lifetime-elision.html; I linked the RFC for historical reference because it's interesting, and has some neat stats.)
Because you can never know how a given website will handle the input, I always leave a space after URL and before punctuation in a text box, like so: http://example.com ;
This leaves the only remaining possibility that references in the return type must have the same lifetime as references in function's arguments. Without a GC, or leaking memory, you can't make a Rust reference/lifetime out of thin air, so there's always such relationship.
When a function takes multiple arguments by reference, the compiler will ask to disambiguate which one is used for the return type. Borrow checker intentionally checks borrowing against function prototypes, not function bodies.
However, Rust lifetimes are stretchy-squeezy, and a free lifetime is basically the same as 'static when it comes to covariant lifetimes, since both are "I can be whatever you want me to be".
However if the lifetime is invariant (e.g. in https://play.rust-lang.org/?version=stable&mode=debug&editio...), it is actually not 'static and will map to whatever lifetime is requested by the context.
Proof: https://play.rust-lang.org/?version=stable&mode=debug&editio...
I'm pretty sure boxed trait objects (maybe all boxed objects?) are assumed to be 'static. That is, Box<dyn MyTrait> is shorthand for Box<dyn MyTrait<'static>). If you want non-static lifetimes you have to do something like Box<dyn MyTrait<'_> + '_>
(I only just learnt this so take it with a pinch of salt.)
TL;DR: Trait<'a> is actually understood as Trait<'a> + 'a.
https://play.rust-lang.org/?version=stable&mode=debug&editio...
[1]: https://github.com/rust-lang/rfcs/blob/master/text/0599-defa...
Also I'm going to point people here next time someone tries to claim that Rust is simple.
help: to declare that the trait object captures data from argument `self`, you can add an explicit `'_` lifetime bound
|
14 | fn things(&self) -> Box<dyn Things<'_> + '_> {
| ^^^^
Edit: having said that, the docs do explain these rules[1], reachable from the trait object documentation[2], while it seems that The Book doesn't.I'm also noticing that my confusion was likely due to `&'r Ref<'q, Trait>` being interpreted as `&'r Ref<'q, Trait+'q>`, which is behavior that was introduced in a follow up RFC[3].
[1]: https://doc.rust-lang.org/reference/lifetime-elision.html#de...
[2]: https://doc.rust-lang.org/reference/types/trait-object.html?...
[3]: https://github.com/rust-lang/rfcs/blob/master/text/1156-adju...
// this function foo
fn foo(s: &str) -> &str {
&s[0..1]
}
// is equivalent to this function
fn foo_desugared<'a>(s: &'a str) -> &'a str {
&s[0..1]
}
// this works too because the lifetime of "" is 'static because
// it is a str that lives in the .data segment, which will live for "any 'a"
fn bar(s: &str) -> &str {
&""
}
fn main() {
// as shown here
bar(&String::new());
}
// but this won't work, because `x` gets `drop`ed at the end of `baz`
fn baz(x: &str) -> &str {
let x = String::new();
&x
}
https://play.rust-lang.org/?version=stable&mode=debug&editio...Keep in mind that lifetimes have the opposite behavior of generic types. A T: Trait type parameter means "something that implements at least Trait, but might implement other things that you won't have access to", while an &'a returned borrow means "something that lives at least up to 'a, but might live less".
I don't know what you're trying to say here, but taken at face value it is wrong, and your own `fn bar` example demonstrates it being wrong. "something that lives at least up to 'a but might live less" is a contradiction.
Perhaps you meant "... but might live longer" ? If so, it would be correct, and also would not be "the opposite behavior of generic types."
fn foo(x: &str) -> &str {
x
}
calling `let z = foo(y);` is valid. The lifetime of `y` conforms to the lifetime of `x` in the `foo` signature, while the lifetime of `z` can be as long as the one for `x` (up to x's lifetime), but it is shorter (unless you try to do `drop(y)`, which would be a borrow checker error).>while an &'a returned borrow means "something that lives at most up to 'a, but might live less".
This paragraph is a bit muddled:
* If you keep the two passes separate, then the second pass is over the token stream. So it's irrelevant that "tokenizing isn't the most expensive operation" because you're never going to do the actual tokenising a second time.
* If you combine the two passes then you're not "saving cycles" because you're still doing both of the loops - it's just that they're now happening in parrallel rather than in sequence. This confusion is repeated later when he says the combined one is "without the runtime overhead of a second loop".
* What you are saving is memory, because you don't have to keep the whole token stream in memory at once.
But really this is a nitpick on a small note. The overall article is great.
You're not avoiding the runtime overhead of a second loop but being able to operate on both streams in parallel would be a far more complicated concept in C/C++. This actually reminds me of that influential Rob Pike talk about writing a concurrent parser in Go.
Maybe you could take advantage of Rust's lifetime guarantees to multithread things more easily than C++. That would potentially be a much more interesting article, if it turns out to be a lot easier than in C++.
At no point does any of these schemes require the lexer to read the whole file into memory, and chop it in one go into a huge list of ready-to-be-consumed tokens.
My experience so far with Rust has been a huge reduction in mental burden/bookkeeping compared to when I write modern C++.
You make a good point about 'a; I could have made that clearer. In practice it works the same way as a generic type (the letter is arbitrary, and it's declared inside the < >), which I assumed people would be familiar with, but I can see how the connection between the two may not have been obvious.
This is wrong. Pretty sure you could do this in ATS for example.
Go, which is garbage-collected, also does the slice-by-reference trick (but fails criterion #1 because it doesn't really have iterators).
fn tokenize<S: AsRef<str>>(code: S) -> ...
code.as_ref()
Edit: whoops, was too quick, apparently this isn't a great idea here.[1]: And arguably a bad suggestion in general. If a fn needs a String, it should take a String. If it needs a &str, it should take a &str, and if the caller has a String they can call the fn with a borrow. fns that are generic on AsRef<str> are usually only found where the ergonomics make it worth it. Eg the standard library's filesystem functions often take AsRef<Path> parameters because string literals are AsRef<Path>, which means you can write
fs::read_to_string("/etc/passwd")
instead of a more cumbersome fs::read_to_string(Path::new(OsStr::new("/etc/passwd")))If they had tried to complete their code and compile it, they would've noticed there's no way the function could return `code.as_ref()`.
In C++ I might do it like this:
struct token {
string::iterator start;
string::iterator end;
};
and then the range in the original source can be recovered via: size_t offset = tok.start - src.begin();
What is the preferred way to do this in Rust? Is there something better than using indexes?Rust doesn't appear to have a way to say "give me the location of this &str in that &str" (even though it seems like it would be safe).
I don't think so.
> Rust doesn't appear to have a way to say "give me the location of this &str in that &str" (even though it seems like it would be safe).
'course there is: str::find (https://doc.rust-lang.org/std/primitive.str.html#method.find)
It looks a bit abstract because it takes a Pattern rather than an `&str`, but a char, a slice of chars, an &str, or a predicate function, are all patterns.
But there's no guarantee that two &str have any sort of overlap, that relationship is not conserved by the language, so it either would be unsafe or would be failible, it'd basically be pointer-twiddling.
What I want is the literal offset of a string slice in another, aka pointer arithmetic in C:
fn derp() -> usize {
let src = "xxxxxxx";
let tok = &src[3..4];
let offset = tok.as_ptr() - src.as_ptr();
offset // should be 3
}
this should be doable safely; I think it's just an annoying hole in the API.I see, I'd misunderstood the original, sorry.
> this should be doable safely
The language doesn't know that there's any relationship between those two strings, so the method would probably need to be fallible, and be more convoluted than this.
src = `let abc = "abc"` token = src[4..7] The 4 and 7 right there are the values we want. And if a str is a pointer and length, then those values should be retrievable with simple arithmetic, rather than a string search
The only issue is if the str isn't a substring. The signature would need to be either
`fn location(&self, &child) -> Option<(usize, usize)>` or `unsafe fn location_unchecked(&self, &child) -> (usize, usize) `
One common trick in C++ is to make source locations just a pointer. This gives you access to the token's contents (just dereference it), but you can also figure out which source buffer it came from, and where in that file, by comparing pointers.
Computing the position might have some overhead, but it's probably better than going over the whole file from the top when you want to turn that pointer into line/column numbers.
As for the content, in Rust this would be in your token structure (or rather in Rust your Token::StringLiteral enum variant), storing start/end pointers for each token seems weird to me.
[edit: Rustc uses byte indexes though, so my intuition might have missed something important for the design of larger compilers: https://github.com/rust-lang/rust/blob/7bdb5dee7bac15458b10b...]
As for the token themselves, they (or at least their position/Span) will be used beyond parsing. Otherwise you won't be able to point at the right place when you have a type error, for example.
Calculating line and column may be complicated if you're doing it for display. Perhaps the input has a flag emoji, with a variation selector, followed by a tab... So it makes sense to store the minimum you can and defer line/column calculation, especially if your parse is perf sensitive.
One place this can go wrong is when you have to show a lot of source locations, e.g. you're printing a backtrace for a stack overflow in an interpreter. Here you can easily get into an N^2 situation. So you want to defer but also memoize.
With all of that said, the approach taken in `rustc` is to have an arena of all `Span`s for everytokens and embed "pointers" (not actual pointers) to them. This is done because we embed `Span`s in everything in order to be able to give reasonable diagnostic messages.
That's a fantastic idea. It seems like it would be not just safe, but quite efficient, too.
fn offset_of(parent: &str, child: &str) -> Option<usize> {
let cp = child.as_ptr();
let pp = parent.as_ptr();
if pp <= cp && cp <= pp + parent.len() {
Some(cp as usize - pp as usize)
} else {
None
}
}
(Edit: as can be expected, my hastily-thrown-together example code has several bugs in it. I think there's an off-by-one error in the ending condition, and I'm really not sure if it handles multi-byte characters properly.)In any case, this scheme is pretty limited, since you probably want other info in your token span like line numbers.
struct Token<'a> {
s: &'a str,
start_index: usize,
}
I believe you can get the len() of a &str, so you could just store the start index and derive the end index.In generally that is my main complaint as I am trying to learn Rust, the syntax seems to be all over the place. Maybe that feeling will fade away as I get more use to the language.
This syntax was taken from OCaml.
Incidentally, nobody really loves it, but nobody could come up with anything clearly better either.
[0] https://www.youtube.com/watch?v=rAl-9HwD858 really helped me understand it.
Of course the Rust variant is not only cheaper (no ref counts), it's also guaranteed thread-safe (which again, C++ would have to use something like the atomic `shared_ptr` to achieve)
TLDR: Even when C++ can kind of do it, you have to give up concurrency safety to get close while adding memory pressure & still not being as fast while Rust gives you guaranteed concurrency safety for free regardless of accidental misuse of your API.
[0] Currently that would be constructor #8 on https://en.cppreference.com/w/cpp/memory/shared_ptr/shared_p... [1] Not only are you always paying for atomic ref counting even when you don't need it, aliased `shared_ptr` will be as expensive to take a ref for as normally constructing it - there's possible no way to `make_shared` with aliasing. This means the aliased shared_ptr will be even more expensive every time you need to access the ref count since the control block loses cache coherency. On the other hand if you structure it more carefully maybe you could vend `const std::shared_ptr<std::string>&` in your iterator so that in the single-threaded case your only overhead is constructing the aliased shared_ptr once & you leave it to the user to take a copy when transferring to another thread.