I would be interested if they're wrong, hehe
I would be interested if they're wrong, hehe
Embedded can be tricky at times; we've had a working group pushing on making stuff great all year. There's still a lot of work to be done. Sometimes you need to opt into nightly-only features. We'll get there!
(One major example is platform support: since we're built off of LLVM, we may not have a backend for more obscure architectures. ARM stuff works really well, generally. AVR is down to some codegen bugs...)
There's one significant thing that's in C that's not in Rust: alloca. If you need that, well, we've talked about it, but haven't accepted a design, and some people never want to gain it directly.
EDIT: another comment pointed out goto; I have no idea what the performance implications are but I’ll have to add that to my list of “stuff C has that Rust doesn’t.”
Sometimes, you need nightly-only features. This means that, in some sense, today's Rust is slower, but tomorrow's Rust may not be.
Sometimes, "equivalence” is the issue. You can turn off bounds checking for array access, for example, but most Rust code doesn't, for hopefully good reason. Sometimes those checks are elided, and so it's exactly the same. Sometimes they're not. Sometimes they're not and that inhibits other optimizations. Is unsafe Rust "equivalent" to the C? Or does that not count? (Unsafe Rust is not always faster than Safe Rust...)
At the theory level, I can't think of anything off the top of my head that would inherently make Rust slower than C, generally speaking. There's also some degree of argument that in theory, Rust should be faster than C, thanks to how much more we know about aliasing, etc. That's a whole other can of worms...
This is something I've personally thought about a lot and I have no idea how I'd actually want it implemented either on the backend or in the syntax, but I think it'd be useful to have some way to say "I want these bound checks to be elided, if they can't be then make that an error". Similar to how you'd decorate a pure function in other languages to say that it can't cause side effects or depend on side effects for what it's doing.
I think that's probably a high level ask that's probably significantly more difficult that it seems initially too.
Edit: Looks like there's been an RFC in the past to add it to rust, but it got punted because the const generics stuff went in first and they wanted to see how that would shake out before trying to fully tackle this. https://github.com/ticki/rfcs/blob/pi-types-ext-2/text/0000-...
On that front, one interesting thought would be to reimplement the CPython bytecode interpreter in Rust and see if you can match the performance of C using computed goto.
In principle, there's no fundamental reason Rust couldn't optimize "loop around match" exactly the same way, without needing the computed goto. (For that matter, so could C.) Doing that would help cases like this.
For safe Rust, sure, there are limitations: you can't turn off bounds checks, for instance. But all Rust is partially unsafe Rust, because the standard library uses unsafe code ⃰. So it's a fuzzy distinction.
⃰ ⃰You wouldn't want it any other way. Implementing low-level functionality as unsafe code allows us to write code, not compiler-code-that-generates-code.
In pseudocode
Goto = Start;
While(Goto!=End) {
Match(Goto) {
...
}
}
It’s a pretty basic obfuscation technique.
Compilers will happily unroll this leaving you with an irreducible CFG.* In C++, using goto: https://godbolt.org/z/YqCMbU
* In Rust, with relooped CFG: https://godbolt.org/z/IIXtKo
The compiler was obviously not able to unroll the relooped code into the original CFG.
A specific example: in principle, the compiler could see the exact target of every continue in the Rust code, so the continue on line 34 could go directly to line 27 (or, better, as in C++ the basic blocks could just be laid out adjacently), but the compiler does not actually do this and there are a bunch of unnecessary tests on the path between those two lines.
This doesn't work in general, and is ... impossible to maintain, but it's not an unreasonable approach for generated code.
I know C++ quite well, and I'm confident that the C++ version is reasonably well optimized, though it probably has a little room for improvement.
In rust, I think I'm doing things in a reasonable fashion but so far the performance is only half of the C++ version. So, not bad, but I was hoping it would be closer. I'd like to know if anyone has any suggestions for resources related to rust optimization.
What’s the situation with the IO? Are you buffering? Are you holding the lock for a long time or rapidly locking and unlocking?
If you can share the code I can take a look.
Code is here: https://gist.github.com/usefulcat/56f334bc58c97edb073b457b68...
There is actually one other file but it only contains a couple of struct definitions for the libpcap file and packet headers. Thanks for having a look!
I didn't think there was room for that much improvement; I'm really impressed.
Edit: HOLY CRAP! I have a program written in Rust, ppbert[1], and I just tried wrapping my StdoutLock object in a BufWriter, and I improved the performance of my pretty-printing by a factor of 2x! I knew to use BufReader for files, I didn't know it was helpful for stdin and stdout! Thank you _so much_ for sharing your experience, I've certainly benefited!
Benchmark #1: ppbert -2 *.bert2
Time (mean ± σ): 3.816 s ± 0.115 s [User: 2.494 s, System: 1.321 s]
Range (min … max): 3.688 s … 4.028 s
Benchmark #2: ppbert-dev -2 *.bert2
Time (mean ± σ): 1.728 s ± 0.045 s [User: 1.493 s, System: 0.234 s]
Range (min … max): 1.678 s … 1.843 s
Summary
'ppbert-dev -2 *.bert2' ran 2.21x faster than 'ppbert -2 *.bert2'
[1] https://github.com/gnuvince/ppbertIt seems like even C projects have a hard time being portable. People even avoid cmake because of the large dependency and the fact that it is cross platform until it isn’t.
Then, you have to make sure that it doesn't break; this means running on that target in CI, somehow. Given that you're already talking about devices that may not even have an OS... emulation can work sometimes?
Then, there's ecosystem stuff. You probably want some sort of HAL and support for not just one platform, but all of the platforms you're deploying to. So that's more work...
Finally, some platforms are proprietary and basically give you their own fork of gcc and so C is pretty much your only option anyway.
First of all, most embedded development makes use of bare metal, where the libraries take the role of an OS, or they use a specialized OS from the hardware vendor.
Just using pure ANSI C isn't possible, because the standard does not expose the hardware features from the underlying platform, so the alternatives are to use Assembly, or language extensions.
Naturally language extensions are more convenient to use, so that is what most developers end up doing.
Also there are many types of embedded platforms, you can be targeting anything between a tiny PIC with 8KB FLASH RAM to a powerful multi-core ARMv8-A with 8GB.
So the toolchain must allow for customization of what actually gets linked into the final binary, and the runtime must be as thin as possible.
Then there this the drivers story, each vendor gives you their own SDK, which most of the time is the only way to access their devices.
It is typical for open source projects to reverse engineer some of those SDKs to get the necessary information for linker maps, compiler flags and driver information.
Regarding Rust, there is an ongoing effort to create a embedded library for hardware drivers, as means to write portable code.
I think you're spot on with regards to tooling being a major concern, and to me, this is one of the promising aspects of Rust, with cargo at the high level and the LLVM tools lower down. It seems like a system that's primed for open source libraries targeting specific families of chips, cleaning up the mess we've got now with each vendor having their own framework, IDE, configuration tool, etc.
Speed and size concerns though, seem to depend a lot on the specific project. In the last couple years, I've worked on systems that had neither/either/both of those as the primary challenge. Of course in embedded, where you try to save money on parts early in a project can determine the technical (and so, schedule) challenges later on.
And, I'd add safety to the list, having spent loads of time finding bugs in legacy C caused by stuff like walking off the ends of arrays, jumping to uninitialised function pointers, fun with enums and unions, etc.
James Munn gave a super interesting talk at this year's RustConf about some of the abstractions they're building for embedded systems: https://www.youtube.com/watch?v=t99L3JHhLc0
I think Rust is in a great position to become incredibly useful in the embedded space.
Also, once you go beyond blinking a couple of LEDs and especially if you care about energy consumption, you quickly get mired in complexity, dealing with interrupts and resulting concurrency. If you're a seasoned embedded developer, you will have an explicit, mostly declarative state machine at this point, possibly nested, and if you have more than 15 years of experience, you will be seriously thinking about better representations and generating all that C code from a Lisp.
I'm have a lot of hope for Rust in the embedded world. I especially hope that it will get embraced by vendors. To take a practical example again, I'm thinking about Nordic Semiconductor, which has really good SDKs and is innovative.
The real questions are whether Rust would improve the performance of things like a Bluetooth Low Energy communications stack given that communication stacks are probably the most complex things that run on a Cortex M0 or M4 class processors.