How to not rewrite it in Rust
adventures.michaelfbryan.com
adventures.michaelfbryan.com
At my employer, the number one reason (I'm aware of) to open source our internal software is to solicit collaboration from other companies. Take away collaboration, and it's going to be a harder sell to open source anything.
So maybe instead of "rewrite it in Rust", the answer is "build a new solution in Rust that meets the needs of a community of people who already use Rust".
I agree with A and B, but C just feels like gate-keeping. Open source license do not require collaboration with the original authors. You can freely fork any opensource codebase and modify it as you wish without asking someones permission.
I think the author means: if you're maintaining changes in your fork indefinitely and advertising this fork to others as superior to the original to use or as a canonical place to file issues, you should have a good reason. Communities are valuable; please don't divide them or throw them away lightly. Try working with the original maintainer first.
And what constitutes a good reason?
The epitome of open source is that you don't need to explain of your fork is really superior, the users will come.
RiiR certainly has it's place though. Ripgrep for me as a user is vastly superior to grep and a large reason for that I see is in the safe optimizations allowed by the borrow checker. They're certainly possible in C, but it wouldn't be maintainable.
Also I'd look at how stable the library is you're wrapping and the size of the code base. It may be lower effort to rewrite than to wrap.
Imagine a well-funded startup trying to make a name for themselves decided to fork the top 10 emerging open source projects, put their company name on them, and then spend millions in PR/marketing so that (A) it seems like they invented them, and (B) they fork the community. Are they free to? Is it right? What will it mean for the original authors and experts who created the software?
I happened to be reading this yesterday: dstat is dead. Red Hat decided to rewrite it using pcp, so the original author has given up. https://github.com/dagwieers/dstat/issues/170
Is Red Hat free to? (probably). Is it right? Now that's where I'd want to see a strong reason.
If the author of the original source had an issue with the above scenario, they would've chosen a more restrictive license that didn't make it possible.
This seems like the opposite of the usual ethical/legal distinction. It’s definitely legal to fork the project and the community, but not necessarily ethical.
I don’t think you can blame authors for not using a license that describes exactly what they are happy with. What constitutes an “ethical” fork is subjective, based on how necessary the fork is, and I don’t think a license could describe that.
Also, consider this analogy. The MIT license doesn’t contain a patent grant. So if some MIT-licensed software uses a patented technique, the software author and patent holder is within their legal rights to sue all forks of their for patent infringement. But this would be unethical, as the MIT license implies that the licensed software is given freely, not that you will sue all users of your software. Would you blame the sued parties for choosing to use MIT-licensed software, the way you are saying you would blame software authors for choosing a license that doesn’t forbid certain forks?
Really smart computer nerds emotionally wrapped up in change they aren’t interested in. They know what the future needs! To stay just like yesterday!
Gen X has become the old geezers who want you to stay off their lawn.
Oh what? It may not be super popular like Linux? We all have to work on the same code bases? Individual drive and curiosity verboten! Your one person project won’t satisfy the general user base? You should just focus on computing as we see it.
Who cares. Generate whatever syntax you want. They don’t have to organize around the semantics if they don’t want.
Gimme a break.
I think there are probably three cases here:
(1) You have perfectly good software and your only motivation to rewrite is infatuation with how great Rust seems. In this case, yes, that's unproductive. (Infatuation-driven design is poor engineering and is a fairly widespread problem in the software industry.)
(2) You already have other reasons to want to rewrite (design could be improved, code is in a poor state, etc.), and you're deciding between languages. In that case, a rewrite could be productive even if you didn't use Rust.
(3) You have mature software that appears relatively free of obvious, known bugs, but it's written in an unsafe language, and it would be valuable to have the additional confidence that a safe language could provide. Security-sensitive or other critical applications are where this is most likely to make sense.
at least Rust -> C ffi is zero cost (usually)
Mature software can have a bunch of hard to fix bugs that you can't motivate anyone to work on any longer. If you want to continue making progress you have to route around the damage somehow.
It would be wise to ignore the language distraction and remember that most urges to rewrite are just flat out wrong. The standard and well respected advice about rewrites doesn't go out the window just because of a fad language, if anything they apply more strongly.
Much of what people perceive as cruft is most of the actual value-- the stored knowledge from years of experience using the software in practice. Few "clean" programs are actually complete.
But I usually append that quote with "but then again, you're not Tridge" :)
Kernighan and Plauger said "don't patch bad code: rewrite it". That is definitely true. And there's a lot of bad code.
There is some concept[6] of converting Java to Rust, but it is far from being as useful as c2rust tool. Would be nice to integrate with c2rust somehow, if someone wants to help.
[1] https://github.com/immunant/c2rust
[2] https://c2rust.com/manual/c2rust-refactor/index.html
[3] https://github.com/tectonic-typesetting/tectonic/issues/459
[4] https://github.com/crlf0710/tectonic/tree/oxidize
[1] https://github.com/duneroadrunner/SaferCPlusPlus-AutoTransla...
[2] https://github.com/duneroadrunner/SaferCPlusPlus-AutoTransla...
> However at best, the temptation to RiiR is unproductive
> A much better alternative is to reuse the original library and just publish a safe interface to it.
Just as models are a lower dimensional representation of a more complex problem, [1] I have to re-iterate that there are no truth(ism)s in software and as in all engineering, there are trade offs. RiiR can and often is a valid choice. The author talks about "introducing bugs", under an engineering approach to a rewrite, more like a language-port, rewrites can _find_ a lot of bugs. One such way is as follows.
1. keep the same interface in the new system, client code should work against either system.
2. have a body of integration tests, capture these from the field OR write a small collection of orthogonal tests, somewhere between unit and integration.
3. use the tooling, bindgen, etc and generate as much of the interface programmatically as possible.
4. iterate on the port, doing differential testing against both systems.
Doing a language port is comparable in work to long term maintenance and refactoring of an existing codebase. As the tooling gets better, RiiR will be more of a smooth oxidization.
If you own the C/C++ that could possibly get rewritten, I'd think one needs to rationalize NOT RiiR. Better tooling (IDE, build), perf tooling is improving, low bug count, increased team velocity, build is improved, etc.
But the biggest reason to RiiR is safety. Integrating a body of C/C++ code into your rust codebase introduces a huge amount of unsafe code, much worse than surrounding your entire rust program with unsafe { }.
At least do what is outlined in the article but ALSO compile the native code into Wasm and run it from within a sandbox.
edit, something like a library for reading (parsing) a file format, should absolutely be RiiR or run from a sandbox. Data parsing and memory corruption vulnerabilities are excellent dance partners.
[1] see PCA https://en.wikipedia.org/wiki/Principal_component_analysis
TBH I don't see a future where RIIR every piece of software would make sense, and I think you sort of put it in there somewhere in your comments.
But yes I do see that for many it would and for those, RIIR would IMHO be an incremental process of oxidizing your project to the point that there's nothing left but Rust.
Writing something from scratch and waiting for it to finish is the biggest problem we face, cause eventually the will to continue with the effort just dies. But creating a meaningful ground not just some "safer" bindings does help in inviting others to help share the effort as well.
So hopefully people will be smart and identify what they should do for their projects, should it be a rewrite or should it be a new feature that you write in Rust. :)
And no I don't think bindings are the solution to anything, they serve no purpose in Rust community as a long term measure, you shouldn't have to sacrifice safety and/or performance incase of high level language.
It doesn't necessarily have to be Rust, and it's a process that may take decades. But I really do think that absolutely everything should be rewritten in a memory-safe language (i.e. not C/C++).
The vast majority of security issues are either stupid misconfigurations or memory-safety issues. And currently most security initiatives are undermined by the fact that they're resting on insecure foundations. Imagine if our computing platforms were truly secure. It would be a revolution, and IMO it's a revolution that's coming.
Edit: For the uninitiated, I am baffled by the fact that people think Java and .NET are memory safe.
Also if you meant that Java and .NET should be replaced by something memory safe.. I am absolutely for it. :)
Not only they have automatic memory management, they have an industry standard memory model for hardware access, adopted by C and C++ standards, used as inspiration for std:atomic<> on CUDA.
While you keep being baffled, where is Rust's memory model specification?
Just scroll up, I explicitly mentioned how it isn't. I just want to say, Java and C# isn't as safe as most people think.
Define industry standard, cause nothing points me towards industry standards actually being better than not industry standards.
Has everyone forgotten these http://cs.oswego.edu/pipermail/concurrency-interest/2017-Dec...
Nevertheless who knows I might not know anything, and I love being proved wrong when it makes most of the world a safer place (quite literally)... :P
Rust statically prevents inappropriate unsynchronized accesses for arbitrary APIs (for instance, the compiler will emit an error when attempting to mutate a non-concurrent object/data structure from multiple threads). Those VMs make unsynchronized mutations of individual memory locations work just enough to not be unsafe, but still allow arbitrary unsynchronized operations. These will likely end up with an incorrect result, even if there's no memory unsafety, and may completely violate any internal invariants of the object or data structure. This latter point means one cannot write a correct abstraction that relies on its invariants without explicitly considering threadsafety (it is opt-in safety), whereas Rust has this by default (opt-out safety).
Rust has essentially adopted the C/C++11 concurrency model.
Not saying that it isn't important though.
For instance, the type representing a handle to the external resource can have limited constructors and limited operations, that will enforce single threaded mutations or access (including things like "this object can only be destroyed by the thread that created it").
Borrow checker doesn't help at all with process IPC.
Plus all ML derived languages have good enough type systems to model network states, while enjoying the productivity of automatic memory management.
That said, I think containing C/C++ to a Wasm sandbox that can be integrated transparently with Rust would be a gigantic win for security, correctness and the Rust ecosystem.
Why you may ask..? The answer is simple cause there is always a cost. So the question is what is easier or when can one say the cost of using a "safer" language out weighs a "less safer" one..
Cause otherwise Rust is also too damn unsafe, just move to something completely safe Idris could be a good starting point. /s
Putting the sarcasm aside. Rust is does not guarantee Memory safety, it tries to help with it. (Check Reference Cycles) The safest feature of Rust would be it's protection against Data Races.
So at the end it's all about the trade-offs and deciding what you are willing to sacrifice and what you aren't...
Rust isn't perfect, and perfect automatically-checked safety probably isn't possible, but it's dramatically safer than C/C++, and cuts down the amount that would need to manually audited to the point where it would be feasible to do it comprehensively.
Of course ideally you get rid of all C/C++, but that may not be feasible. Combining Rust with some C/C++ still adds the benefit of Rust safety to the new code you write. The contained C/C++ is like the unsafe section of Rust code, a place where you have to be more vigilant, but you're slowly constraining the space where these issues can arise. And you can iteratively only replace those parts that cause a lot of issues.
Using standard library types, actually integrating sanitizers into CI, enable bounds checking (even on release builds) and above all avoid C style coding.
Naturally this doesn't work out for third party libraries that one doesn't have control over.
So from a business point of view it boils down how much one is willing to spend re-writing the world vs improving parts of it.
C++/CLI, C++/CX, GC pluggable API introduced in C++11, C++ Builder VCL in ARC mode, Unreal C++ managed classes.
https://docs.microsoft.com/en-us/dotnet/standard/language-in...
https://docs.microsoft.com/en-us/cpp/dotnet/managed-types-cp...
Then you can also use regular low level stuff, but in that case you will get mixed mode Assemblies (other forms are now deprecated).
https://docs.microsoft.com/en-us/cpp/dotnet/mixed-native-and...
C++/CLI is basically a set of language extensions, just like clang and gcc have theirs.
Right now it supports up to C++14, if I am not mistaken.
Many don't seem to realise that CLR started as the next evolution of COM, with the required machinery to support VB, C#, J# and C++, alongside any other language that could fit the same kind of semantics.
https://source.android.com/devices/tech/debug/asan
https://android-developers.googleblog.com/2019/05/queue-hard...
Also FORTIFY has been being enabled across the codebase and will become default going forward.
https://android-developers.googleblog.com/2019/10/introducin...
Future Android devices on ARM will make use of memory tagging.
https://security.googleblog.com/2019/08/adopting-arm-memory-...
Guess what, we even used to have a famous magazine called "The C/C++ Users Journal".
100% agree. We had a C++ service that made heavy use of libcurl. A particular release of libcurl introduced some memory safety problems that caused frequent segfaults for us. These memory safety bugs were eventually fixed in another release, but it scared us enough that we investigated rewriting the service in Rust.
After successful experimentation with some prototypes, we eventually rewrote the service in Rust, auditing the unsafe code in our dependencies (of which there was very little). No segfaults ever since.
Side benefit: Since Rust networking libraries tend to have strong async support, and since some of our C++ libraries performed synchronous networking operations, we saw a big improvement in performance. The number of threads needed dropped by 5x.
This is not a good reason if there are cheaper ways to get safety (at least for existing codebases). Indeed, there are sound static analysis tools (like TrustInSoft) that guarantee no undefined behaviour in C code. Using them is not completely free, and it may even require adding annotations to the code or even changing the code, but it does seem significantly cheaper than a rewrite (in any language). Such sound static analysis tools are already being used for safety-critical systems in C, and I believe they are more popular than a rewrite in Rust in that domain (where there are reasons not to rewrite in Rust other than just cost, though).
> and it's far from clear that fixing all the bugs, or even all the undefined behavior, in an existing C library will be less effort than rewriting it
It's pretty clear to me.
> Consider that one of the worst security bugs in history
I'm not sure what bug you're referring to, but if it's Heartbleed, than that was an undefined behavior bug. Of course, a functional bug can be introduced at any time, including during a rewrite in another language.
As for the cost of rewrites, there's a lot of evidence from software project metrics that the cost of modifying software can easily exceed the cost of rewriting it; see Glass's Facts and Fallacies of Software Engineering for details and references. Also, though, it should be intuitively apparent (though perhaps nonobvious) that this is a consequence of the undecidability of the Halting Problem and Rice's Theorem — it's impossible to tell what a given piece of software will do, which means that the cost of reproducing its existing behavior in well-understood code is unbounded.
Except that Heartbleed is the only OpenSSL bug I've heard of :) Also, I don't know who Kurt Roeckx is.
> there's a lot of evidence from software project metrics that the cost of modifying software can easily exceed the cost of rewriting it
But we're not talking about arbitrary modification, but about, at worst, fixing undefined behavior, which requires only local modifications (or Rust wouldn't be able to prevent that either). As an ultimate reduction, you could choose to rewrite the software in C and still use sound static analysis to show lack of undefined behavior.
> which means that the cost of reproducing its existing behavior in well-understood code is unbounded.
Yes, but that still doesn't mean that a rewrite is cheaper. Also, while your conclusion is correct, your statement of Rice's theorem is inaccurate: it's impossible to always tell what ever piece of software would do. It's certainly possible to tell what some software would do, at least in some cases, or writing software would be impossible to begin with.
BTW, if you're interested in the theory of software correctness, you might be interested in this blog post of mine, that lists relevant results: https://pron.github.io/posts/correctness-and-complexity
Your blog post looks very interesting indeed! I will read it with care.
I do think there's a subtle point about modifying software. Not just any modification of the software that lacks undefined behavior will do; we want a modification that preserves the important aspects of the original software’s behavior. Not only is this easy to get wrong—as shown spectacularly by the OpenSSL bug (which you've presumably looked up by now), but also, for example, by the destruction of the first Ariane 5—but there is no guarantee that it can be done with purely local modifications, even if the final safety property you wanted to establish can be established with chains of local reasoning.
I do agree that sound static analysis of C that is written to make that analysis tractable is just as effective as rewriting in Rust. Not only can such analysis show the absence of undefined behavior, it can show arbitrary correctness properties, including those beyond the reach of Rust’s type system. Probably the strongest example of this kind of analysis is seL4, although now its proofs verify not only the C but also the machine code, thus eliminating the compiler from the TCB.
As to seL4, it isn't exactly similar to sound static analysis, as the work was extremely costly. All of seL4 is 1/5 the size of jQuery and it's taken years of work. But it also includes functional verification, not just memory safety. In fact, it is among the largest programs ever functionally verified to that extent, and yet it was about 3 orders of magnitude smaller than oridnary business software, roughly the same verification gap we've had for decades. We don't yet know how to functionally verify software (end-to-end, like seL4) of any size that's not very small.
Anyway, Rust offers a much more limited form of assurance, and sound static analysis tools offer the same, and at a lower cost for existing codebases.
Also, did you catch the part where the point is about how expensive it is?
Who is 'we'? And yes, we can predict exactly what the output of a given program for a given input is, for the vast majority of cases. All you have to do is run the program.
> The mere fact that compilers exist is a meaningless form of "knowing what the program will do", only superficially relevant.
You think static analysis, type checking, intermediate representation, optimization, the translation of the program with exact semantics into another language, etc. - is 'superficially relevant' to understanding a program?
> You can't solve the halting problem with compilers
Now that's pretty irrelevant.
> Also, did you catch the part where the point is about how expensive it is?
Did you catch the part where I was only commenting on a specific part of the comment? But tell me, how expensive is it?
Iiuc tools like TrustInSoft are for situations where you need not just safety, but "reliability" (i.e. no crashes, no exceptions). That doesn't really scale so well to larger applications.
> at best, the temptation to RiiR is unproductive (unnecessary duplication of effort)
For tiny, mature libraries, this might be true.
But a rewrite in Rust can also vastly improve maintainability. If a library is under heavy maintenance (and will continue to be), your investment in rewriting it is likely to pay dividends by saving developers -- especially ones who are new to the library -- a lot of time.
Yes you could rewrite them in Rust, but considering these are tools which are used all over the ecosystem and have massive momentum behind them, your time would be better spent building cool things on top of existing work than reinventing the wheel and trying to convert the ecosystem to something which will (initially, at least) be an inferior product.
The original statement was deliberately opinionated and extreme, but I still feel like the pragmatic approach of reusing existing libraries instead of rewriting them is the best one for the short/medium term (jury's still out on the long term costs/effects).
But reading code is hard, it's easier to project your sense of confusion onto your predecessor than own it yourself, and once you rewrite the code in Rust and start to understand the true complexity that the previous code had to work around, you'll leave for another job (this time with Rust on your resume).
Doubly true if you're not all-in on one of the major Bigcorp silos (Java, C#).
I occasionally check their LinkedIn and they're still spending 12-15 months every time at companies, being CTO or Big Data Strategist or whatever.
Mostly good experiences when it's a specialist gig that happens to require a specific technology. Like some parts of a stack aren't interchangeable and aren't easy to pick up overnight.
Sometime I catch them a few weeks before they quit and I get to ask why they developed a low level http server in java interpreting controller code written in javascript that nobody understand but them. "It's much more efficient than using php or nodejs or higher level java" "you benchmarked ?" "I have to go do more knowledge transfer, bye"...
I dont know what I'd do if they were allowed by the company to do Go or Rust :D Guess I'd redo their wheel reinvention in the majority language of the team... increasing the overall constant migration I see again and again...
I wonder if, as a stereotype, it's unique to the software industry.
If you could rewarded with an incredible salary for sticking at the same company for 10 years and making a great, reliable platform using "boring technologies" then more people would probably do it.
Instead to raise your salary you gotta jump jobs every 2 years
> Instead to raise your salary you gotta jump jobs every 2 years
I believe the majority of America stays at their employers. It just seems to be the en vogue style for this group. I've made careers at my past two employers. I like to think I've made out very well by sticking to the same company for 4 years, going on 5, much better off than if i joined a startup working slave labor for much less total compensation only to have my shares diluted once an exit happens.
I've done it all at this point in my career, and sticking to a stable, cash-positive business is the best option at this point, IMO of course. Everyone likes to think they'll be the special 1% to make that big exit but then again our generation (millennials) were raised to believe we were special so it makes sense why people go chasing the dollar.
Your employer takes a little more time to adjust, but in the end, people end up where they should be over time.
The problem is that the elite quitter is usually incompetent. They are capable of learning some new things. But they are incompetent in the sense that they are incapable of seeing that the new thing is often the same as or worse than the old thing. They are incompetent in the sense that they have a low attention span and do not follow up on the things they do so they never learn from their mistakes. They often love complexity etc etc
So the hiring team cannot see these details directly. It's up to them to have competent interviewers to flush this out. But of course they are often in awe of the latest fad themselves and so on it goes.
Developers have a much longer lifetime of continuous learning, opportunities to get better, and focus on being the best they can be.
But at the end of the day, when a developer or player no longer offers their organization value, they will be cut loose. This is hugely destructive to that person's ability to care for themselves and their families. Is it any wonder that they want to do whatever they can to secure a stable future for themselves?
I think I must reject the premise a bit, unless I misunderstand how Rust works.
If you write a safe wrapper around unsafe code, is it not still the case, that if the unsafe code bombs out, it will take the safe Rust down with it, engulfed in shared flames?
(Unless you spawn unsafe code in its separate process or something elaborate like that.)
(Then again, from a practical standpoint the reasoning may be good. Maybe, even likely, the C library is very good and very battle tested. And very likely, my novice would be re-implementation in Rust would not be very good.)
This seems like the cleanest way to move a codebase forward without too much regression work. It reminds me of how TensorFlow 2.0 still contains tf.compat.v1 APIs so devs can write forward-facing code while maintaining existing modules.
You'll need to do this anyway if you want an actual library that can be independently packaged, because that means you have to use the system ABI in the library itself. Rust does not provide a stable ABI.
Do we admit that this is a factor in the decision to rewrite vs. reuse?
Allow me to introduce myself ;) I mostly work on new products and rewriting to me just a waste of time. I am interested in creating features that give actual ROI. The only time I rewrite a piece of code is if said piece presents a major problem (bugs, performance etc)
Wrapping your buggy C library in Rust isn't going to give you the kind of security against malicious data that a rewrite in Rust would give you.
https://medium.com/dwelo-r-d/using-c-libraries-in-rust-13961...
Autoconf creates a portable shell script, right? No need to have it installed just to build the project.
I'm just being pedantic. It's a problem...
If you want to natively compile on a fairly recent linux with the common standard libraries autotools works well. Even though autotools was written to work around these differences, in most cases the developer didn't hook into autotools detecting that difference and so it doesn't work.
CMake based projects almost always cross compile easily.
Why they published it like this is beyond me, however. The whole point of autotools is not making your users install them (unlike CMake and virtually every other build system for C there is).
No, parent answers the question, "What is the difference between the git tree and a dist tarball?" Which is self-evidently (and unhelpfully) "you don't need autotools installed with the dist tarball". My questions are: Why does this difference exist? Why would you not track everything necessary to build a library in git? How is the dist tarball built differently than just zipping the current git tree?
"It" in my post referred to the Rust cargo package in question. The package ships configure.in, but not the generated configure. You do need autoconf for configure.in -> configure. Similarly, it ships Makefile.am, but not the generated Makefile.in for use with configure.
Note how I said "The whole point of autotools is not making your users install them".
Did you happen to read the book? We cover strings early because this is a common pain point.
Also, if you have something more detailed than "get strings out of libraries", I can give better advice. It's tough to tell what the actual issue is.
use sha2::{Sha256, Digest};
fn main() {
let mut hasher = Sha256::new();
hasher.input(b"hello world");
let result = hasher.result();
??
}
How do I get a string with the hash value? Which chapter of which book should I read?The result of a hash is never a string, so it makes sense that you'd need another step to print it as one.
In this case, the hash ends up being raw bytes. So the question is, what are you trying to do with this hash? Being a bunch of bytes, this may not be valid UTF-8 (and in this case, is not ASCII, let alone UTF-8). One option is to encode to base64:
use sha2::{Sha256, Digest};
use base64;
fn main() {
let mut hasher = Sha256::new();
hasher.input(b"hello world");
let result = hasher.result();
let encoded = base64::encode(&result);
assert_eq!("uU0nuZNNPgilLlLX2n2r+sSE7+N6U4DukIj3rOLvzek=", encoded);
}
But yeah, you can't exactly "get a string" out of a random bag of bytes unless you know how you want that bag of bytes to be represented, encoding wise. That's not exactly a satisfying answer, but such are strings!EDIT: sorry I don't know which docs I was looking at, I can't find it in the current version.
There are some examples of printing output using formatting strings in the RustCrypto readme:
let hash = Blake2b::digest(b"my message");
println!("Result: {:x}", hash);
https://github.com/RustCrypto/hashes#usageIf you're looking to get the actual string value rather than just print it out, then take a look at the format macro. It uses the same format strings / traits etc as println, so wherever you see println you could drop in format instead:
let hash = Blake2b::digest(b"my message");
let text = format!("{:x}", hash);
https://doc.rust-lang.org/std/macro.format.htmlRust does have a pair of traits that can be used, Display [1] and Debug [2]. Display must be implemented manually by the author of a type, while Debug can be automatically derived if all the members of a type implement Debug. Like this:
#[derive(Debug)]
struct Foo {
}
[1] https://doc.rust-lang.org/std/fmt/trait.Display.htmlIt is unfortunate that you’ve been downvoted for asking a question.
It is unfortunate how HN is nowadays but I don’t care if I can still learn a lot from people like you.
In Rust a string is a UTF-8 sequence. Typically human readable. A byte is generally represented by the u8 type, and a collection of bytes is either an array or vector (e.g. [u8] or Vec<u8>). To create a human readable form of an object in Rust you'd typically use the Display ({} format specifier) and/or Debug traits ({:?}). A string (e.g. &str or String) is NOT a collection of bytes. There are exceptions however with OsString/OsStr representing something closer to a collection of bytes and CString/CStr representing a bag of bytes.
Your comments seem a bit XY-ish to me. What are you trying to solve by converting things to strings? For debugging or human readable output Array, Vec, and u8 all implement the Debug trait. u8 also implements UpperHex and LowerHex so you can get a hex formatted version as well (e.g. format!("{:02X}", bytes)).
For the (MD5) hash case, as others have pointed out you're looking at a base 64 encoding which you'll often have to do on your own depending on the library you're using.
Now if you're struggling with owned vs borrowed strings that's a whole other matter. You can insert some magic into your functions with generics and trait constraints and into your structures with the Cow type (clone on write).
str::from_utf8(result.as_slice())
https://docs.rs/generic-array/0.8.2/generic_array/struct.Gen...
let my_rust_string = str::from_utf8(result).unwarp()
If there is a possibility of an error, (you read it from file, network ...) you can do: let my_rust_string = match str::from_utf8(result) {
Ok(str) => str,
Err(err) => panic!("Invalid utf8: {}", err),
};
Or you can do lossy conversion (will have that question mark on chars that it cant decode) let my_rust_string = String::from_utf8_lossy(result);
edit: Ok so i read sha2 api wrong. result() is raw bytes and result_str() is hex encoded string. Above wont be much use for either. And others have already explained what to use.I just stared programming in rust several months ago, and that was my pain also, since where i live we still have some other encoding in use.
What many low-level data formats deal with are raw bytes ("byte strings" in some contexts). The output of a hasher is random binary data.
What most humans are accustomed to are some form of encoding that makes it easier to spell out the octets. Hex and base64 are common. But they are not the native representation that the machine deals with.
With almost everything being explicit in rust you need an explicit conversion from bytes to hex, base64, urlencode or some other format.
Those conversion methods (with their contract often expressed as a separate trait) may be provided by the same crate or you may have to pull them in from a different crate.
A hasher does not need to provide this itself because its output types can be enhanced by foreign traits.
The O'Reilly book is good too.
Nonetheless a cookbook for converting strings for different use cases isn’t a bad idea at all.
https://web.archive.org/web/20191023134247/http://adventures...
What are examples of the unheard of things?
In another language you'd risk having a pointer to invalid memory when the function returns if the pointer escapes the function, and so it is advised to use pointers to objects on the stack sparingly if at all.
In rust, the compiler will statically determine if a pointer to an object on the stack could escape the function, and if so will fail to compile.
Just a nitpick, this is only true for non-garbage-collected (or refcounted) languages. :-)
A garbage-collected language wouldn't let you allocate such items on the stack in the first place.
GC languages tend not to let you choose where to allocate in the first place, and have obligatory heap semantics for reference types. Some compilers (including Go and Java, to my knowledge) will attempt to optimize the implementation to stack allocation when possible using escape analysis, but this is less precise, and more opaque, then Rust's allocation. And you won't get feedback if it stops working.
So, by GP's comment of "unheard of (or just plain dangerous)", the "unheard-of" applies to GC languages and "plain dangerous" would apply non-GC languages.
:-)
> A garbage-collected language wouldn't let you allocate such items on the stack in the first place.
I get your point, but I don't think this is entirely true.
When people think of call stacks in the general sense, there's a major distinction between CPU stacks and managed stacks.
For example in both Java and Go, stack allocations are determined at compile time, and heap allocations at runtime. (Obv, stacks here are not actual CPU stacks, just managed memory stacks depending on the implementation.)
Modula-3, Oberon, Oberon-2, Active Oberon, Mesa/Cedar, Component Pascal, Eiffel, Sing#, System C#, C#, Swift, D, Nim
Reference-counting, a form of GC very commonly used in Rust and C++, is inefficient compared to more intrusive schemes, particularly where there may be contention for the count, but may be applied selectively, e.g. never on critical paths, so that the inefficiency has zero impact on overall system performance.
Obligate-GC advocates like to point at custom benchmarks showing overhead at a small percentage. They are invariably lying, by reporting only the time that the profiler says the program counter is pointing to GC code, and hiding the much-larger results of loss of cache locality on the system as a whole. Typically they don't even know they are lying, which should not give one much confidence in their engineering judgment.
Yep, really lying.
https://www.reddit.com/r/programming/comments/3t50xg/more_in...
They admit it wasn't quite full-GC stuff. It was close to the C++. It did use the language safety and parallelism to its advantage on top of the Midori OS. They ended up improving performance over the original.
Still worth mentioning even if not fully-GC. I mean, they could've always used a real-time GC if that was important. There's already commercial and academic ones. For some reason, these projects never try to do that. I think even those developing GC'd systems might not know about RT designs.
e: Also, many languages allow the user create user-defined value types, which are copied rather than referenced. Variables containing these can be stack-allocated.
At least in the C/C++ world, this is not true. Where possible, we prefer stack allocated objects. Stack allocation is fast and usually gives you better cache performance than heap allocation.
Rust's lifetime/borrow checker provides compiler support for well-established best practices amongst professional C and C++ users.
Ummm.... any pointers as to what the author is talking about?
This is how you avoid growing as a programmer.
Our job is not to program crap again and again, it's to deliver some value for the high fees we charge. Rewritting a working library to migrate language is really hard to defend...
You may question the usefulness of having “x but in Rust” and that is fair. However, people keep answering the reasons why you might do something like this and yet like amnesia the same bad opinions come out again next time.
— Sincerely, someone who also ships.
To be clear I’m not suggesting you should always rewrite something versus just wrap it, not at all. I’m just saying the reasoning for not doing it being “you didn’t write the original so what do you know” is very offensive to me; how are you supposed to learn?! People write NES emulators all the time even though they already exist. While most are purely for practice it doesn’t really matter: they wanted to write it, they wrote it, they learned. Discouraging people from writing something that they want to write is upsetting to me.
Most programmers are programmers for hire, and it would be a waste of everyone's time to re-implement (for example) libdispatch in Rust. If someone wants to write it on their own time for their own edification and no other purpose, fine -- but half-assed internal implementations of common functionality are a significant drag on the lives of other working programmers.
Also, people build dependencies on effectively “hobby” projects all the time. Open source in particular works well this way because the dependents have a good reason to contribute back; the users become stakeholders and contributors. The same may not be true of internal software, but again, this has little to do with the concept of rewriting something in Rust.
This is really diverging into a discussion that has nothing to do with the original paragraph I addressed.
Rewriting in some language is a matter of skill and preference. To even consider a rewrite there should be a rational reason, i.e. it is broken or can't work in the ecosystem.
For fun and learning you can do what you want, of course.
So, an article titled, say, "how not to paint your house" would be expected to outline things to avoid when painting your house, and "how to not paint your house" could just be "go for a bike ride instead."
'How not to x' articles are typically self-deprecating articles that do x, but are critical of how they approached it and lessons learnt
'How to not x' is more explicit, i.e. How to avoid doing x
It also isn't the wording used in the original article title.
I think this kind of title is trying to do exactly that, or at least, that's how I took it.
> It also isn't the wording used in the original article title.
This is a good reason to switch it, for sure. I didn't realize this, if you had led with that, I wouldn't have said anything :)
I took the exact opposite: This article is describing an alternative to RiiR, so that RiiR is avoided when not needed.
What I was trying to say is that I think the title is making a joke; it's saying "how to re-write something in Rust badly", because well, it's not re-writing it in Rust at all.
Actually, this is the case. It describes the bad way (and bad reasons) to rewrite something in Rust.
How to not X = How to avoid doing X
How not to X = The wrong way to do X
...this is tongue-in-cheek comment. I do agree with the author, "rust rewrite" alternative of wrapping c library is way better in many cases for well established or legacy libs. However it depends on the momentum you can bring - as pure rust implementation for widely used areas can benefit from compiler grinding through it. I can see rust written low level libraries used across all high-level languages in the near future.