Self Modifying Code as an alternative to macros
matklad.github.io
matklad.github.io
The author knows full well, and even acknowledges in the very first sentence, that this article is not about self modifying code as commonly understood. This article is about generated source code.
I wish there was a simple way in Rust to make proc macros optionally generate code that gets checked in, allowing for most of the macro costs to be shifted to dev-dependencies. While some macros might be generating code for very dynamic parts of applications, others like clap (CLI parser) are less likely to change and could benefit from it.
[0] https://epage.github.io/blog/2019/10/speeding-up-rust-builds...
It's not stabilized yet, but RFC3086[1] introduces macro metavariable expressions, which make counting trivial!
The example macro becomes:
macro_rules! define_error {
($($err:ident,)*) => {
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum Error {
$($err,)*
}
impl Error {
pub fn from_code(code: u32) -> Option<Error> {
match code {
$( ${index()} => Some(Error::$err), )*
_ => None,
}
}
}
};
}
You can see this in practice on the playground[2].[1]: https://github.com/rust-lang/rfcs/blob/master/text/3086-macr... [2]: https://play.rust-lang.org/?version=nightly&mode=debug&editi...
macro_rules! count {
() => (0);
($odd:tt ($a:tt $b:tt)*) => ((count!($($a)*) << 1) | 1);
(($a:tt $b:tt)*) => (count!($($a)*) << 1);
}
It uses bit twiddling to avoid generating very big AST and eventually crashing the compiler. Though I have to admit using `${count()}` is better.This describes an interesting approach that I hadn’t seen. I personally wouldn’t have considered putting the result in the same file because I usually like to gitignore any and all generated code, mostly to reduce phantom merge conflicts. However, some colleagues have the opinion that those conflicts are trivially resolved by rebuilding and thus don’t matter, and having diffs of the generated code show up in pull requests adds value.
What do others think, check in generated code or put it in gitignore?
This is how mistakes happen, and when dealing with things like serialisation is an absolute _minefield_
> nd having diffs of the generated code show up in pull requests adds value.
They do have a point here, but is there a reason that the generated code diff couldn't be done separately to being commited?
> What do others think, check in generated code or put it in gitignore?
Always ignored, no questions asked. Generated code is a build artifact, the same way an object file or a jar is. It just happens to be human readable.
Your build system should be 100% bulletproof and pick up changes to your input files and correctly regenerate the output files for you, anything less should be considered a p0 bugfix. If you need to cache these files for whatever reason, do so at a different level (sccache/container/whatever).
I'm not persuaded of the value of having diffs in source control, but i haven't looked at concrete examples. This would be an interesting post for someone to write.
For performance, i'd rather follow a distributed caching approach. Hash the true source, look for /some/nfs/mount/or/s3/path/${hash}.tar.gz, if it exists, unpack it into the generated source directory, if not, generate the source, tarball it, and upload that to the said path. Takes some care to hash the right things (needs to include the version and configuration of the generation tool, etc). Again, you could have a CI job checking that the caches are correct if you're worried.
Simpler to get started running the code, less compute wasted running generators (this is especially nice when doing something like bisect), and it's nice to be able to just see all your code, right there in the repo.
I actually think the biggest downside is the diffs -- often there's no easy way (e.g. in GitHub) to hide all generated code, which can be annoying.
Merge conflicts don't bother me at all, since they're so easy to resolve; I agree with your colleagues there.
(Though that would be trickier if you had files mixing generated and hand-written code! Contra the original article, my instincts/advice would be to always keep generated code well-isolated from hand written code, ideally using some standard convention [e.g., all generated stuff is in /generated/ path].)
Commands to take snapshots of every build ("git add *", "git commit", "git rev-parse" on the actual source repository, "date" etc.) are easy to script.
I have been relying on cargo expand for macro debugging, let me check what utilities rust-analyzer has for this use case. Thanks for sharing.
EDIT: wow, it works exactly as you advertised, thanks a bunch again for sharing, such a nice productivity boost!
I think this is a great investigation into a problem that's somewhat self-inflicted by rust, kind of like how so many of those GoF "design patterns" were simply a response to the lack of expressivity of Java.
For example, I think this problem doesn't even really make sense in Zig, where types are simply runtime values, and you have the full language available to you in a "compile time" pass. That approach is not without issues of its own (e.g. the article's point about IDE insight into the code, but writ even larger), but it's neat to see a different exploration in the design space, and how certain choices can make some issues just melt away.
Disclaimer: I'm both a rust and zig newb, and while I think I know roughly how I'd do this on the Zig side, I might be missing something.
I'm not too familiar with Rust so I'm not sure if the test code would be simpler than the macro solutions or the hack, but I think it would be.
Edit: you might could also do some shenanigans using integers, `as` conversions, and #[repr(u8)], but it seems like it would get nasty and potentially unsafe
The concerns of undefined behavior, data races, use-after-free, etc. seem a lot less consequential in test code.
unsafe { transmute::<u32, Error>(error_code) }
than it is to verify that the weird macro code in the article is correct. There is no data races, use-after-free, etc. in that block, it does exactly what the article needs. To me, Rust programmers seem to have a weird notion that just the word "unsafe" is enough to conjure demons and monsters into otherwise totally normal code.Personally waiting for his version of "Goto Statement Considered Harmful" but for the macros.
Contrived example:
import std.stdio;
mixin template Func(string s, string r)
{
mixin("string " ~ s ~ "()" ~
"{" ~
"return \"" ~ r ~ "\";" ~
"}");
}
mixin template Comb(string name, char c=',')
{
mixin(function string(){
import std.conv;
string r;
enum counter(size_t x=[__traits(allMembers, mixin(__MODULE__))].length)=x;
string suffix = to!string(counter!());
r ~= "mixin Func!(\"Foo" ~ suffix ~ "\", \"Hello\");";
r ~= "mixin Func!(\"Bar" ~ suffix ~ "\", \"World\");";
r ~= "string " ~ name ~ "() {";
r ~= "return Foo" ~ suffix ~ "() ~ \"" ~ c ~ " \" ~ Bar" ~ suffix ~ "();";
r ~= "}";
return r;
}());
}
void main()
{
mixin Comb!("Mixer");
writeln(Mixer());
}
It is just that the language has more features than C so they aren't needed as often. But in my eyes at least the mixin statement is basically C macros on steroids :-P.That industry is founded purely on fear and very little actual expertise. They should not fulfill an example role in your mind and nobody should actively listen to them when considering the future.
Unrelated, but I can't help but think this should use try_from or even potentially panic. If you get an error type you don't recognize, that could pose a risk.
At least with the manual approach you can decide if particular ranges are safe to ignore or not and do something different. Granted, I'm dreaming up scenarios that are very likely outside the author's intention, so I don't think my thoughts on the matter necessarily refute anything being said.
Personally, I’ve had cases where three or four places had to be kept in sync, and the list was changing frequently enough that something like this became useful.
Even then, it’s still a judgement call whether code generation or (if available) macros are worth the complexity. The decision might also heavily depend on the experience of the team.