Parsing bitstreams with Nom
blog.adamchalmers.com
blog.adamchalmers.com
It's not quite as inspired, but I wrote a minimal bitstream cursor/parser from scratch in Rust a few months ago[1]. It's developed with LLVM bitstreams in mind, but the API itself is suitable for just about any bitstream.
[1]: https://docs.rs/llvm-bitcursor/latest/llvm_bitcursor/index.h...
https://docs.rs/bitmatch/latest/bitmatch/ works nicely but only for a single integer.
https://internals.rust-lang.org/t/pre-rfc-binary-patterns/89... didn't get much discussion but seems the closest to Erlang.
This is much easier than trying to design the entire parser on a whiteboard/paper and implement it in one big go, I think.
I've written 4(?) different parsers with different major versions of nom, and each has been easier than the previous one.
The main thing I'd like to see improved related to it would be `cookie-factory`, its basically symetric serialization library.
Then it hit me: If we ignore the odd choice of 3 & 5 bit fields, this is very similar to ASN.1
This was extremely convenient for me in the past.
Highly recommend both libraries!
use bitvec::prelude::*; // BitView trait, Msb0, etc
let my_input = vec![0b11100000u8];
let bitslice = my_input.view_bits::<Msb0>();
let my_num = bitslice[..3].load_be::<u8>();
// continue parsing after my_num
let nextslice = &bitslice[3..];
You would use Lsb0 instead of Msb0 if you wanted to parse bytes starting at the low bit.It would be more convenient with a peek/take style wrapper over the slice, but I don't think one is included in the crate.
I think Nom might be more useful as the format gets more complicated... the combinators can really save you quite a bit of time. Having things like many0, or delimited_pair, or length_count already tested and documented saves time compared to implementing + testing it yourself. It's cool seeing how many different approaches you can use to solve the same problem :)
> I think Nom might be more useful as the format gets more complicated... the combinators can really save you quite a bit of time. Having things like many0, or delimited_pair, or length_count already tested and documented saves time compared to implementing + testing it yourself.
Totally agree -- bitvec works for things that are easy to hand write a parser for. Nom is a lot more powerful.
I wonder if some unit test / fuzz tests in a less strict language could lead to the same amount of code quality yet at a higher level of work satisfaction while writing the code.
"Less strict" inevitably means not caring about a bunch of edge cases. It's a bit like "overflow is undefined behavior in C"; nobody wants to care about how "i++" is a statement that's (a) absolutely everywhere and (b) potentially invokes undefined behavior at runtime, so the only way to guarantee that out involves a lot more proof work somewhere.
That's an unnecessarily broad generalization to a concrete statement, and if it were true more people would write their programs in dependently typed programming languages or whatever.
> nobody wants to care about how "i++" is a statement that's (a) absolutely everywhere and (b) potentially invokes undefined behavior at runtime, so the only way to guarantee that out involves a lot more proof work somewhere.
In many domains, the high standards you're displaying here will almost immediately throw you out of business. In particular "i++" for most applications is just fine to write without any second thoughts. It's not even a good idea to write a test. Just run the program and fix it in the rare case something it doesn't work as expected (it's almost a safe bet the issue won't UB).
You're not taking a tank to go to the supermarket either, are you?
Programming language choice is greatly influenced by ecosystem factors. If you wanted to write a program for the browser, you had to write it in Javascript - so people went to considerable lengths to fit typed front ends onto Javascript.
> Just run the program and fix it in the rare case something it doesn't work as expected (it's almost a safe bet the issue won't UB).
> You're not taking a tank to go to the supermarket either, are you?
No, but if you're going into a live fire hostile environment you might want a tank, and the internet is such an environment. If there is a bug in your parser, and the parser is exposed to untrusted data, then that's a potential avenue for exploit.
Less of an issue in managed-runtime languages, but it's still possible to have semantic level exploits. I remember there was one for Apple which exploited the fact that there were different proplist parsers. https://daringfireball.net/linked/2020/05/02/psychic-paper
The idea of fighting ("[being] in a constant battle with") the compiler seems weird to me, the compiler is my automated helper, the stricter it is, the better I work and I thank it for correcting me.
No need to thank me.