Here's my OCaml library influenced heavily by Erlang: https://bitstring.software/examples/
Here's my OCaml library influenced heavily by Erlang: https://bitstring.software/examples/
Consider, for example, how you might use a routine like memchr[1] if all of your public APIs are generic iterator adapters.
And then once you get into things like regex engines, modifying them to work on Iterator<Item=u8> is a highly non-trivial affair. It is of course possible to write a regex engine that works on such things, but it's going to be limited in performance or capabilities. The way to make regexes and streaming work together is probably something more like Iterator<Item=&[u8]> (which is perhaps roughly analogous to what Hyperscan does). You really want blocks of bytes, not one-byte-at-a-time.
[1]: https://github.com/BurntSushi/memchr/blob/8037d11b4357b0f07b...
So I don't see how your links are relevant here.
However, I do agree that Erlang's Bit Syntax is, by far, the best handling I've ever seen for slicing and dicing bits and bytes at the very lowest level.
See if you can guess correctly what this produces:
<<1:5, 2:4, 0xFFOO:15, "foo":binary>>You use a hex literal syntax that doesn't exist--16#FF00 is Erlang and how on earth did you get "O" instead of "0".
Once I fix the brokenness:
10> <<1:5, 2:4, 16#FF00:15, "foo">>.
<<9,127,0,102,111,111>>
Looks ... unsurprising? Maybe it would be nice to get a truncation warning on that 16#FF00:15, but otherwise ... fine?
This looks like an advertisement that Elixir kind of sucks and you should use straight Erlang instead?
Note from downthread: this is broken even in Elixir, so I really don't know what your point was.
If you mean this equivalent bit of Elixir, then I got it right:
<<1::5, 2::4, 0xFFOO::15, "foo"::binary>>
I think it's all good and sensible, except that "0xFF00::15" should raise an error. The bit syntax silently masks out some stuff which can hide some errors.
Ok. Maybe I'm just dumb. But I was really surprised that the bit order goes from msbit to lsbit (while the byte order goes from lsbyte to msbyte). Since the 15 bit element crosses over the byte boundary (and overflows it) the arrangement -- 00s in byte 1 and 0x80 0x3F in bytes 2 and 3... is not obvious, to me).
In particular this makes it tricky to reinterpret packed structs coming from c, when using NIFs
Specifically, what I would have wanted, was lsb-msb in bytes, and an overflow error, or syntax in the descriptor modifier that is required if you might overflow.
Or, this would all be solved if we used an RTL layouts like the divine one intended.
You get the bits in the order you ask for in the expression, and then it encodes each valuee as you ask. Note that the default encoding is big endian, so the value 0xFF00 comes out as the bit sequence 1111111100000000 wherever it is.
Now converting from a C-struct you absolutely do have to think about byte order, but erlang has you there too. Let's say you've got this:
#pragma pack(1)
struct {
uint16_t a;
uint8_t b;
uint8_t c;
uint32_t d;
};
Then all you do is: <<a::16-native, b::8, c::8, d::32-native>> = my_block_of_data
And then all of your values (a, b, c, d) will be correct.As I said above, erlang defaults to encoding/decoding each value as big-endian (e.g. <<0xFF00::16>> is equivalent to <<0xFF00::16-big>>. This is the default because big endian is used for the vast majority of comms protocols (in that context it's also called 'network byte order'). When interpreting values in memory though it can be either, depending on your processor (although mostly -little these days). Using -native as yoru endianness means "use big or little as appropriate for my processor". It's also a good signal that you're interpreting something from memory (as supposed to a protocol).