Construct-JS – A library for creating byte level data structures
github.com
github.com
Both implementations are very optimized and fast. Also DataView offers working with 64-bit values using the new BigInt type [2].
Here [3] a short example on how you can quickly build C-like structures in JavaScript
[0] DataView: https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
[1] TextEncoder: https://developer.mozilla.org/en-US/docs/Web/API/TextEncoder
[2] BigInt: https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
[3] C-like Structures: https://jsfiddle.net/5tz0hsjq/
We looked into this problem years ago as part of our spreadsheet parser/writer library (open source version https://github.com/sheetjs/js-xlsx) and ultimately opted for lower-level functions that act on Buffers/ArrayBuffers/Arrays. Performance-wise, creating a new ArrayBuffer / Buffer for each field (which you are currently doing in https://github.com/francisrstokes/construct-js/blob/master/s...) is incredibly expensive. When we did performance tests years ago, allocating in 2KB blocks and manually orchestrating writes was nearly 10x faster than individual field-level allocations and concatenations both in the browser and in NodeJS.
Ironically, the ZIP file example referenced in the README alludes to another pitfall in the approach: the actual DEFLATE algorithm used in compression actually requires unaligned bit writes. See section 5.5 of the current APPNOTE.TXT for more details: https://pkware.cachefly.net/webdocs/casestudies/APPNOTE.TXT
As for DEFLATE, I'm working on an automatic bit level structure at the moment which would allow for unaligned structures. Should be in the lib in a couple of days.
Things have probably improved in Buffer/DataView/TypedArray land thanks to the push for WASM, but the allocation overhead still will be a lot higher and involve a lot more work to allocate them.
const localHeader = Struct('localHeader')
.field('header', DoubleWord(0x04034b50))
.field('sharedHeaderInfo', sharedHeaderInfo)
.field('filename', filename);
There's already a construct for key-value pairs in JS, it's called an object: const localHeader = Struct('localHeader', {
'header': DoubleWord(0x04034b50),
'sharedHeaderInfo': sharedHeaderInfo,
'filename': filename
});The interface that would be most flexible would probably be a list of alternating key values, e.g. ['fieldname', datastruct(), 'fieldname'] since then you could do all the list things to it.
[1] https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
In practice the order will be maintained in the major engines, but that's still not a good reason to do rely on it.
I might consider a static `Struct.fromMap(name, mapObject)`, but as a complimentary API.
- By providing a 'builder' you add a clear API to the struct object's constructor. It doesn't make any more or fewer claims of capibility than it needs to. By exposing the internal storage you break that level of abstraction.
- Even assuming that the JS object type _is_ ordered, the ADT of the Struct type may be subtly different to the ADT of the object. By coupling the two you may prevent future flexibility.
- To that point, the 'field' method _does_ do something with each invocation (forceEndianess) as it adds it to the internal storage. If you passed in a map object it would still need to traverse each key to do that work, which (knowing JS objects) might make it more complicated.
[0] https://github.com/francisrstokes/construct-js/blob/master/s...
In other words, it's still possible to pass in a separate js object instance containing the configuration.
Yet I'm not sure about the whole of your arguments, so I can't tell if it would actually make sense to wrap the default js builder pattern implementation (the javascript object initializer syntax aaaaaaaaaaab refers to).
However, I do method builders often overused in API's where a simple singular builder method combined with builder classes suffice and simplify.
I'm reluctant to spend too much time arguing over a fine point. It could be done one way or the other. I was just addressing aaaaaaaaaaab's question.
One thing I love about Rust is that it uses u16, u32, u64 etc for unsigned and i16, i32, i64 etc for signed, which is about perfect - clear, concise and future-proof. That would be perfect for this library.
Yes, even good old C has uint16_t and int16_t for this. I use these exclusively for embedded work because we care about the size of everything. Also agree that Rust gets it right by using a single character with the size: u16, i16.
It's funny because C opted to leave the number of bits machine dependent in the name of portability, but that turns out to have the opposite effect.
60 bit words, 18 and 60 bit registers, 6 bit characters.
It would interesting to see a CDC 6400 implementation of "like C, but we ignore the bit about character set and also the part where the compiler is a evil genie trying to screw you over", though.
That depends on what you consider the portable part. In the era of proprietary mainframes operating on proprietary datasets, data portability probably didn't matter as much as code portability to perform the same sort of operation on another machine.
C's size-less `int`, `float`, etc allows the exact same code to compile on the native (high speed) sizes of the platform without any editing, `ifdef`s, etc.
(Side note: That's what bothers me a lot about the language wars -- the features of a language are based on the trade-offs between legibility, performance, and the environment from the era they were intended to be used in. Often both sides of those spats fail to remember that.)
Firstly, no, good old C doesn't. These things are a rather new addition (C99). In 1999 there was decades of good old C already which didn't have int16_t.
It is implementation-defined whether there is an int16_t; so C doesn't really have int16_t in the sense that it has int.
> It's funny because C opted to leave the number of bits machine dependent in the name of portability, but that turns out to have the opposite effect.
Is that so? This code will work nicely on an ancient Unix box with 16 bit int, or on a machine with 32 or even 64 bit int:
#include <stdio.h>
int main(void)
{
int i;
char a[] = "abc";
for (i = 0; i < sizeof a; i++)
putchar(a[i]);
putchar('\n');
return 0;
}
Write a convincing argument that we should change both ints here to int32_t or whatever for improved portability.Yes, it does. The C99 standard has been around for 20 years. I started giving up on compilers that don't support it 10 years ago. I consider it a given for C.
>> It is implementation-defined whether there is an int16_t
No, it's not. That type is part of the C99 standard.
>> Is that so? This code will work nicely on an ancient Unix box with 16 bit int, or on a machine with 32 or even 64 bit int:
That's cool, your example only needs to count to 3. Any size integer will do. The problems arise when you go to 100,000 and that old 16bit machine rolls over at 65536 but the newer ones don't. Other times someone (me) may want things to roll over at the 16 bit boundary and we need to specify the size as int16_t rather than figure out if int or short or char is that size for each architecture. (and yes I know rollover is undefined behavior)
>> Write a convincing argument that we should change both ints here to int32_t or whatever for improved portability.
In your example it doesn't matter. I'd argue that at least giving the size of your integers some thought every time is a good habit to get into so you don't write non-portable code in the cases where it does matter. You are free to argue that it's too much effort or something, but I'd invite you to argue that "i16" is more effort to type than "int".
There, we are running into the question of: on that system, can we define that array at all, if it has 100,000 characters.
> That type is part of the C99 standard.
Unfortunately, the standard isn't what translates and executes your code; that would be the implementation. The standard allows implementations not to provide int16_t, if they have no type for which it can be a typedef name. A maximally program can use int16_t only if it has detected that it's present. If an int16_t typedef name is provided then <stdint.h> will also define the macro INT16_MAX. We can thus have code conditional on the existence of int16_t via #ifdef. (Since we know it has to be nonzero if provided, we can use #if also).
> Other times someone (me) may want things to roll over at the 16 bit boundary and we need to specify the size as int16_t rather than figure out if int or short or char is that size for each architecture. (and yes I know rollover is undefined behavior)
Someone who knows C knows that unsigned short is 16 bits wide or wider, as is unsigned int. A maximally portable wrap-around of a 16 bit counter, using either of thise types, is achieved using: cntr = (cntr + 1) & 0xFFFF. That will work on a machine that has no 16 bit word, and whose compiler doesn't simulate the existence of one.
We can do it less portably if we rely on there being a uint16_t; then we can drop the & 0xFFFF. (It isn't undefined behavior if we use the unsigned type.) The existence of uint16_t in the standard is encouraging programmers to write less portable; they reach for that instead of the portable code with the & 0xFFFF. Code relying on uint16_t, ironically, is not portable to the PDP-7 machines that Thompson and Ritchie originally worked with on Unix; those machines have 18 bit words.
In computing, "word" is understood to be machine dependent: "what is that machine's word size?"
I used it in the past to read a proprietary file format and it worked well, but they also have quite a few predefined formats in their gallery. [2]
Edit: this provides a good explanation http://doc.kaitai.io/faq.html#_google_protocol_buffers_asn_1...
Seems weird that it's not even mentioned. I thought it would be a port.
[1] https://github.com/zandaqo/structurae
[2] https://github.com/zandaqo/structurae#RecordArray
[3] https://blog.usejournal.com/structurae-data-structures-for-h...
The idea is a good one for another reason, because you can use this API to define a data-structure which can be 'rendered' in different ways - which could be useful for porting data-structures between node, DOM and WASM.
Bravo!
Mnemonist [1] is a library which looks more like what you're looking for. There are some performance benchmarks here and there on the web, comparing the performance of their data structures with other....
If you are not the original developer, most likely source maps are not available to you. Then, wasm can be studied with already existing reverse engineering tools, like IDA Pro, binary ninja or radare2/cutter. Last one is even FLOSS.
https://developer.mozilla.org/Web/JavaScript/Reference/Globa...
- https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe... - https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe... - https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe... - https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe... - https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe... - https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe... - https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe... - https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe... - https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe... - https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
JITs can also use integer types internally when possible (small integer optimization).
Or is this more for the heavy lifting and backend uses of JS?