Tny: A simple data serializer in C
github.com
github.com
The positives (like tny) were that the code was very straightforward and easy to use. And one of the negatives were that it wasn't "self describing" (which JSON is to some extent).
Of course the ultimate in binary self describing data is ASN.1 syntax :-) but lets not go there.
As such I doubt "the author read the XDR code", though perhaps you weren't wondering that at all, but rather wanted to mention XDR and ASN.1, in which case you should've really just done that. As it's written, it makes "the author" look like either an ignorant imbecile who doesn't know such basics as XDR or as someone who piggy-backed on XDR without acknowledging it.
A first step to improving this library would be to add a SAX style parser so I could build my own representation.
If you're into C++, take a look at benejson's pull parser. There is a C core but its presently very ugly and only intended for library, not application developers.
Often you can avoid almost any memory overhead. E.g. any structures that are allocated on opening a document, and deallocated on closure, can be allocated from larger buffers without keeping any information about the individual allocations. That can save anything up to 16-20 bytes per allocation with many malloc() implementations, and reduces typical allocation cost to incrementing a pointer and checking whether or not you need to allocate a new buffer (and the allocation cost for new buffers might be amortised over anything up to thousands of small allocations).
For a practical example, there's a font library that calls malloc() to allocate structures for every single glyph when opening a font. Most of the (thousands) of allocations are 4-8 bytes. Changing that to using an arena for an application I did, cut memory usage per font to about 25%-30% and cut load time for fonts to <10% by avoiding the malloc() calls.
You can achieve the same by throwing abstractions out the window and putting stuff in arrays etc. But arenas is often a very effective way of keeping the abstractions while effectively getting almost the same performance and memory usage.
Optimizing the usage of malloc and free is no different than optimizing the run time of any algorithm. The only difference is people believe they can just abstract allocation away and rely on somebody else's code to handle it.
My second question is, can you actually avoid relying on somebody else's code when you call malloc? Due to virtual memory being comprised of pages, address randomization, and other things, heap is not a continuous memory space in hardware ready to be used at the moment program runs. So, what is the cost/benefit of "individually called malloc" for each struct, as opposed to pre-allocating some space? I could be totally out of my depth, in which case ignore me ;)
malloc maintains a shared data structure that is used across cpus and must lock for access to that data structure. There are optimizations there being used with some alternative allocators such as per-cpu or per-thread pool but these only reduce the rate of locks taken.
In addition after you have that virtual memory in hand you need to have physical memory behind it which entails a page fault which switches into the kernel to get you a physical page, this happens per-page.
There is also the effect of physical memory layout on performance, if your memory is perfectly contiguous and your data is laid-out properly you can reap benefits from that compared to the quite likely fragmented nature of data in a heap-allocated setting.
This obviously means that the applications needs to manage its memory allocations on its own but when you really care about performance these become important concepts and I too tend to write my libraries for such reuse and avoid memory allocations inside them where possible.
malloc is a general purpose tool and as such needs to cater to many different environments, including multi-threaded programs so it has to lock its access to the shared data structures it holds to main
The problem is that malloc() needs to be generic. It's not about never calling malloc(), but about deciding when it is necessary for your use case vs. specialized solutions tailored to your use.
In many cases you may know that the allocated structure will always be de-allocated at the same time, for example, in which cases pools/arenas can outperform malloc() by a magnitude or more in the right circumstances, and can also use less memory.
{
"hello": "world"
}
Tny converts this JSON into a (usually, but not always) smaller form that's faster to traverse and encode/decode. Tny's alternative, BSON, is being used by MongoDB [1], so that might be a good place to learn more.Deleted comment
"Tny is a simple library to serialize data in C. It can be seen as a kind of binary JSON but unlike BSON it also supports arrays as root elements."
After a short review of the add function for handling of complex structs i found a possible bug as they miss to perform a deep copy of objects containing references to other objects. they would only copy the value of the pointers of a struct like that: struct { obj* ptr; } and not the content of ptr.