HNHacker News
TopNewBestAskShowJobs

jerrinot

343 karma · joined October 8, 2017

submissionscomments
jerrinot··on Designing a new concurrent data structure
Hello, the author here. It feels great to see my blog on HN!

It was quite a journey, at first I thought I invented a novel concurrency schema. However, it turns out that it was simply a mix of my ignorance and hubris! :-)

Still, I had a lot of fun while designing this data structure and I believe it made a nice story. Ask me anything!

jerrinot··on Designing a new concurrent data structure
DirectByteBuffer exhibits an intriguing behavior: it deallocates its backing memory during the finalization process, which occurs after garbage collection cycles. This poses an issue if your system is conservative with on-heap allocations, leading to infrequent GC cycles. In such cases, there could be a significant delay between the time the memory becomes unreferenced and when it is actually deallocated. This behavior could, in some respects, mimic a memory leak.

This is why some libraries hacks into DirectByteBuffer to deallocate memory explicitly, bypassing the finalizer altogether. For instance, the Netty library has implemented such a workaround: https://github.com/netty/netty/blob/795db4a866401aa172757b95...

jerrinot··on Fuzz testing: the best thing to happen to our application tests
it's the usual spectrum of tests:

1. correctness: from small units tests to relatively complex integrations tests. they typically populate a test database and query it via various interfaces, such as REST or the Postgres protocol. we use Azure Pipelines to execute them - testing in MacoOS, Linux (both Intel and ARM) and Windows.

2. performance: we tend to use the TSBS project for most of our performance testing and profiling. fun fact: we actually had to patch it as the vanilla TSBS was a bottleneck in some tests. Sadly, the PR with the improvements is still not merged: https://github.com/timescale/tsbs/pull/186

edit: I thought I would link some of the more interesting tests: Since QuestDB supports the Postgres wire protocol we have to gracefully handle even various half-broken Postgres clients. But how do you write a test with mimicking a client generating invalid requests? No sane client will generate broken requests. So we use Wireshark to record network communication between the broken client and our server and then use this recorded communication in tests. Example: https://github.com/questdb/questdb/blob/3995c31210c70664d4b3...

jerrinot··on Fuzz testing: the best thing to happen to our application tests
Hello, I'm Jaromir, one of the core engineers at QuestDB team. I just noticed this blog is trending! Andrei - the author - lives in Bulgaria and he is probably already sleeping. Happy to answer any question the blog left unanswered.
jerrinot··on The Mullvad Browser
Vanilla Firefox beats it too if you set `privacy.resistFingerprinting` to `true`.

I assume Mullvad browsers has this on by default.

jerrinot··on Unstable Builds and Open Source Infrastructure
Hi, it's the author here.

I thoroughly enjoyed my troubleshooting session and thought that others might enjoy it too. How often do you debug the internals of a build tool? Therefore, I wrote this post to share my story and also to express my gratitude to all maintainers of essential software infrastructure.

jerrinot··on No, QuestDB is not Faster than ClickHouse
The article is not _just adding an index_. They are embedding one of the search fields in a table _primary key_. That likely means the whole physical table layout is tailored for that single specific query.

While it can help to win this very benchmark it's questionable whether it's usable in practice. Chances are an analytical database serves queries of various shapes. If you only need to run a single query over and over again then you might be better off with a stream processing engine anyway.

jerrinot··on Show HN: Jet – in-memory, fault-tolerant, distributed stream processing
The truth is the original Hazelcast replication protocol was not a good fit for some data-structures. We took the analysis seriously. I know every project and vendor claims that. Here is what we did in recent years:

1. Re-implemented concurrency primitives on top of Raft protocol. This includes Distributed Locks, Semaphores, AtomicLong, etc. Raft provides linearizability and that's what you usually want for concurrency primitives. See our epic blog post about locking: https://hazelcast.com/blog/long-live-distributed-locks/ or our Jepsen testing story: https://hazelcast.com/blog/testing-the-cp-subsystem-with-jep...

2. Added a FlakeID generator. This is on the opposite side of the consistency spectrum - it's a k-ordered Available (wrt CAP) ID generator. It won't generate duplicates even when there is a split-brain. See: https://docs.hazelcast.org/docs/4.0.2/manual/html-single/ind...

3. PNCounter - CRDT-based eventually consistent data structure, suitable for .. well, counting things:) See: https://en.wikipedia.org/wiki/Conflict-free_replicated_data_...

4. Significantly extended documentation, to be more explicit about Hazecast replication models and guarantees. The goal is clear: Avoid Surprises. See: https://docs.hazelcast.org/docs/4.0.2/manual/html-single/ind...

Disclaimer: Obviously I am biased as I work for Hazelcast.

jerrinot··on Sub-10 ms Latency in Java: Concurrent GC with Green Threads
I didn't know about fully concurrent root scanning, thank you!

How does root scanning work wrt to Loom? Are stacks of virtual threads treated as roots? I guess there is no other option?

jerrinot··on Sub-10 ms Latency in Java: Concurrent GC with Green Threads
Modern Garbage Collectors do concurrent scanning.

I believe most GC implementations have non-concurrent "initial marking" phase, but that's typically fairly quick. It has to scan roots of your object graph, think stack, JNI, etc.

← PreviousPage 2 of 2