HNHacker News
TopNewBestAskShowJobs

willothy

69 karma · joined May 5, 2024

submissionscomments
willothy··on Show HN: Whirlwind – Async concurrent hashmap for Rust
Yeah, after looking into this more I think this was a big oversight on my part. Working on a (hopefully) better way of doing this right now - I'm thinking per-shard waker queues and only falling back to spinlocking like this if the queues are full.
willothy··on Show HN: Whirlwind – Async concurrent hashmap for Rust
This is an interesting idea. I am gonna try this out - especially with dashmap, I think that could perform very well.
willothy··on Show HN: Whirlwind – Async concurrent hashmap for Rust
Oh this is really good to know, thank you!
willothy··on Show HN: Whirlwind – Async concurrent hashmap for Rust
Hmm, imo it's definitely better than directly spinlocking to have many spinlocks running cooperatively, but you're right that it may not be ideal. Thanks for pointing this out. I'll see if I can find a better way to coordinate the polling/waking of lock acquisition futures.
willothy··on Show HN: Whirlwind – Async concurrent hashmap for Rust
I don't believe that waking the waker in `poll` synchronously waits / runs poll again immediately. I think it is more likely just adding the future back to the global queue to be polled. I could be wrong though, I'll look into this more. Thanks for the info!
willothy··on Show HN: Whirlwind – Async concurrent hashmap for Rust
I agree with this take a lot. I think having lots of custom implementations is inevitable for systems languages - the only reason why Rust is different is because Cargo/crates.io makes things available for everyone, where in C/C++ land you will often just DIY everything to avoid managing a bunch of dependencies by hand or with cmake/similar.
willothy··on Show HN: Whirlwind – Async concurrent hashmap for Rust
There are a few reasons - For one, I'm not sure BTreeMap is always faster in Rust... it may be sometimes but lookups are still O(log(n)) due to the searching where with a HashMap it's (mostly) O(1). They both have their uses - I usually go for BTreeMap when I explicitly need the collection to be ordered.

A second reason is sharding - sharding based on a hash is quite simple to do, but sharding an ordered collection would be quite difficult since some reads would need to search across multiple shards and thus take multiple locks.

If you mean internally (like for each shard), we're using hashbrown's raw HashTable API because it allows us to manage hashing entirely ourselves, and avoid recomputing the hash when determining the shard and looking up a key within a shard.

willothy··on Show HN: Whirlwind – Async concurrent hashmap for Rust
This is a great point too - I definitely want to run some more varied benchmarks to get a better idea of how this performs in different settings. We'll also be using it in prod soon, so we'll see how it does in a real use setting too :)
willothy··on Show HN: Whirlwind – Async concurrent hashmap for Rust
I think there's a pretty big difference between committing to semantic versioning and saying "do not use this until some unspecified point in the future." Maybe I'm just not clear enough in the note - I just mean that the API could change. But as long as a consumer doesn't use `version = "*"` in their Cargo.toml, breaking changes will always be opt-in and builds won't start failing if I release something with a big API change.
willothy··on Show HN: Whirlwind – Async concurrent hashmap for Rust
Tokio is just used for async tests and the examples, the crate doesn’t depend on any specific async runtime :)
willothy··on Show HN: Whirlwind – Async concurrent hashmap for Rust
Hey, I think the name is cool!

Fair point though, there would definitely be some benefit to having some of these things in the stdlib.

willothy··on Show HN: Whirlwind – Async concurrent hashmap for Rust
I use a multiple of `std::thread::available_paralellism()`. Tbh I borrowed the strategy from dashmap, but I tested others and this seemed to work quite well. Considering making that configurable in the future so that it can be specialized for different use-cases or used in single-threaded but cooperatively scheduled contexts.
willothy··on Show HN: Whirlwind – Async concurrent hashmap for Rust
Definitely a good point. I used dashmap's benchmark suite because it was already setup to bench many popular libraries, but I definitely want to get this tested in more varied scenarios. I'll try to add a benchmark for a single key only this week.

Regarding your edit: damn I hadn't thought of that. I'll rerun the benchmarks on my Linux desktop with a Ryzen chip and update the readme.

willothy··on Show HN: Whirlwind – Async concurrent hashmap for Rust
Yep, someone suggested loom on our Reddit r/rust post as well - I'm actively working on that. Somehow I'd just never heard of loom before this.
willothy··on Show HN: Whirlwind – Async concurrent hashmap for Rust
The blocking mainly occurs due to contention - imo most of the performance gain comes from being able to poll the lock instead of blocking until it's available when acquiring locks on shards.

In all honesty I was quite surprised by the benchmarks as well though, I wouldn't expect that much performance gain, but in high-contention scenarios it definitely makes sense.

willothy··on Show HN: Whirlwind – Async concurrent hashmap for Rust
Good point, thanks! I'll look into adding that crate to the benchmarks.
willothy··on Launch HN: Fortress (YC S24) – Database platform for multi-tenant SaaS
Those are absolutely things we're looking to support, especially in the case of monitoring / accounting. Isolated tenants don't face the noisy neighbor problem here, but tenants in shared databases still may at the moment. One possible solution we're looking at there is to move towards a serverless approach where every tenant has isolated storage and ephemeral worker VMs perform the queries.
willothy··on Launch HN: Fortress (YC S24) – Database platform for multi-tenant SaaS
We support scaling to near-zero because of Aurora serverless, but we definitely are looking into other solutions that could be cheaper to run or self-hosted.

Some regional regulations (GDPR, etc.) require local and/or isolated hosting. Most companies indeed solve this with either a tenant id column, dedicated tenant databases or both. We want to simplify those architectures, and a proxy layer is exactly our idea there - we're working on a solution that handles connection pooling and routing to remove the need to cache connections on the client.

willothy··on Launch HN: Fortress (YC S24) – Database platform for multi-tenant SaaS
Hey Will here, another Fortress cofounder.

We're still thinking about this a lot. We're using AWS Aurora currently which auto-scales compute and storage, and are looking into other options such as distributed databases (Cockroach, etc.) and Kubernetes operators.