HNHacker News
TopNewBestAskShowJobs

cuno

81 karma · joined September 27, 2022

submissionscomments
cuno··on I've been writing ring buffers wrong all these years (2016)
This stuff dates way, way before 2004.

For non-power of two, just checked our own very old circular byte buffer library code and using the notation from this article, it is:

  entriesAllocated() { return ((wrPtr-rdPtr+2*bufSize) % (2*bufSize)); }
  remainingSpace() { return bufSize - entriesAllocated(); }
  isEmpty() { return (entriesAllocated()==0); }
  isFull() { return (entriesAllocated()==bufSize); }
  incWr(int n) { wrPtr = (wrPtr+n) % (2*bufSize); }
  incRd(int n) { rdPtr = (rdPtr+n) % (2*bufSize); }
The 2*bufSize gives you an extra bit (beyond representing bufSize) that lets you disambiguate empty vs full. And if it is a constant power of two (e.g. via C++ template), then you can see how this just compiles into a bitmask instead, like the author's version. You read and write the buffer at (rdPtr%bufSize) and (wrPtr%bufSize) respectively.
cuno··on Can a model trained on satellite data really find brambles on the ground?
So after transforming multispectral satellite data into a 128-dimensional embedding vector you can play "Where's Wally" to pinpoint blackberry bushes? I hope they tasted good! I'm guessing you can pretty much pinpoint any other kind of thing as well then?
cuno··on C stdlib isn't threadsafe and even safe Rust didn't save us
We ended up overriding and replacing with our own thread-safe version years ago when we also hit this.
cuno··on Launch HN: Regatta Storage (YC F24) – Turn S3 into a local-like, POSIX cloud FS
Good to see you too! Lets catchup sometime - tried to connect with you separately but no luck.
cuno··on Launch HN: Regatta Storage (YC F24) – Turn S3 into a local-like, POSIX cloud FS
Founder of cunoFS here, brilliant to see lots of activity in this space, and congrats on the launch! As you'll know, there's a whole galaxy of design decisions when building file storage, and as a storage geek it's fun to see what different choices people make!

I see you've made some similar decisions to what we did for similar reasons I think - making sure files are stored 1:1 exactly as an object without some proprietary backend scrambling, offering strong consistency and POSIX semantics on the file storage, with eventual consistency between S3 and POSIX interfaces, and targeting high performance. Looks like we differ on the managed service vs traditional download and install model, and the client-first vs server-first approach (though some of our users also run cunoFS on an NFS/SMB gateway server), and caching is a paid feature for us versus an included feature for yours.

Look forward to meeting and seeing you at storage conferences!

cuno··on Dynasaur: Userspace syscalls that are super fast
Dynasaur intercepts syscalls in Linux entirely in userspace, superfast and across static, semi-static and dynamic binaries. It even works inside containers and doesn't need admin or any special privileges like PTRACE, or need a VM like GVisor. We built it for our cunoFS virtual filesystem and we want to know if others are interested in buliding on top of it for other use cases?
cuno··on The Rise and Fall of Silicon Graphics
Yes Bali was the next gen architecture and incredibly scalable. It consisted of many different chips connected together in a network that could scale. The R chip was so big existing tools couldn't handle it and ppl were writing their own tools. As a result it was very expensive to tape out so many hefty chips and I think that's why when it came time, and with a financial crisis, upper management pulled the plug.

Yes there were separate teams working on the lower-end graphics.

cuno··on The Rise and Fall of Silicon Graphics
Nintendo wasn't loyal to the company it was loyal to the team, so when they just decided to leave and form ArtX they took the customer with them... SGI was happy with the Nintendo contract. They earned $1 in additional royalties for every single N64 cartridge sold worldwide. Losing the team was a big blow.
cuno··on The Rise and Fall of Silicon Graphics
I worked at SGI on the next generation (code named Bali) in 1998 (whole year as an intern) and 1999 (part time while finishing my degree, flying back and forth from Australia). Bali was revolutionary. The goal was realtime Renderman and it really would. I had an absolute blast. I ended up designing the highspeed data paths (shader operations) for world's first floating point frame buffer (FP16 though we called it S10E5) with the logic on embedded DRAM for maximum floating point throughput. It was light years ahead of its time. But the plug got pulled just as we were taping out. Most of the team ended up at Nvidia or ArtX/ATI. The GPU industry was a small world of engineers back then. We'd have house parties with GPU engineers across all the company names you'd expect, and with beer flowing sometimes maybe a few secrets could eh spill. We had an immersive room to give visual demos and Stephen Hawking came in once pitching for a discount.

For team building, we launched potato canons into NASA Moffet field, blew up or melted Sun machines for fun with thermite and explosives. Lots of amazing people and fond memories for a kid getting started.

cuno··on S3 is files, but not a filesystem
Sorry for the late response - I didn't see your comment until now.

Our aim is to unleash all the potential that S3/Object has to offer for file system workloads. Yes, the scale of AWS S3 helps, as does erasure coding (which enhances flexibility for better load balancing of reads).

Is it suitable for every possible workload? No, which is why we have a mode called cunoFS Fusion where we let people combine a regular high-performance filesystem for IOPS, and Object for throughput, with data automatically migrated between the two according to workload behaviour. What we find is that most data/workloads need high throughput rather than high IOPS, and this tends to be the bulk of data. So rather than paying for PBs of ultra-high IOPS storage, they only need to pay for TBs of it instead. Your particular workload might well need high IOPS, but a great many workloads do not. We do have organisations doing large scale workloads on time-series (market) data using cunoFS with S3 for performance reasons.

cuno··on S3 is files, but not a filesystem
The problem with these approaches is that the data is scrambled on the backend, so you can't access the files directly from S3 anymore. Instead you need an S3 gateway to convert from scrambled S3 to unscrambled S3. They rely on a separate database to reassemble the pieces back together again.
cuno··on S3 is files, but not a filesystem
IOPS is a really lazy benchmark that we believe can greatly diverge from most real life workloads, except for truly random I/O in applications such as databases. For example, in Machine Learning, training usually consists of taking large datasets (sometimes many PBs in scale), randomly shuffling them each Epoch, and feeding them into the engine as fast as possible. Because of this, we see storage vendors for ML workloads concentrate on IOPS numbers. The GPUs however only really care about throughput. Indeed, we find a great many applications only really care about the throughput, and IOPS is only relevant if it helps to accomplish that throughput. For ML, we realised that the shuffling isn't actually random - there's no real reason for it to be random versus pseudo-random. And if its pseudo-random then it is predictable, and if its predictable then we can exploit that to great effect - yielding a 60x boost in throughput on S3, beating out a bunch of other solutions. S3 is not going to do great for truly random I/O, however, we find that most scientific, media and finance workloads are actually deterministic or semi-deterministic, and this is where cunoFS, by peering inside each process, can better predict intra-file and inter-file access patterns, so that we can hide the latencies present in S3. At the end of the day, the right benchmark is the one that reflects real world usage of applications, but that's a lot of effort to document one by one.

I agree that things like dedupe and compression can affect things, so in our large file benchmarks each file is actually random. The small file benchmarks aren't affected by "write bigger blocks" because there's nothing bigger than the file itself. Yes, data consistency can be an issue, and we've had to do all sorts of things to ensure POSIX consistency guarantees beyond what S3 (or compatible) can provide. These come with restrictions (such as on concurrent writes to the same file on multiple nodes), but so does NFS. In practice, we introduced a cunoFS Fusion mode that relies on a traditional high-IOPS filesystem for such workloads and consistency (automatically migrating data to that tier), and high throughput object for other workloads that don't need it.

cuno··on S3 is files, but not a filesystem
It depends on if you want to expose filesystem semantics or metadata to applications using it. For example random access writes are done by ffmpeg, which is a workhorse of the media industry, but most things can't handle that or are too slow. We had to build our own solution cunoFS to make it work properly at high speeds.
cuno··on S3 is files, but not a filesystem
For example, you can do separate parallel ListObjectV2 for files starting a-f and g-k, etc.. covering the whole key space. You can parallelize recursively based on what is found in the first 1000 entries so that it matches the statistics of the keys. Yes there may be pathological cases, but in practice we find this works very well.
cuno··on S3 is files, but not a filesystem
For AWS, we're comparing against filesystems in the datacenter - so EBS, EFS and FSx Lustre. Compared to these, you can see in the graphs where S3 is much faster for workloads with big files and small files: https://cuno.io/technology/

and in even more detail of different types of EBS/EFS/FSx Lustre here: https://cuno.io/blog/making-the-right-choice-comparing-the-c...

cuno··on S3 is files, but not a filesystem
Actually we've found it's often much worse than that. Code written against AWS S3 using the AWS SDK often doesn't work on a great many "S3-compatible" vendors (including on-prem versions). Although there's documentation on S3, it's vague in many ways, and the AWS SDKs rely on actual AWS behaviour. We've had to deal with a lot of commercial and cloud vendors that subtly break things. This includes giant public cloud companies. In one case a giant vendor only failed at high loads, making it appear to "work" until it didn't, because its backoff response was not what the AWS SDK expected. It's been a headache that we've had to deal for cunoFS, as well as making it work with GCP and Azure. At the big HPC conference Supercomputing 2023, when we mentioned supporting "S3 compatible" systems, we would often be told stories about applications not working with their supposedly "S3 compatible" one (from a mix of vendors).
cuno··on S3 is files, but not a filesystem
We and our customers use S3 as a POSIX filesystem, and we generally find it faster than a local filesystem for many benchmarks. For listing directories we find it faster than Lustre (a real high performance filesystem). Our approach is to first try listing directories with a single ListObjectV2 (which on AWS S3 is in lexicographic order) and if it hasn't made much progress, we start listing with parallel ListObjectV2. Once you start parallelising the ListObjectV2 (rather than sequentially "continuing") you get massive speedups.
cuno··on Show HN: cunoFS – mount S3/AZ/GCP storage *without* FUSE, 60x faster than s3fs
We've spent a lot of time identifying bottlenecks and fixing them, up and down the stack, with FUSE being just one of them, even the AWS SDK itself introduces its own set that we've addressed. cunoFS can also be used with FUSE, and we find that it is roughly half the speed of non-FUSE (but thanks to our other optimisations, this is still much faster than alternatives).
cuno··on The miracle of modern chip manufacturing (visual)
This is a nice simple visual guide to chip fabrication, chiplets, and the advantage of back-side power.
cuno··on The Myth of RAM (2014)
Funnily enough I published something similar as part of my PhD (2010). Essentially communication costs dominate over computation costs - an addition operation is practically free compared to the time and energy of moving data within a chip to the ALU to do the addition, let alone between chips. This is now the reverse situation to historical VLSI - whereby the computation was slow and expensive compared to practically "fast and free" on-chip communication. The implications extend far beyond just RAM. But yes, depending on the access patterns involved and the (physical) spatial arrangement of data, binary tree traversal (on 2D CMP or 2D cross-chip layout) is O(sqrt(N)) or O(log(N)sqrt(N)), and this result is not dependent on the system boundaries of memory hierarchies - it is true for dedicated on-chip scratchpad memories as it is for giant wafer-scale chips (such as Cerebras) as well. There are some special cases where it can be O(1) or O(log(N)) on average that I won't go into, but you can read more here if interested (sections 2.1 and 8.3):

https://www.cl.cam.ac.uk/~swm11/research/greenfield.html

A key way to think about things is that algorithms are not really running in some Platonic realm, each executed instruction occurs at some physical location and time, and data needs to move from one place and time to another place and time. To this end both physical wiring (or networks) are required to move the data spatially, and memory is used to move the data temporally. Together an algorithm's executed instructions are situated somewhere spatio-temporally and RAM (both on-chip and off-chip) serves as a temporal interconnect, but one that itself takes up physical space.

cuno··on Is POSIX really outdated?
Hi sorry I missed this. Yes, you can run Samba for SMB, and Ganesha for NFSv4. Feel free to email us if you think there's a better way though.
cuno··on Is POSIX really outdated?
Hi author here, sorry I missed this post. The performance benchmarks and cost comparisons are for comparing S3 vs EBS (ext4 formatted), EFS, FSx Lustre and others within the same datacenter (i.e. LAN use case rather than WAN use case). That means if you have an EC2 instance running in, say AWS Ohio, and are comparing those storage options also within AWS Ohio, then cunoFS is both cheaper and higher throughput than those other options. It's a different story over WAN. In that case, your own local NVMe storage is going to be cheaper and generally faster that remote storage over a WAN. But that local NVMe storage (on say your solo laptop) isn't going to have anywhere near the Enterprise-grade redundancy, availability and scalability that AWS S3/Azure Blob/Storj/Wasabi/etc has.
cuno··on Is POSIX really outdated?
Hi, author here.

Yes agree that you can't "just" put a POSIX API on S3, but that doesn't make it impossible. For the sake of keeping the article to a reasonable size, I left a lot of things out. There are tradeoffs that occur between POSIX semantics, consistency and performance. Each application/process has different needs but the great thing about running right inside the process is that we can see what those needs are and adapt. For example, many applications have no need for random access writes - the only libc calls and syscalls exposed are purely sequential. Some processes have both random access writes and POSIX record locks around them to protect them from other concurrent processes - and we can see that. That means we treat these applications/files differently, with some corresponding performance implications. This is very different to a normal filesystem that has to treat every process the same way because it is a "black box".

You're right that AFS, and for that matter NFS, can in principle return error on close which many existing applications unfortunately aren't written to handle. However, that doesn't mean that NFS isn't practical - it is very widely used.

Our customers mostly run workloads in the same region as the object storage (whether in cloud or on-prem) typically with very high availability. As an essentially networked file system, you're right that it can't make much stronger guarantees than the NFS protocol itself does, but operating inside cloud infrastructure you typically see 4 9s availability.

cuno··on Is POSIX really outdated?
Depends on what you mean by metadata. MP4 metadata is data inside the file - and is modified by server-side-copy semantics that replaces only the bits that are changed. If you mean POSIX metadata, we avoid storing that in the object, and for performance store that elsewhere (it's encoded and compressed in the actual filename of hidden files).
cuno··on Is POSIX really outdated?
Hi, author here.

Funny you should give the stream-encoding MP4 example, because yes that is what people's experience has been for S3. We've solved that - no temporary local file needed - all streamed directly to S3, for example using ordinary ffmpeg. The trick is a deeper understanding of how multi-part upload works, and if necessary, server-side copy semantics on those parts.

cuno··on Is POSIX really outdated?
Hi, author here.

Yes we have many large companies (Fortune Global 500) down to small organisations using our software with this kind of interception (see for example https://cuno.io/about-us/). It took us a decent sized team a lot of years to get right, because it is so very hard a problem to crack. But we think it is worth it. And for those who don't want to use such interception, we do offer a FUSE layer as well that still offers much higher performance than alternatives.

cuno··on Is POSIX really outdated?
Hi, author here.

Yes, we've had to do a lot of things to deal with object storage latency. Since cunoFS is running inside a given process itself, it has greater visibility into what actions it can take and is likely to take. This means we can make huge improvements in our prediction logic, so that we can prefetch within and across files much better, thus hiding latencies. For POSIX metadata, we have a caching mechanism which is shared between processes, and we've invented a better way to encode this metadata so that it is retrieved alongside the LIST operation.

cuno··on Is POSIX really outdated?
Hi, similar to nvm0n2 who already commented below, object storage systems often directly write to block storage rather than through a VFS. While minio, for example, can run on top of a VFS, even they say you shouldn't be - you should instead run it directly on block storage across multiple drives / nodes.
cuno··on Is POSIX really outdated?
Hi, author here.

The problem is that it's actually really hard to write you own specialized glue layer,and so most developers who do that have poor implementations with low performance and/or incompatibilities with non-AWS solutions. In the context of object storage, we've found lots of applications have tried to add S3 support, and while they work with AWS S3 and have basic functionality, they fail on a lot of other S3-compatible solutions. So they end up tied to AWS, and you can't use them on Microsoft Azure Storage, or often on Google Cloud Storage (despite its S3 gateway), or others.

For instance, a key workhorse of the genomics field is `samtools`, which works with AWS S3 in some ways, but not others (like Amazon Resource Names[1]). Our approach works across vendors transparently, and on S3, is much faster compared to such native implementations.

[1]: https://docs.aws.amazon.com/IAM/latest/UserGuide/reference-a...

cuno··on Is POSIX really outdated?
Hi, author here.

Yes, we're a startup. With cunoFS, we've made it possible to leverage the lower cost and higher throughput of object storage like S3, and make it work transparently with functions like e.g. mmap(), chmod(), execve(), renameat2().

Page 1 of 2Next →