Apache Helix – Near-Realtime Rsync Replicated File System
helix.apache.org
helix.apache.org
Edit: it seems to invoke an a rsync process but doesn't parse stdout for progress. A bit disappointing but I suppose that's all which is needed for this application.
Here's an example of library that wraps rsync https://metacpan.org/pod/File::Rsync - you can provide callback functions that will be called for each line of stdout/stderr.
Granted it's in Perl - and chances are your code isn't written in Perl. So if you can't find something similar for your language of choice, and if you can't be bothered with checking out implementation in Perl and rewriting it...
I would suggest searching & checking out "Inline::" namespace (e.g.: https://metacpan.org/pod/distribution/Inline-Java/lib/Inline..., and there's also JS, C ...etc).
And next to those - Perlito can "Compile Perl to Java/JavaScript/Python/Ruby/Go/etc" https://github.com/fglock/Perlito. I think we had/have some Perl code running in production Hadoop/JVM.
The link here is to the docs of version 0.6.8 while 1.0.1 is current (after 0.7, 0.8 and 0.9).
This also explains the heavy weight architecture, you probably are running a ZooKeeper anyway in that environment.
If a file gets renamed or copied, modified or not, rsync will transfer the whole file.
They added --fuzzy to try improve on this situation, but it generally only helps with unmodified files copied/renamed within the same directory.
You can build robust solutions around rsync that work efficiently most of the time, but pathological cases such as users having a huge file regularly regenerated using something like a $(date --iso-8601=seconds) filename (think mysqldump backups) are surprisingly common once you have enough users. Even if you can convince them to uncompress such files for delta transfer sake, something as common as a versioned filename prevents rsync from automatically finding the relationship for use in the delta transfer.
I believe you can use journalling to detect file name change events but I don't think it scales well
For a filesystem the developers are free to innovate on any form of filesystem-wide block checksumming and content-addressed deduplication approach. This kind of global block-level content awareness is entirely decoupled from filenames and paths.
Using rsync to underpin a filesystem strikes me as a gross hack done by developers punting on the real challenges of modern filesystem development.
WRT rsync doing better than --fuzzy, the options are more limited since it's a userspace tool working within the confines of filesystem APIs like POSIX.
Fyi have a look at zfs. It has some neat traits/behaviours. So my thinking is, someone should write a database that uses gazzilions of files (instead of a file per table), a tiny file for each cell/field in a table. Then use zfs for versioning of data (or git). Zfs can already do software mirroring to drive pools on the machine, so maybe add rsync for remote pools. Optimize it for SSD's and you have a filesystem-database-replicating-versioned -based store.
Maybe just need a nice cli/gui for zfs to make it easier/safer to expose more of it's interesting features.
That sounds a bit like IPFS
If these things were written in C, Go or Rust then I'd leap at the opportunity to explore them and maybe use them in production.
But Java dependency just brings so much baggage with it. Its all well and good if you've got a Java pro or two on your team. But if you don't, you usually end up spending hours and days troubleshooting obscure JVM problems or somesuch.
(Note that there are a lot of very high quality Apache java projects.)
I know the ASF isn't as trendy these days, but that's never really bothered those of us who volunteer. The focus has always been about a stable, open source ecosystem, primarily, but not exclusively for the internet and its underlying infrastructure. The goal is measured in decades.
It wasn't about being "enterprisey" or anything like that. The ASF as a foundation doesn't really care about the language or tech stack. Being a more mature non-profit, many corporations have found it easier to interface with the ASF than many other open source organizations out there, which does lend itself to seeing a lot of enterprisey donations.
All the worst and most time-consuming troubleshooting experiences of my life have involved Java. Whether staring at pages of obscure stacktraces that the JVM decided to vomit all over my screen, or hours on end with Oracle tech support troubleshooting why Java decided to throw a tantrum (followed by editing obscure Java XML config files).
These days I just stay well away from Java.
It's the app whose stack traces you are looking at. And the app that decided to use XML config files.
I really haven't seen all that much XML in Java recently outside of legacy applications... there's libraries for loading config data from JSON, YAML, etc etc.
Google's scale means that performance actually matters a lot versus productivity. When you've got millions of servers the costs add up. So making something twice as fast but taking 25% more dev time is probably a good tradeoff for them. For most other companies that's not the case.
Have you ever heard the phrase "the dog that didn't bark?" You appear to be using an absence of evidence to prove evidence of absence, not to mention an argument from ignorance (which is always possible in our post-Gödel world).
> java is very easy to troubleshoot compared to c, go, and rust
>> Can you expand on this? Why do you think that?
So the comparison is C, Go, and Rust. Yes, I'm familiar with all 3 of them and Java is easier to troubleshoot than either of them because it has better introspection out of the box.
For example, also having experience with all of them, the overall advantage over C is clear-cut but Rust is a lot less clear since stronger type system and better culture around package management and complexity avoids a lot of problems that otherwise require runtime troubleshooting.
Granted, a key part of that is really the question of whether you’re talking about a modern Java project or the more common enterprise Java sort with layers of accreted complexity and probably architecturally frozen at Java 8, where the problem is cultural rather than the language itself.
Consider that you've gotten more useful information from my response than what you've paid for. If you want the whole picture, buy the book. Your sense of entitlement to more of my time seems misplaced.
All the above can be done in other environments, but generally with specially prepared deployments or executables, with choices being made along the way as to how to do it. I think the two key things are that in the JVM the heap is actually a quite structured database, which allows for introspection without recompilation or special tooling, and there's a standard mechanism for exposing detailed performance data.
Whether or not that translates into actual productivity differences is tricky. In my experience large companies tend to build their own equivalent and better-targeted tooling, and smaller companies increasingly pass their performance diagnostics to SAAS companies. It takes time to learn diagnostic tools and procedures in general, and in my experience a lot of Java teams don't know the tooling they have. I'd say the main productivity gain would be in quickly diagnosing production issues. You could argue that the JVM leads to software thinking that increases production issues by relying on long-running processes that need the stability to survive, but in my experience there are definitely niches where that is the only way to meet your performance goals.
I realize, writing this, that for someone unfamiliar with the ecosystem this sounds very abstract, but MBeans (managed beans) are simply counters, operations, stats, faults that the app makes available to other apps/users via a standard protocol.
Most new projects are related to existing ASF projects, or use ASF libraries, or are created by someone already in the ASF that maybe use/prefer Java.
But there is no requirement on a project being Java.
See Airflow, Arrow, Log4net, httpd... it's just a question of people creating the proposals that must show a possible community of users/devs to use/support it.
Not saying that either is good, but it is just fact of life.
Languages like C and Rust do not have any equivalent to JVM tuning.
In Java's case if you have all the class files, all you need is the JVM for your architecture. You don't need the program's author to compile it for your architecture.
And finally, JVM tuning does not need to be performed by the user, but can be performed by the user. The most frequent tuning needed to be done by the user was the maximum memory that the JVM could allocate, but that hasn't been needed for a long time (since Java 8, I think). Now the JVM, by default, has the sensible behavior to allocate as much memory as the OS will let it, when asked by the application.
My experience is not that you can run 99% of Java applications without running out of memory or something like that. That's why I avoid Java these days when I have the option. I don't think things would have changed much that much in recent years, making my experience no longer valid.
I've been using Java daily for years at this point. The only time I've had that sort of issue is when the program I was writing dealt with large files in memory, and it was solved with a simple -Xmx16G.
Admittedly even with a C program the size of a problem requiring 16GB might not run well in a 2GB machine, but the fact remains that Java applications are often much more memory hungry than comparable C/C++ implementations. And if you have the memory they will just use it without any Xmx. That sounds really like computing in the 1960/70s that you need to allocate memory in advance. (I'm old enough to have seen such stuff, although not when it was new.)
Is there a concept that the search machine would find, which has brought the radical change you mention?
Regarding OpenJDK, other JVM vendors have improvements of their own.
G1 has had ~700 improvements since JDK 8 and is now very different: https://archive.fosdem.org/2020/schedule/event/g1/
ZGC, the low-latency collector (which gives a ~1-2ms max pause times on heaps of up to 16TB in JDK 16) was introduced in 11 and made production-ready in 15.
In general, the way the GCs work is just very different from 8.
Plus, there have been lots of improvements to CDS and startup time in general due to how the VM loads classes and initialises data (~30 ms to run Hello, World on a cold VM in 15) -- https://cl4es.github.io/2019/11/20/OpenJDK-Startup-Update.ht..., https://twitter.com/cl4es/status/1311335253139771393, https://www.morling.dev/blog/building-class-data-sharing-arc...
And, of course, now everything lies on top of the module system.
I have run operations for programs written in Java, C++, and Go. I've done it both at small companies and large companies, for everywhere from a rack of dedicated servers running Apache Tomcat to thousands of virtual machines running a mixture of C++ and Go microservices.
The constant thread is... over the last ten years... the people running JVMs complain about tuning the JVM, and always have stories to tell about it. The JVM is fantastically tuneable, and when people online complain about GC pauses or memory problems, there's often some way to tune the JVM to fix those problems. The JVM is amazing, it's a marvelous piece of technology.
But there's also a bunch of people who don't know how to do it and just turn knobs. Like, oh, customers are calling in to complain about latency, so I'll increase the size of the Java heap. (Which, for those not familiar with GC, will make throughput better but make latency worse.) Running services on the JVM means that "working with the JVM" is now a skill you need to select for whatever operations team you have, and if you're a company running a mix of different services (like, I don't know, some databases, memcache, load balancers, etc) then "knowing how to tune the JVM" competes with several other skills you'd like your team to learn.
Or those that complain about RDBMS queries being slow, without having normalized their data or written proper indexes.
There is always bunch of people that don't know how to do things, and then there are those among them, that care to improve their skill set and get to know how to turn those knobs.
It’s not a question of whether there are knobs to turn. It’s about the typical experience of an actual person running a JVM app versus, say, C or Go, and those experiences are quite different.
And we’re also talking about running someone else’s app. If your database queries are slow, that’s a conversation between the DBA and the devs. If you’re running some Apache app, you’re probably not talking to the devs.
This hasn't been my experience at all after running java on production servers for over a decade.
That said you probably can’t get very far in c/c++ without worrying about memory the whole way through. So, you only worry about tuning when writing your code.
Most of the time defaults just work.
(You can read a disk 10 times in series or 100-1000 times in parallel in 1ms; it’s practically an eternity.)
I found the comments really clueful but also appreciate when others like e.g. @_msw_ are very overt with disclosure of their employment and interests whilst talking about $employer's stuff on HN (and in his case that'd be AWS and EC2).
Edit: also, did they replace the JIT with AOT?
No, OpenJDK 15 defaults to G1 (since JDK 9), but it is very different from G1 in JDK 8 and rarely requires tuning: https://archive.fosdem.org/2020/schedule/event/g1/
Plus, ZGC, the low-latency collector (which gives a ~1-2ms max pause times on heaps of up to 16TB in JDK 16) was introduced in 11 and made production-ready in 15.
> Edit: also, did they replace the JIT with AOT?
Replace? Why would anyone want to do that? But there is an AOT compiler available: https://www.graalvm.org/reference-manual/native-image/
They absolutely do. It’s called “managing memory.”
You're likely talking about work on custom memory allocators and impact on real time performance. In the hosted services domain this isn't so much an issue. Nobody's doing that in Java, and you likewise wouldn't care much if writing the same type of service in Rust.
Rust very much enables you to write high level, Java-style code. It's not bad code, nor is it difficult to write.
In my opinion it’s the worst of both worlds at that point—you’ve got to manually manage memory and tip toe around the runtime.
You really don't. Anything you could do it, say, Go, you can do in Java without doing any JVM flags. The only time you need to do JVM tuning is when you have performance requirements that are simply impossible in most languages.
Does't this argument cut both ways?
i can write my C programs without ever thinking about java. it is irrelevant to me. however, C is very important for java since the JVM it is running on top of is written in C.
i think, personally, the outcome this type of thinking is that “therefore every java program is in fact just cruft on top of C” which personally as someone who does 80% of their job in SQL i am ill-equipped to object on java’s behalf.
https://en.wikipedia.org/wiki/List_of_Java_virtual_machines
Most of them don't have a single line of C, rather C++.
Then there are a couple of them like JikesRVM and GraalVM that are meta-circular implementations.
I’ve deployed things written in many languages at extremely large scale. They’ve been written in Java, C, C++, python, go, sh, perl, and other languages I’ve forgotten. Java is, by far, the least operable of those languages.
I’m an expert Java developer, along with most of those other languages (where expert is defined by me as “over 10 years experience”).
In fairness to java, it’s not my least favorite language on the list.
Most technology is useful, depending on what you're optimizing for. I urge you not to optimize for being comfortable.
See sibling comment: https://news.ycombinator.com/item?id=24899595
All the worst and most time-consuming troubleshooting experiences of my life have involved dealing with a legacy VB6 app with a truly terrible MSSQL database structure - that's not an emotional statement, it's a fact.
Do you feel any powerful emotion when you think about that experience? Would you be indifferent if you would need to do it again?
I'm not doubting that at all. Was it VB6's fault or MSSQL's?
I was just trying to make the point that it's not necessarily emotional for someone to not want to work with a certain tool without an expert available because of terrible past experiences where that was the case.
I'd say it's almost impossible to imagine any past event someone has present at without having some emotional response.
Having spent much of the first half of my career writing C and Perl and then Java, I was very happy to set aside the latter two - for completely differing reasons.
Our shop has about the same Cassandra DBAs, as JVM experts to keep it running.
Every time a node goes down because of JVM shenanigans the ScyllaDB migration discussion pops back.
You can just treat these appliances as a black box if you can't be bothered to learn the tooling.
As soon as I see Java I know it's going to have a whole set of properties that will make it highly manageable to deploy and run. And when they go wrong I have a lot of hooks and tools to understand and resolve the issue that are completely unavailable to me with native applications. You apply the pejorative "baggage" to these as if they are valueless but sometimes you just have to develop enough experience before the utility of things becomes apparent.
(Writing an application in Java however ... I would go straight for a JVM language like Kotlin/Groovy/Scala)
Ironically it pushed me back towards Groovy because I found I was constantly hitting unexpected performance bottlenecks by implicit conversions that unexpectedly jumped in and executed things I didn't even realise were happening. Groovy was completely unelegant but pragmatic option that actually did what I expected most of the time.
If it was years ago, then I can say that things have changed a lot since the last time you touched Scala, and in the right direction. Most of the criticisms from, say, 5 years ago, are being addressed.
For example, if by "operator overloading" you mean "symbolic method names", then those have largely fallen out of fashion. sbt has seen huge improvements (some people dislike it, some people like it, but it's one of the most powerful build tools out there). Compilation times also have improved.
As far as I am concerned, it is my favorite language, and there is so much going for it: great tooling, the upcoming 3 version, binary compatibility improvements, rock-solid JavaScript compilation, native compilation under development, and, of course, all the benefits from a language with one of the most powerful types systems around.
> If these things were written in C, Go or Rust...
Do these languages have similar ecosystems?
https://github.com/draffensperger/go-interlang
> Rust libraries can only be used from Rust
https://rust-embedded.github.io/book/interoperability/rust-w...
And any language that supports an FFI interface can call rust code. For example you can write a rust lib and call it from PHP using it's recently released FFI https://www.php.net/manual/en/book.ffi.php, e.g. https://dev.to/verkkokauppacom/introduction-to-php-ffi-po3
I love using things written in Rust, because if nothing else I know it's going to be resource efficient and the deployment will be nice and easy.
Java applications: I hope there's enough spare memory, I hope the setup instructions aren't littered with "set this obscure JVM config in some obscure way, that isn't the application config, but we won't tell you how/where, you've just got to figure that one out"
I currently use Gluster which has its downsides, but is what I consider near-realtime, but with full filesystem features in a nice fuse wrapper. I'd love it if someone would contrast the performance and more of the features.
I asked him if we would get "real time results?"
He stopped, and looked at me puzzled with a tilted head... "What other kind of time is there?"
Your question is highly reasonable, but I heard a colleague today say that the issue could happen only in a very, very short time period of, say, 2-3 minutes. We routinely work with corner cases of 1-2 ms; so the context is relative :-)
It's also thrashing a bit development wise; there's a lot of features being pruned, and I'm a little concerned what its future looks like as Red Hat push harder down the Ceph route, since they're the main source of code contributions to Gluster.
Curious to know more about your production use case?
Maybe this could be an alternative for those!
The approach described in this page does make some attempt to provide consistency. The master generates a single stream of file updates, and all of the replicas consistently apply those changes in the same order. But rsync is not guaranteed to observe changes to files on the master in the same order that they were written.
In particular, SQLite (by default) uses a rollback journal to recover from failures. While a transaction is in process and some parts of the database might have been partially updated, the journal stores the old contents of the modified pages. Since the old data is fsync'ed to the journal before the new data is written to the database file, the data on disk at any moment in time is recoverable.
But there's no clean way for rsync to read an atomic snapshot of both the database file and the journal at the same instant in time. If it reads them at different times, a database change might be partially applied or partially rolled back, which has a high likelihood of corrupting the database.
The property you are looking for in a distributed filesystem is called strong consistency, or cache coherent (different name, same effect).
If a distributed filesystem does genuinely provide that property, it should be ok for SQLite and other applications, because it means the filesystem behaves the same as if it were multiple processes on the same local system accessing it.
Most distributed filesystems don't provide that guarantee though. It's a very nice guarantee, but it comes at a complexity cost, usually a performance cost (although there are clever ways to approach the perforance of a non-consistent system), and in particular it means the filesystems can't be used when the network connection is down.
There is another property called durability, which you might also care about. That affects whether you get corruption when devices fail or are rebooted suddenly.
It goes over my head a bit exactly why (related to POSIX lock mechanisms?) but sqlite doesn't work reliable on NFS/ceph/gluster/networked filesystems, even when only accessed from a single host, distributed or not.
> https://sqlite.org/forum/forumpost/4b340b81eb
I don't see anything in that forum thread which says SQLite is unreliable on sufficiently POSIX-ish network filesystems. Single host or not.
The linked post says it will be slower to commit write transactions than is possible using other methods, but I don't see anything there saying it's unreliable.
(What looks unreliable to me is Joelmo's proposal to remove the WAL lock for better performance. But the lock is there for a reliability reason, you can't just remove it. To get the boost in performance Joelmo would like requires significantly different techniques.)
Which distributed file systems do provide this guarantee? I can't think of one, so hopeful to learn something new today.
I actually don't know of any distributed (multi-master, fully replicated) filesystems that are published with this property. I only know it's possible because I designed one that isn't published.
The basic principle is similar to CPU MESI caching but with predicate-scopes suited to a filesystem rather than cache lines. (Predicate-scopes are similar to predicate-locks in databases). As you know, multi-core CPU systems remain fast despite sharing memory, incurring significant overhead only for the changes that need to be communicated between cores. Same applies on a network.
So, what happens after 32bit transactions without re-election?
For purely server-side "cluster" sync I've found csync2 to be a reasonable solution.
You'll adopt the oft-repeated "RAID isn't a backup" mantra as soon as any of these issues hit you (like several have hit me): https://photostructure.com/faq/raid-is-not-a-backup/#why-isn...