How to Use the Foreign Function API in Java 22 to Call C Libraries
ifesunmola.com
ifesunmola.com
I have bundled shared libraries for five or six platforms in a java library that needs to make syscalls. It works but it is a pain if anything ever changes or a new platform needs to be brought up. Checking in binaries always feels icky but is necessary if not all targets can be built on a single machine.
The problem with the new api is that people upgrade java very slowly in most contexts. For an oss library developer, I see very little value add in this feature because I'm still stuck for all of my users who are using an older version of java. If I integrate the new ffi api, now I have to support both it and the jni api.
There are several, including SWIG.
All the good bindings I use are generated by custom systems (ie, usually some Python scripts) tailored for the specific way their library works.
> It works but it is a pain if anything ever changes or a new platform needs to be brought up. Checking in binaries always feels icky but is necessary if not all targets can be built on a single machine.
It is definitely a pain when you cannot test all changes on a single local machine. But I would argue that it is true whenever multiple platforms (or maybe even multiple glibc versions) are involved, regardless of what languages/libraries/tools you use.
[1] https://libgdx.com/wiki/utils/jnigen [2] https://github.com/libgdx/libgdx [3] https://github.com/gudzpoz/luajava
Every ffi example I've found seem to operate on the assumption that you want to invoke syscalls or libc, which (with possibly the exception of like madvise and aioring) Java already mostly has decent facilities to interact with even without native calls.
I was on the assumption that it was dynamically linking the libarary with the OS dynamic linker, which in no OS I'm aware of is capable of loading libraries inside of zip files.
Not sure where I got that notion. Maybe I was overthinking this.
I have some extremely unwieldy off-heap operations currently implemented in Java (like quicksort for 128 bit records) that would be very nice to offload as FFI calls to the corresponding a single-line C++ function.
Stay in the swamp :)
Except if we're talking about some college student or hobbyist picking their first language and exploring the language space...
But even if it was needed, such records can be commonly represented by the same structs both at C#'s and C++'s sides without overhead.
An array of such could be passed as is as a pointer, or vice versa - a buffer of struts allocated in C/C++ can be wrapped in a Span<Record128> and transparently interact with the rest of standard library without having to touch unsafe (aside from eventually freeing it, should that be necessary).
You can just load as a resource. We do this internally since much of network stack is C. But we use JNI because code is older than Java 22.
Stackoverflow is full of "copy it into a temp file" solutions. ChatGPT keeps saying "sorry" but still insists on copying it into a temp file :)
[0] - https://docs.oracle.com/en%2Fjava%2Fjavase%2F22%2Fdocs%2Fapi...
new FileOutputStream(tmpFile)
Apologies.Decompiled class file:
try {
var4 = File.createTempFile("libzstd-jni-1.5.0-4", "." + libExtension(), var0);
var4.deleteOnExit();Though in the example given, I do see your point now. You'd have to make sure the DLL was unloaded before the delete-on-exit happened.
Runtime.getRuntime().addShutdownHook(...)1: https://docs.oracle.com/en/java/javase/22/docs/specs/jni/inv...
https://github.com/java-native-access/jna/blob/40f0a1249b5ad...
Do NOT force the class loader to unload the native library, since
that introduces issues with cleaning up any extant JNA bits
(e.g. Memory) which may still need use of the library before shutdown.
Following the blame back to 2011, they did unload DLLs before https://github.com/java-native-access/jna/commit/71de662675b... Remove any automatically unpacked native library. Forcing the class
loader to unload it first is only required on Windows, since the
temporary native library is still "in use" and can't be deleted until
the native library is removed from its class loader. Any deferred
execution we might install at this point would prevent the Native
class and its class loader from being GC'd, so we instead force
the native library unload just a little bit prematurely.
Users reported occasional access violation errors during shutdown.Without it 1.6-2MiB-sized AOT binaries would not have been possible (most space is occupied by standard library/runtime bits and GC)
[1]: Technically, this is the only distribution model because all Java runtimes as of JDK 9 are created with jlink, including the runtime included in the JDK (which many people use as-is), but I mean a custom runtime packaged with the application.
Java libraries are still obtained from Maven repositories via Maven/Gradle/Ant/Bazel/etc.
For example, each these jars named "native-$os-$arch.jar" contain a .dll/.so/.dylib: https://repo1.maven.org/maven2/com/aayushatharva/brotli4j/
JNA will extract the appropriate native library (using os.name and os.arch system properties), save the library to a temp file, then load it.
> JNA will extract the appropriate native library ..., save the library to a temp file, then load it.
JNA does this?FYI: JNA = Java Native Access project: https://github.com/java-native-access/jna
Because you would use ffi to interact with libraries that don't have Java wrappers yet: IE, you're writing the wrapper.
Using syscalls or libc is a way to write an example against a known library that you're probably familiar with.
I’ve met many perfectly reasonable developers who do know all those steps can be done but can’t put them all together - maybe because it just hasn’t clicked that you can store a library in a jar. It feels like something tutorials should cover, but I think falls into the, “surely everyone can work it out?” category.
i wonder if there's a way to do this entirely in memory? Because some deployment scenarios might not have disk space at all.
Edit: glibc proposal which was never accepted: https://sourceware.org/bugzilla/show_bug.cgi?id=11767
actually you delete it immediately (after load) on anything that's not windows... even then but it's likely to return false.
deleteOnExit just stores the path to delete and uses a shutdownHook to actually call delete. Nothing really special about it
- Find all the shared libraries in your JARs or configured app inputs (files in your build/source tree)
- Sniff them to figure out what OS and CPU arch they are for
- Bundle them into the right package for each platform that it makes, in the right place to be found by System.loadLibrary()
- Sign them if necessary
- Delete them from the JARs now they are extracted. Optionally extract them from library JARs, sign them and then put them back if your library refuses to load the shared library from disk instead of unpacking it (most libs don't need this)
- JLink a bundled JVM for your app for each platform you target, using jdeps to figure out the right set of modules, and combine that with your shared libs.
When building Debian/Ubuntu packages it will also:
- Read the .so library dependencies, look up the packages that contain those other shared libraries and add package dependencies on those packages, so "apt install" will do the right thing.
So that makes it a lot easier to distribute Java apps that use native code.
In C# I can just do something like this conceptual code:
```
// FILE *fopen(const char *filename, const char *mode)
[DllImport("libc")] public unsafe extern nint fopen([MarshalAs(UnmanagedType.LPStr)] string filename, [MarshalAs(UnmanagedType.LPStr)] string mode);
// char *fgets(char *str, int n, FILE *stream)
[DllImport("libc")] public unsafe extern nint fgets([MarshalAs(UnmanagedType.LPStr)] string str, int n, nint stream);
// int fclose(FILE *stream)
[DllImport("libc")] public unsafe extern int fclose(nint stream);
```
So much less code, and so much more precise than any of the Java JNI and FFI stuff.
var text = "Hello, World!"u8;
write(1, text, text.Length);
[DllImport("libc")]
static extern nint write(nint fd, ReadOnlySpan<byte> buf, nint count);
(note: it's recommended to use [LibraryImport] instead for p/invoke declarations that require marshalling as it does not require JIT/runtime marshalling but just generates (better) p/invoke stub at build time)I'm sure someone will come along and write annotations to do exactly as you describe there. The Java language folks tend to be very conservative about putting stuff in the official API, cuz they know it'll have to stay there for 30+ years. They prefer to let the community write something like annotations over low-level APIs.
Anyway, the GraalVM folks don't have quite the same limitations as Java, so they have annotations already (https://yyhh.org/blog/2021/02/writing-c-code-in-javaclojure-...):
@CStruct("MDB_val")
public interface MDB_val extends PointerBase {
@CField("mv_size")
long get_mv_size();
@CField("mv_size")
void set_mv_size(long value);
@CField("mv_data")
VoidPointer get_mv_data();
@CField("mv_data")
void set_mv_data(VoidPointer value);
}But if you want to use SDL2 from something higher-level, you will be much better served by C# which will give you minimal FFI cost and most data structures you want to express in C as-is.
When I played with this new java api. I wasn't worried about the FFI cost. It seemed fast enough to me. My toy application was performing about 0.77x of pure C equivalent. I think Java's memory model and heavy heap use might hurt more. Hopefully Java will catch up when it gets value objects with Project Valhalla. Next decade or so :)
In improbable case you may want to try it out, then all it needs is
- SDK from https://dot.net/download (or package manager of your choice if you are on Linux e.g. `sudo apt-get install dotnet-sdk-8.0`, !do not! use Homebrew if you are on macOS however, use .pkg installer)
- C# extension for VS Code (DevKit is not needed)
- SDL2 abstraction: https://github.com/dotnet/Silk.NET (there are all sorts of alternate bindings depending on your preferences)
And once you have enough momentum, switching isn’t usually worth it.
(As someone who has done Perl, C, Java, C#, Kotlin, JS, and Python professionally - god help me. Maybe a million lines of code all in now?)
Why I chose Java boils down to two reasons:
- runs on linux (I know there is some version of c# that eventually opened up, but I kind of expect it to have lot of conditions for being cross platform, I assume that standard c# code is not crossplatform due to some reason (e.g. Com usage might be standard way of doing stuff), which would make finding crossplatform answers tedious)
- whole ecosystem is more open source and more involved parties (which I interpreted as abit less controlled by the corporate overlord, so if corporate overlord went rogue, greater chance that language would survive somehow)
Never needed to call into C though..
Neither point was ever true in the last ~10 years when it comes to gamedev (or where you want to use SDL) where Java was and continues to be a much weaker choice.
Since the open .NET is pretty young, and they still have trouble with community perception due to their past actions, finding high quality FOSS libraries may pose a problem depending on what you're doing. Pretty much everything from MS is open and high quality, but they don't provide everything under the sun.
And with Java you always have alternative runtimes in case this Oracle deal goes sideways for any reason.
So you're all good, don't worry about it.
GObject (GTK4 and similar): https://github.com/gircore/gir.core (significantly better and faster than Java alternatives, this is just one example among many)
Young: first OSS version was released 8 years ago
Solaris: might as well say "it runs COBOL but not .NET"
It's funny that everyone missed the initial context of the question and jumped onto parroting the same arguments as years ago, without actually addressing the matter at hand or saying anything of substance. Unsurprising show of ignorance by Java community. Please stay this way - will help the industry move on faster.
The premise is always the same - if something is missing in {technology I don't like}, it's a deal-breaker, and when it's not or was always there - it never mattered, or is harmful actually, that is, until {technology I like} gets it as well.
I wasn't advocating java for gamedev. Just pointing that, this new api is a nice addition. And I am glad that jvm ecosystem is improving.
To be fair, if I was starting a game project I wouldn't stay in Java/C# level. Depending on the project, something like C, C++, zig might be more practical. Ironically I believe they would be easier for iterating ideas and deploy into different platforms (mobile, wasm etc.).
*C/C++ tooling and verbosity pains, what PL dev progress is for? C# feels more modern than some of the "modern" alternatives but eh.
C# also has function pointers (managed/unmanaged) and C exports with NativeAOT.
Might just build a SDL2-wrapper for ffi just as an FFI and FMI-exercise.
[0] https://gist.github.com/raysan5/17392498d40e2cb281f5d09c0a4b...
Very useful, especially the prebundled platform bindings.
The Swift folks have put a lot of effort into attaining a stable ABI that's native to their language. They can achieve that because Swift is the officially endorsed language for development on Mac OS and iOS, so it (together with the platform itself) can set a standard that other languages will have to live with.
In a way, software VM's like the JVM and CLR can also be said to define 'ABIs' of sorts within their runtime, that every language implementation on these runtimes will have to deal with.
I can also point to the GCC Inline Assembler as an excellent way to call arbitrary functions whether they implement the standard C procedure call standard or not. By providing the list of arguments and what register they correspond to, along with the clobber list, you know everything you need to know to call the function. So it's more suitable for "fastcall" type functions where you need the arguments to correspond to particular registers.
But of course, ASM isn't portable.
It describes how to adopt memory from C and have C adopt memory you allocate, and gives control over how memory is allocated in an arena.
The arena has lifecycle boundaries, and allocations determine the memory space available. Java guarantees (only) that you can't use unallocated memory or memory outside the arena, and if you access via a (correct) value layout, you should be able to navigate structure correctly.
The interesting stuff is passing function pointers back and forth - look for `downcall method handles`.
Tried it again yesterday on java 22, and the helpers from jextract are waaaay better. I actually completed a MVP implementation this time in a couple of hours. This could perhaps be released as library if I find the effort to wrap it in a meaningful way!
We currently wrap this in java by calling the binary with subprocesses, which has been working great at some latency overhead. The big bonus of this though, is that we can kill the process from java when it misbehaves. Putting this C code inside Java again, means we likely lose that control.
Glad to see things are progressing!!
This is similar, except more boilerplate and much, much slower.
[0] Objects that need pinning are pinned(by toggling a bit in object header), byrefs are pinned by simply storing them on the stack, arguments that need marshalling involve calling corresponding marshalling code. That code can allocate intermediate data on heap, on stack or call NativeMemory.Alloc/.Free C-style.
[1] Overhead can be further reduced by 1. annotating FFI calls with [SuppressGCTransition] which saves on possible arguments stack spills and GC helper call, replacing the call with a single flag check and optional call into GC in epilog, 2. in NativeAOT, p/invokes can be "direct" which saves on initialization checks and indirections (though they are reduced in JIT as it can bake data directly into codegen after static init has finished on recompilation). This has a tradeoff as system's dynamic loader will be used at application startup instead of regular lazy initialization and 3. direct p/invokes can be upgraded to static linking, which transforms them into direct calls identical to regular C calls save for the same GC flag check in post-condition. This comes with compiling .NET executables and libraries into a single statically linked binary (well, statically linked for the native dependencies the user has opted into linking this way).
While a step closer to Valhala, the whole dev experience is still quite lacking versus what .NET offers.
Currently is too much like making direct use of InteropServices.
Doubtful, given that this is something we worked hard to avoid. To be efficient, a P/Invoke-like model places restrictions on the runtime, which inhibits optimisation and flexibility and this cost is worth it only when native calls are relatively common. In Java they are rare and easily abstracted away, so we opted for a model that offers full control without giving up on abstraction, given that only a very small number of experts (<1%) would directly write native calls and then hide them as implementation details. I'm not saying this approach is the right one for all languages, but it's clearly the right one for Java given the frequency of native calls and who makes them.
Of course, you can wrap FFM with a higher-level P/Invoke-like mechanism, but it won't give you as much control.
I will rather keep writing C++ with JNI, instead of enduring the current boilerplate, specially if I already need to manually create header files to feed into jpackage, for basic stuff like struct definitions, which I don't feel like writing by hand.
As for performance, this is something I agree with neonsunset, unless we see Techpowerbenchmarks level of Panama beating P/Invoke, it is pretty much theoretical stuff at the expense of developer convience.
s/jpackage/jextract/gI would encourage those who think that we're consistently making suboptimal choices for Java compared to choices made by significantly less successful languages to consider whether it is possible that their preferences are not aligned with those of the software market at large. Java is and aims to continue being the world's most popular language for serious server software, and that requires tailoring designs to a very large audience.
I always notice a certain lack of respect on forums such as HN for the world's most consistently successful and popular languages -- JS, Java, and Python. Different programmers have different preferences and I'm all for rooting for the underdog now and again, but you simply cannot consistently make wrong decisions over a very long period of time and yet consistently win. What we do may not be everyone's cup of tea (no language is), but it is clearly that of a whole lot of people. We work to offer value to them.
[1]: E.g. the design of native interop has significantly impacted that of user-mode threads (or lack thereof: https://github.com/dotnet/runtimelab/issues/2398) in both .NET and Go, and we weren't willing to make such tradeoffs in either performance or programming model.
That is it, other use cases, have other programming stacks.
As such our native libraries are written in consideration to be consumed at very least, across .NET (P/Invoke, C++/CLI, COM), Java (JNI), nodejs (C++ addons), Swift.
So to move the existing development workflow from JNI to Panama, it must be an easy sell why we should budget rewrites to start with.
Also in regards to "hate", if all decisions were that great there wouldn't be needed to create a new library support group to help Java ecosystem actually move forward and adopt new Java versions, as I learned from JFokus related content.
To elaborate just a bit more on what I wrote in my previous comment, to get a straightforward interop with C you need to place certain restrictions on the runtime which limit your ability to implement certain abstractions such as moving GCs and user-mode threads. Because native interop requires special care anyway due to native memory management, which makes it significantly more complex than ordinary code and so less suitable for direct exposure to application developers -- so it's best done by experts in the area -- and on top of that native calls in Java aren't common, we decided not to sacrifice the runtime in favour of more direct interop. As a result, native interop is somewhat more elaborate to code, but as it requires some special expertise and so should be hidden away from application developers anyway, we decided it's better to place the extra burden on the experts doing the interop rather than trade off runtime capabilities and performance. We think this is the better tradeoff for Java. Consequently, we have both compacting collectors and no performance penalty for native calls on virtual threads. Other languages made whatever tradeoffs they thought were right for them, but they did very clearly sacrifice something.
In general, being successful and popular had little to do with how well a PL is designed. Visual Basic, PHP, and even C are some historical examples that I have plenty of personal experience with.
Perhaps, but it is fairly easy to design a product for a small, self-selecting group of fans who find the aesthetics appealing and so declare the design good for their taste. Unless a language becomes heavily used in codebases that are maintained for years by a large variety of programmers, it's hard to tell how well it is actually designed as a mass-appeal product.
Two of the three languages you mentioned weren't able to attain nearly the same success as Java for as long a duration. I'd give C a similar success score because what it lacks in popularity it still makes up for in longevity, being almost twice as old. There are good reasons for why C is still as popular as it is. For example, in its domain -- which requires compilation to exotic architectures -- "good design" entails being able to easily implement efficient compilers.
And to be clear, I'm not advocating for aesthetics here. It's not like C# is a model of purity, either; but I would say that their choices over the years have been more pragmatic overall from the perspective of someone who needs to write readable, good quality code without jumping through too many hoops or getting lost in the verbiage.
I personally don’t agree with that, C# is very “impulsive” at adding new features, which sounds cool in isolation, but makes the language significantly more complex to understand, and has non-intuitive interactions with other features.
I think C# is quick at going the C++ way, and there is no return from there if we guarantee compatibility.
I much prefer Java’s approach, where yeah, at times one might lack some syntactic sugar/nicety (often greatly overcome by IDE/tooling’s advancements), but over time they do add important ones, but only commit to features that have been earnestly tried and sustainable.
To be fair to Microsoft, they have favoured rich, complex languages for a long time now. They were probably the biggest champions of C++ and TypeScript also doesn't seem to be going down a particularly minimalistic route (an understatement). They're fans of rich languages, and while such languages are not my cup of tea, they do have a large audience (although I think it's a large minority audience).
My point is that when people say that one language is technically superior to another, what they really mean is that it's superior in the technical aspects that they themselves value more than the aspects where the other language is technically superior. This is all fine, except that these personal preferences aren't distributed equally. This is a little like the Betamax vs. VHS debate. Sure, Betamax had a superior picture quality that some valued, but VHS had a superior recording time, which others valued but that latter group was bigger.
As for C# -- strong disagree there. I think they're making the classic mistake of trying to solve every problem in the language and soon, resulting in a pretty haphazard collection of features, quite a few of them are anti-features, making up a pretty complicated language. For example, they have both properties and records, while in Java we figured that by adding records we'll both direct people toward a more data-oriented form of programming and at the same time make the problem of writing setters so much less annoying to the point it shouldn't be addressed by the language (while properties have the opposite effect of encouraging mutation). They've painted themselves into a very tight corner with async/await (the same with P/Invoke, which constrained their runtime's design making it harder to add user-mode threads), and I think they've made a big security mistake with interpolation -- something we're trying to avoid with a substantially different design. Also, while richer languages do have a lot of fans, all else being equal more people seem to prefer languages with fewer features. Our hypothesis is that it's better to spend a few years thinking how to avoid adding a language feature (like async/await or properties) than to spend a few months adding it.
Also, every feature you add constrains you a little in the future (and every language makes this tradeoff early when it's trying to acquire users, but once it's established you need to be more careful). That's why we try to keep the abstraction level high at the expense of a quicker and tighter fit to a particular environment. This delays some things, but we believe it keeps us more flexible to adapt to future changes. It's like having an adaptation budget that you don't want to fully spend on your current environment (I think P/Invoke and properties are such examples of overfitting that pays well in the short term and make you less adaptable in the long term). The complexity budget is another you want to conserve. Add a language feature to make every problem easier, and over time you find yourself not only constrained, but with a pretty complex language that few want to learn.
Green threads experiment proved net negative in terms of benefit but the follow-up work on modernizing the implementation details of async/await itself was very successful:
Issue https://github.com/dotnet/runtime/issues/94620
Technical details https://github.com/dotnet/runtimelab/blob/feature/async2-exp...
The result is such that regardless of p/invoke existence green threads would have been a worse tradeoff.
It also seems that common practices in Java indicate that properties are not a mistake as showcased by popularity of Lombok and dozens of other libraries to generate builders and property-like methods (or, worse, Java developers having to write them by hand). In addition, properties existed in C# since its inception, that's...not a few years.
Not entirely sure about string interpolation but if you are alluding to `var text = $"Time: {DateTime.Now}";`, then it's a non-issue - APIs that care about it in complex contexts like querying a DB or logging can handle it with interpolated string handler API which allows to pass string interpolation expression to methods accepting interpolated string handler types, which can then, for example, generate parametrized query with sanitized inputs, without any friction for the user. Something that Java does not seem to sufficiently appreciate.
Example: https://learn.microsoft.com/en-us/ef/core/querying/sql-queri...
We put a lot of thought into which features we want to add to the Java platform and in what form, and also consider what other languages have done. Sometimes we choose to make different tradeoffs based on what we think are the right tradeoffs for most Java users (a tradeoff that's right for language X may be wrong for language Y [1]), and sometimes we disagree on aesthetics or technical merit. But the choices we've made have worked well for Java. We're well aware of differing opinions, but it seems that we're managing to align with the majority opinions (don't confuse "popular" with "majority"; something like Lombok is quite popular in absolute terms, but is still liked by a minority, i.e. it is less popular than not using it; Kotlin is also quite popular, but it is still more than ten times less popular than Java so does that mean we should follow its decisions?). At the adoption levels enjoyed by JS, Python, and Java, something could be hugely popular in absolute terms yet liked by a minority.
In our primary domain of serious server-side software, no other language has done better (or as well), and we and our users are happy, for the most part, with the choices we've made (except maybe for choices made very early on, but that's true for all languages). The mere fact that sometimes not everyone agrees with our choices (let's be honest, programmers rarely agree on anything) doesn't mean we should change them, especially as languages that go a different way don't seem to be doing as well. Still, different programmers will continue liking different things, and most will continue insisting that their preferences -- however popular -- are somehow "objectively" better with or without bottom-line metrics to support their beliefs.
In general, thinking about a programming language from the perspective of a programmer situated in specific circumstances can be quite different from thinking about a programming language from the perspective of the language maintainer, who needs to take into account different and often conflicting needs of many programmers situated in a variety of different circumstances. The wider the market you're targeting, the more aspects there are to consider and the closer attention needs to be paid to the distribution of programmer preferences.
[1]: E.g. the technical constraints that impact the design and performance of user-mode threads in Rust or C++ are fundamentally different from those that affect Java (re. e.g. the cost of allocating memory, and where pointers are allowed to point). The constraints around async/await in JS -- where a lot of code is already written under the assumption of no intervention -- are also very different from those in Java, where threads have existed from day one.
Ah, that must be why I see FooMethod and FooMethodAsync side-by-side in C# all the time.
Native calls are rare in Java because they're such a pain. If it wasn't so hard to do native calls in Java, it would be common even for non-experts to make use of non-Java libraries.
Note that interaction with native libraries often requires a more careful management of native memory that, though much easier now with FFM, is still significantly trickier (and more dangerous in terms of introducing undefined behaviour) than interacting with Java code regardless of how that interaction is declared in code. In Java, as in Python, interaction with native code -- in the vast majority of cases -- is best encapsulated inside a Java library and not often directly exposed to application programmers.
That's JNI, which really was truly terrible. Java 22 introducing FFM is finally an admission that JNI was crap and a dead end.
There's a reason they're calling it the 'FFM' API and not JNI v2. The API devs were correct in rethinking the approach to native interop.
This just proves my point; being crappy ON PURPOSE is why it's a dead end; it's very difficult to improve something that's been deliberately designed badly.
Besides that, no Java dev in their right mind is going to continue to use JNI once they upgrade to Java 22 and realise FFM exists.
Secondly, our libraries also land on Android applications, and lets see if FFM ever lands on ART.