WebAssembly for the Java Geek
javaadvent.com
javaadvent.com
Validation is faster due to it being structured, but stackmaps make it fast for class files as well. It is lower level, but is that really a good thing? It is quite trivial to just assign a huge ‘long’ array to some jvm byte code for the same result, and while I will look at the linked GC proposal, because haven’t been kept up-to-date on that, I think having a good GC is the single most important thing, as that is much harder to do than just providing a linear array. Which will in turn decimate the implementations (or at least mark a toy vs prod ready divide).
Also, will that structured nature not make it a less than great compilation target? This means that certain, more niche languages will have to use ugly (and slow) hacks to get implemented. And it being lowish level it is much harder to optimize it well (in that you express the exact semantics).
> It is lower level, but is that really a good thing? It is quite trivial to just assign a huge ‘long’ array to some jvm byte code for the same result
That's basically what the WASM implementations on the JVM currently do. But it is actually _higher_ level from a dev POV due to language choice (at a not-huge performance expense).
Why hasn't this been done by anyone yet?
Also confused why this does not have a binary windows release yet: https://github.com/bytecodealliance/wasm-micro-runtime
Edit: Epsilon is not the answer here you can stop mentioning that.
This is the way I code C now allready.
I mean Java isn’t the worst tool out there, but even C# outclasses it as a language. The biggest plus for Java is the massive, mature ecosystem which I imagine mostly evaporates when you’re only using a self restrictive subset of the language itself.
But you may want to ask yourself why you want to do that. OpenJDK's GCs have become really, really good in both throughput and latency, and the main thing you pay in exchange is memory footprint.
Maybe object pooling (which helped performance in old JVM's in the 90s) will make a comeback? ;)
Basically only static atomic arrays (int/float) in classes will be allowed, for cache and parallelism, and AoS up to 64 bytes encouraged to avoid parallel cache invalidation.
And I'm even considering dropping float, and only have integer fixed point... but then I'll need to convert those in shaders as GPUs are hardcoded to float.
On mobile here would reference on PC.
The bytecode is only meant to run in my VM.
[0] https://github.com/real-logic/simple-binary-encoding/wiki/De...
If you keep your object allocation in check you can use this to guarantee no GC pauses.
What would be the point? Java programs assume GC (or at least, they assume they can allocate memory and not worry about when it will be freed). If you don't have GC then you're not going to be compatible with extant Java programs, so what's the point in trying to be "Java compatible" at all?
only users of long running programs assume GC, the program itself doesnt care.
6) int arrays do not have cache misses (but you might need to pad them to avoid cache invalidation) 7) parallel atomic multicore works on int arrays out of the box!
From my perspective I am happy with the tradeoff of hanving gc pauses and not needing to manage memory manually.
2) VMs avoid crashes upon failures that cause a segmentation fault in native opcodes, with a VM you can keep the process from crashing AND get exact information where and how the problem occurred avoiding debug compiles with symbols and reproducing the error on your local computer.
Right now I use this with my C++ code to get somewhere near the feedback I get from Java (but it requires you to compile with debug): http://move.rupy.se/file/stack.txt
The question you really should ask is why are people using C++? Performance is only required in some parts of engines, to have a VM without GC on top should be default by now (50 years after C and 25 years after Java).
I don’t have anything against Java, but it seems like you lose a lot of the benefits of using when you take away allocation.
Java + WASM is probably what I'll end up with, don't like the AoT step from a distance we'll see.
I'm looking at all Risc/stack op/byte-code, and it's disheartening in how many ways humans can copy the same thing differently.
C#, RISC-V, Java, WASM, ARM, 6502 ASM, uxn, lox the list goes on and on...
2) Why exactly? 3) that’s why you have a GC. But if I really want to understand what you mean (I assume you meant non-memory resources not being explicitly closed?), try-with-resources and cleaners solve the problem quite well.
4) Does it really matter if the last 2 decades were spent on optimizing allocation to the point where it is literally a pointer bump and like 3 basic, thread-local instructions? Deallocation can be amortized to be practically zero cost (moving GC), at RAM’s expense. Where it really really matters though, you could always do ByteBuffers, the new MemorySegment’s or just straight sun.misc.Unsafe pointer arithmetics. That string allocation will be escape analyzed and be stack allocated either way.
I don’t even know what do you mean by static allocation. You mean like in embedded, having fix sized arrays and exploding when the user enters a 32+1 letter text? I really don’t miss that. 6) value classes are coming and solving the issue. Though it begs the question, what’s the size of your average list? Also, how come it doesn’t matter for all the linked lists used extensively in C? Also, see point 4. 7) I don’t get what you mean here, do you mean not doing atomic instructions, because java don’t have out-of-thin-air values? There are rare programs that can get away with that, but I think Java has quite a great toolkit for synchronization primitives to build everything (most concurrency books use it for a reason).
The first and last bottleneck of computers is and will always be RAM. Both speed, memory size and energy. A 256GB RAM stick uses 80W!!!! Latency is increasing since DDR3 (2007) and we have had caches to accommodate for slow RAM since 386 (1985) (3 always seems to be the last version, HL3 confirmed? >.<):
You need to cache align everything perfectly: 1) all data has to be in an array (or vector which is a managed array but i digress) 2) you need your types to be atomic so multiple cores can write to them at the same time without segmentation fault (int/float). 3) You need your groups/objects/structs to perfectly fill (padded) 64 bytes. Because then multiple cores cannot invalidate each others cache unless they are writing to the same struct.
So SoA vs. AoS never was an argument! AoS where structures are exactly 64 bytes is the only thing all programmers must do for eternity! This is the law of both X86 and ARM.
So an array of float Mat4x4 is perfect and I suspect that is where the 64 bytes came from. But here is another struct just as an example:
struct Node {
int mesh, skin;
Vec3 spot, pace;
Quat look, spin;
};Most applications have plenty of objects all around that are rarely used and are perfectly fine with being managed by the GC as is, and a tiny performance critical core where you might have to care a tiny bit about what gets allocated. This segment can be optimized other ways as well, without hurting the maintainability, speed of progress etc of the rest of the codebase.
Cachelines are 64 bytes on all modern hardware.
They will probably never change this value ever.
Everything fits into 64 bytes if you make the effort.
And if it doesn't you have to use two Arrays of 64 byte Structures and pad the last.
This is non negotiable and I'm completely baffled nobody has mentioned this yet.
I call this law: Ao64bS (did I invent my first law?) :D
But since it’s AMP and not SMP, sharing work across cores doesn’t necessarily work how you expect it to.
128 bytes is perfect 2 x 64! So even if the risk of cache invalidation goes up even if two cores are not writing to the exact same structure the alignment still works!
Good job Apple!
This seems very high to me, since most high capacity server sticks have no heatsink. Have you got a source?
How would you handle the ~ 1000 java.x classes that even a minimal HelloWorld programs uses. There are plenty of random allocation.
That depends what you mean by "this." The Java spec requires you to support `new A()` by returning a fresh object or throwing a VMError (like OutOfMemoryError). If by "this" you mean that allocations would succeed until some fixed amount of memory is exhausted, then, as others have pointed out, this has been done even in OpenJDK. If by "this" you mean that the allocation of some particular set of objects -- say, only those allocated during class initialisation -- would succeed regardless of memory consumption and all others would fail with an OutOfMemoryError, then I guess that it hasn't been done in that particular way because people haven't found it particularly useful as most Java programs would fail, but you can give it a try.
RTSJ, the specification for hard-realtime Java [1][2], actually goes further than that and supports both "static" memory (ImmortalMemory) and arenas (ScopedMemory). So if that's what you mean, then it has been done.
On the .NET side, the language is more suitable because you can define your own value types (coming soon in Java AFAIK), and there is a lot of people using C# for game dev which is a big use case for GC optimization.
Unity has been integrating unmanaged allocators in their engine, which lets developers skip the GC much more easily (without having to manually preallocate and reuse memory, which isn’t anyone’s favorite workflow, and doesn’t save you much compared to a fast native allocator). I’ve also seen a couple of roughly equivalent projects, eg. Smmalloc-CSharp, github.com/alaisi/nalloc.
Short lived objects are extremely cheap on the JVM with the right GC. Almost stack allocation cheap. So if you write code that avoids too many long lived objects there is almost no overhead to a GC.
For long lived objects GC based compaction can give you even some performance advantage over manual allocation due to better memory locality. But that heavily depends on the application of course.
In my experience people blame the GC way too early. Most often the application is just poorly written and some small tweaks can fix gc spikes.
I used to do Java apps which didn't allocate memory in the long run, and so did some banks for high frequency trading, by reusing objects on the spot or using pools for objects with a life cycle. It imposes some programming style/patterns, in particular for APIs design, but it's perfectly doable, so I guess people prefer to just do that. In what context that wouldn't be enough? For proof?
to a very large extent, it is open till LLVM upstream is more functionally complete to support GC.
wasm doesnt support GC to the way Java and Graal does. Which is why Graal is seriously impressive piece of tech. so it is not wasm vs java. Because that is really apples verus oranges.
it is just a matter of time before java gets compiled to wasm. that's what makes me puzzled about this article. Maybe its really about rust or golang vs java.
0: https://extism.org/docs/integrate-into-your-codebase/java-ho...
When running a GC language in its own runtime, usually there is at least 1 GC thread, that runs concurrently with the app code and pauses/preempts the main thread.
Afaik, in WASM, your application only has CPU time when you are inside a WASM function and threading is incredibly hacky and may be disabled due to security reasons related to side-channel attacks (read up on `SharedArrayBuffer` and cross-origin site isolation).
New():
ref <- allocate()
if ref = null
collect()
ref <- allocate()
if ref = null
error "Out of memory"
return refHowever, true "Java Geeks" should try the JavaScript support in TeaVM. It is mature and battle-tested. The resulting code runs great in a browser.
Live TeaVM-based game: https://frequal.com/wordii
Getting started with TeaVM: https://frequal.com/TeaVM/
Performance Comparison: https://frequal.com/java/TeaVmPerformance.html