Jazelle DBX: Allow ARM processors to execute Java bytecode in hardware
en.wikipedia.org
en.wikipedia.org
But Java and similar languages extract more freedom-of-operation from the programmer to the runtime: no memory address shenanigans, richer types, and to some extent immutability and sealed chunks of code. All these could be picked up and turned into more performance by the hardware; with some help from the compiler. Sort of like SQL being a 4th-gen language, letting the runtime collect statistics and chose the best course of execution (if you squint at it in the dark with colored glasses)
More recent work about this is to be found on the RISC-V J extension [1], still to be formalized and picked up by the industry. Three features could help dynamic languages:
* Pointer masking: you can fit a lot in the unused higher bits of an address. Some GCs use them to annotate memory (refered-to/visited/unvisited/etc.), but you have to mask them. A hardware assisted mask could help a lot.
* Memory tagging: Helps with security, helps with bounds-checking
* More control over instruction caches
It is sort of stale at the moment, and if you track down the people working on it they've been reassigned to the AI-accelerator craze. But it's going to come back, as Moore's law continues to end and Java's TCO will again be at the top of the bean-counter's stack.
more like Wirths law proving itself still
Ironically when one dives into computer archeology, old Assembly languages are occasionally referred as bytecodes, the reason being that in CISC designs with microcoded CPUs they were already seen that way by hardware teams.
In theory JIT should be higher performance, because it benefits from statistics taken at actual runtime. Given a smart enough compiler. But as a piece of code matures and gets more stable, the envelope of executions is better known and programmers can encode that at compile-time. That's the tradeoff taken by Rust: ask for more proofs from the programmers, and Rust is continuing to pick up speed.
That's also what the Leyden project / condensers [1] is about, if I understand correctly. Pick up proofs and guarantees as early as possible and transform the program. For example by constant-propagating a configuration file taken up during build-time.
Something I've pondered over the years: a programmer's job is not to produce code. It is to produce proofs and guarantees (yet another digression/rant: generating code was never a problem. Before LLMs we could copy-paste code from StackOverflow just fine)
In the end it's only about marginal improvements though. These could be superseded by changes of paradigm like RAM getting some compute capabilities; or programs being split into a myriad of specialized instructions. For example filters, rules and parsing going inside the network card; SQL projections and filters going into the SSD controller; or matrix-multiplication going into integrated GPU/TPU/etc just like now.
[1] https://openjdk.org/projects/leyden/notes/03-toward-condense...
Android has learnt to have both, and thanks to PGO being shared across devices via Play Store, the AOT/JIT outcome reaches the ideal optimum for a specific application.
Azul and IBM have similar approaches on their JVMs with a cluster based JIT, and JIT caches as AOT alternative.
Also stuff like GPGPU is a mix of AOT and JIT, and is doing quite alright.
I am not so confident with LLMs, when they get good enough programmers will be left out of the loop, and will have to contend to similar roles as when doing no-code SaaS configs or some form of architects.
A few programmers will remain as the LLMs high priests.
That's interesting.
It's controversial to say that in 2024, but not all opinions have the same value. Some are great, but some are plain dumb. The current corporate right opinion is to praise LLMs as end-all be-all. I've been asked to advise a private banking family office wanting to get into LLMs. For advising their clients' financial decisions. I politely declined. Can there be a worse use case? LLMs are parrots with the brain size of the internet. With thoughts of random origin mixed together randomly. It produces wonderful form, but abysmal analysis.
IMHO as LLMs will begin to be indistinguishable from real users (and internet dogs), there's going to be a resurging need to trace origin to a human; and maybe to also rank their opinions as well. My money is on some form of distributed social proof designating the high priests.
When we read about history of Fortran, there are several remarks on the amount of work put into place to win over those developers, as otherwise Fortran would have been yet another failed attempt.
LLMs seem to be at a similar stage, maybe their Fortran moment isn't yet here, parrots as you say, but it will come.
One stereotypical (but not the best) example would be regexes: you basically want to compile some AST into a mini-program. This can also be done with a tiny interpreter without JIT, which will be quite competitive in speed (I believe that’s what rust has, and it’s indeed one of the fastest - the advantage of the problem/domain here is that you really can have tiny interpreters that efficiently use the caches, having very little overhead on today’s CPUs), but I am quite sure that a “JITted rust” with all the other optimizations/memory layouts could potentially fair better, but of course it’s not a trivial additional complexity.
https://www.cpushack.com/2016/05/21/azul-systems-vega-3-54-c...
If you're building hardware masking, it should be viable for low bits too. If you define all your objects to be n-byte aligned, it frees up low bits for things too, and might not be an imposition, things like to be aligned.
Java itself got very good. Though Oracle was blocked to leech money, or have return for their investment, depending on the viewpoint.
ART runs on devices for 1B+ users and is more relevant for the world population as Oracle. Although we can speculate likely Android would have switched to something else if Oracle were to win in the court.
Ironically Android more realised Java’s original light client vision “Write once, run everywhere” if you consider “everywhere” as all around the world, by every human, with various device architectures.
Lol, virtually all business server applications are running on th e JVM in the corner of the world that I see.
Google could have acquired Java, after screwing Sun, and decided to take a bet on not doing it.
ARM is quite capable in vapourware generation. 64bit ARM was press-released (https://www.zdnet.com/article/arm-to-unleash-64-bit-jaguar-f...) a decade before ARMv8 / aarch64 became a thing.
(I'd love to learn more)
It couldn't have, as Dalvik VM is distinct from JVM.
https://source.android.com/docs/core/runtime/dalvik-bytecode
Relegated to the dustbin of history.
I find Java Card pretty puzzling. You go from high-level interpreted languages on powerful servers, to Java and C++ on less powerful devices (like old phones for example), to almost exclusively C on Microcontrollers, and then back to Java again on cards. If. it makes sense to write Java code for a device small enough to draw power from radio waves, why aren't we doing that on microcontrollers?
There have been several more-or-less successful attempts at running higher-level languages on microcontrollers, e.g. .Net Micro Framework and CircuitPython. In all of these cases, though, you tend to struggle with all the native device behavior being described/intended by the vendor for use with C or C++ and the BSP for the higher level environment being an afterthought.
More from Ars (1999) https://archive.arstechnica.com/cpu/4q99/majc/majc-1.html
ART is another matter, though.
To everyone who wants to write: but I didn't read that thread and I find this quite interesting; you are free to find it interesting, but I did read about it 2 days ago and to me it looks like karma farming.
To me it seems like Hackernews, as a whole, goes off on the same kinds of thought-tangents as I do, and that makes the site more interesting. And I was one of the commenters about Jazelle on the thread you mentioned.
Reposting that tangent as a separate submission increases the chance of me finding the conversation.
Even so, why does it affect you so much to see other people getting these internet points? Or that they were somehow unearned?
Who cares.
To simply allow these posts and having them hit the front page when they get upvotes is a valid position. But I think it contributes to a website that is less interesting.
I don't think these posts should be removed, but they should at least be frowned upon, and/or linked to the original comment thread.
Retaining and releasing an NSObject took ~6.5 nanoseconds on the M1 when it came out, comparing with ~30 nanoseconds on the equiv gen Intel.
In fact, the M1 _emulated_ an Intel retaining and releasing an NSObject fast than an Intel could!
One source: https://daringfireball.net/2020/11/the_m1_macs
During years when this instruction set was relevant (though apparently unutilized), Oracle still had very limited ARM support for Java SE, so having a fast interpreter could have been desirable -- but it makes no sense on beefier ARM systems that are able to support decent JIT or AOT support available nowadays.
I never heard of anyone actually using Jazelle, though - I assume JIT ended up working better.
So on a smartcard ... write software in a (uncommon, and when compared with ARM which is a very "rich" assembly language) form of low-level instruction set, and pay both Sun and ARM top$ for the privilege - nevermind the likely "runtime" footprint far exceeding the 256kB RAM you planned for that 5$ card - why? Writing small singlethreaded software in anything that compiles down to a static ARM binary has been easy and quick enough that going off the ARM instruction set looked pointless for most. And learning which parts of "Java" actually worked in such an environment was hard, even (or especially?) for developers who knew (the strengths of) Java well. Because developers and specifiers expected "rich Java", and couldn't care less about the Bytecode. JITs later only hoovered up the ashes.
But the reality was that JIT allows code to get faster over time, as the JIT improves.
Things like Jazelle let chip manufacturers paper over a paper objection.
There were those "LISP machines" in the early 1980s but when Common Lisp was designed they made sure it could be implemented efficiently on emerging 32-bit machines.
Ehh .. PGO is only somewhat better for JIT than AOT. More often for purely-numerical code the win is because the AOT doesn't do per-machine `-march=native`. It's the memory model that kills JVM performance for any nontrivial app though.
Contrast this with the problem of specialization in AOT languages, which can easily result in bloated binaries (PGO does help here quite a lot, that much is true). For example, generics might output a completely new function for every type it gets instantiated with - if the function is not that hot, it actually makes sense to rather try to handle more cases with the same code.
For a while Linus Torvalds, of the Linux kernel fame, worked for a company called Transmeta, https://en.wikipedia.org/wiki/Transmeta, who were doing some really interesting things. They were aiming to make a highly efficient processor, that could handle x86 through a special software translation layer. One of the languages they could support was picoJava. IIRC, the processor was never designed to run operating systems etc. natively. The intent was always to have it work through the translation layer, something that could easily be patched and updated to add support for any x86 extensions that Intel or AMD might introduce.
It’s more for embedded use. Might give someone ideas, though.
On modern Cortex-A systems, there are enough resources to make JIT feasible. On smaller systems, AOT is a reasonable alternative.
This is so old that its replacement, ThumbEE, had already been deprecated as well.
The JVM is stack-based, right? So it'd be an interpreter (in microcode)? Unless there's some kind of "virtual" stack, as spec'd for picoJava.
I'm less clear on how Jazelle would implement only a subset of the bytecodes.
Am noob. And a quick scholar search says the relevant papers are paywalled. Oh well; now it's just a curiosity.
Stack-based CPUs are cool, right? For embedded. Super efficient and cheap, enough power for IoT or secure enclaves or whatever.
But it seems that window of opportunity closed. Indeed, if it was ever open.