I have written a JVM in Rust
andreabergia.com
andreabergia.com
Also note that I cloned the repo and tried to run `cargo test` every test fails with 'should be able to add entries to the classpath: InvalidEntry(".../vm/rt.jar")' vm/tests/integration/real_code_tests.rs:15:10
[1] https://github.com/andreabergia/rjvm/blob/be9c54066c64a82879...
[2] https://manishearth.github.io/blog/2021/04/05/a-tour-of-safe...
[3] https://without.boats/blog/shifgrethor-iii/
[4] https://coredumped.dev/2022/04/11/implementing-a-safe-garbag...
There is a performance cost for a VM having its own virtual callstacks like this, but it makes GC tracing much simpler. (It also makes implementing interesting concurrency and control flow primitives like coroutines or continuations much easier too.)
[1] https://github.com/andreabergia/rjvm/blob/main/vm/src/native...
[2] https://github.com/andreabergia/rjvm/blob/be9c54066c64a82879...
[3] https://github.com/andreabergia/rjvm/blob/be9c54066c64a82879...
I guess the solution would be to add an explicit API to create a GC root, invoked by native methods (which is a bit complicated by the fact that I use a moving collector).
Many years ago I was using SpiderMonkey in a c++ project and I seem to remember there were some APIs for native callbacks to invoke that rooted values. Same problem and similar solution. :-)
This is why I do in the Wren VM. Any time a native C function has the only reference to a GC-managed object and it's possible for a collection to occur, it calls a function to temporarily add the object to a list of known roots.
One thing struck me as a bit odd:
> In particular, it does not support: generics
What kind of support is there for generics in the JVM? Maybe I'm too naive to assume that due to type erasure on bytecode level everything is just an Object, ie. a reference type? Or do you mean the class definition parser - but then, you don't really have any checks in place to see if the class file is valid (other than the basic syntax)?
OTOH lack of string interning is super strange [it's trivial to implement], and w/o it JVM is not a thing. String being equal by reference is important, and part of JLS.
Lack of thread makes the entire endeavor a toy project.
But, as the previous commenter attempted to point out to you this project is a *self-described* toy project.
On the very page that is linked, the author of the JVM specifically says:
"I want to stress that this is a toy JVM, built for learning purposes and not a serious implementation."
Thus, absolutely no one disagrees that its a toy JVM. They just want you to stop being dismissive of someone's toy project by repeatedly pointing out its a toy project and not a "not a thing"
SQL (ACID) over multiple non-cache-coherent nodes is extremely difficult to pull with regards to consistency, though.
Thats... why it's a toy! I'm really not sure what you're missing here.
> Moreover, a string literal always refers to the same instance of class String. This is because string literals - or, more generally, strings that are the values of constant expressions (§15.28) - are "interned" so as to share unique instances, using the method String.intern.
And later:
> Literal strings within different classes in different packages likewise represent references to the same String object.
Source: https://docs.oracle.com/javase/specs/jls/se8/html/jls-3.html...
But that does - as far as I can see - say nothing for non-literal strings.
[edit]: formatting
Nothing much to think -- distinct objects must have distinct references [e.g. new String("a")!=new String("a')], literals must have the same references for the same values [e.g. "a"=="a"].
if("abc" == "abc") { System.out.println("correct"); }
if(new String("abc") != new String("abc")) { System.out.println("correct"); }
So, not having proper string interning support means that you mis-execute certain programs.> say nothing for non-literal strings
yes, of course.
https://docs.oracle.com/en/java/javase/11/docs/api/java.base...()
Also interning strings to optimize equality checks to be able to use pointer comparison is dangerous for external inputs since iirc at some point interned strings could permanently be stored (unless implemented by a WeakSet) and attackers could fill up your heap (or cause other GC issues since the entire interning functionality is a cache) by filling up your interning lists with crap.
I never said String must be equal by reference when their content is. However string literals must be equal by reference. I thought Mentioning the JLS would make it obvious, esp. having 'intern' in the context
Not certain about whether `String.intern` is permanently stored; I rather suspect that it sweeps the existing strings since iirc the java string has a hash associated with it anyway.
yeah, as stated by the author in the line that says "I want to stress that this is a toy JVM, built for learning purposes and not a serious implementation."
This is generated when you do something like: final Main value = list.get(0);
http://henrikeichenhardt.blogspot.com/2013/05/how-are-java-g...
About the generics - some people have pointed out the same on reddit, and yeah, you are correct. The only thing that should be done is to read the Signature attribute that encodes the generic information about classes, methods, and fields (https://docs.oracle.com/javase/specs/jvms/se7/html/jvms-4.ht...)
As a matter of fact, I just did a test and the following code works! :-)
public class Generic {
public static void main(String[] args) {
List<String> strings = new ArrayList<String>(10);
strings.add("hey");
strings.add("hackernews");
for (String s : strings) {
tempPrint(s);
}
}
private static native void tempPrint(String value);
}IMAO, Android kind of achieve that...kind of. They write lots of OS logics in Java (or Kotlin) but mixing lots of system services written in native code at the same time, interconnected by the famous (or infamous?) Bind IPC.
[1] https://en.m.wikipedia.org/wiki/Java_Card
[2] https://superuser.com/questions/362567/are-there-any-credit-...
[3] https://www.oracle.com/java/java-card/
All of this is just as easy as, if not easier than, using ChatGPT. It's unclear that such a tool even serves this purpose (retrieval of basic facts) adequately, so it should probably be avoided in the future.
(fwiw I have a bit of prior experience here)
Even fully passive NFC tags contain logic that needs power to talk NFC back to the reader, there’s no such thing as just reading data via NFC.
You know those 4 lights on a contactless card reader? They indicate different transaction stages between the card and the terminal. I don’t find them that useful because it’s so fast they all appear to light at the same time, but that’s what they are!
If they were just passive tags, they wouldn’t be very secure, they have cryptographic processors onboard with private keys that can sign stuff for the terminal and your bank. The specs for the interaction are all public if you’re interested (I wouldn’t be!) and lookup contactless EMV.
https://www.cardlogix.com/product/cardlogix-credentsys-lite-...
And on another source:
Visa became the first large payment company to license JavaCard. Visa mandated JavaCard for all of Visa’s smartcard payment cards. Later, MasterCard acquired Mondex, and Peter Hill joined as their CTO, licensed JavaCard, and ported the Mondex payment platform to JavaCard.
Source: https://javacardforum.com/2022/07/28/the-birth-of-javacard/
It's almost bizarre what that thing does.
Though mind you, it is a very limited subset of Java, not the standard one.
respect yourself enough to look at primary sources
It didn't help any that it was clear from the initial post you were questioning someone with domain knowledge, which was later gently indicated to you
Engineering is all about tradeoffs, and ‘works’ is pretty high praise frankly.
Ideological purity is rather subjective, and has an unfortunately poor track record of real world success.
And perhaps proposes a concrete alternative that matches those constraints better?
Don't forget to include things like long term support, developer time, interoperability, etc.
Kind of, yes. Since instead of evaluating based on an understanding of the technology you're saying "In this universe was chosen for a project, the project was successful, show me the universes where the alternative decision was made".
I'm saying that if you think something is crap and doesn't meet customer needs, at least propose a concrete alternative you believe is better so someone can respond meaningfully! Or concretely what concrete needs are not being met!
Currently, we have one example of something that all evidence leads us to believe fits the universe as it exists, at least in that specific niche.
If you think it doesn't, how doesn't it? Or if you're saying there is something better, are you saying that is hand written assembly? Or TurboPascal? Or ADA? Or some as yet not designed system?
I'm not asking for an alternative universe. I'm asking you to support your statement with enough details it can be assessed in the current universe.
You: If it is the convenient alternative than it is a good choice
Me: That is a bad justification for something being a good choice
You: What criteria then?
Me: An understanding of the engineering principals/ domain
Am I accurately summarizing this conversation so far? This isn't about alternatives, or what else they should have done. Something can be a bad fit and the right choice.
As an example, I could say "Java dominates that section due to historical artifacts of business, not technology. Java is a bad fit for this type of work otherwise because of the complexity involved in implementing a Java VM in hardware". I can then also say "Java is the only real choice because of those historical artifacts so I have to recommend that you use it unless you're willing to build your own hardware from scratch".
I actually don't have to propose any alternatives at all, hopefully you can see that - we can just evaluate Java as a language (complex VM, assumes a heavy runtime) against the constraints (custom hardware, low energy) and see that the fit is weak. Obviously people overcame that and made it work, and because of that Java is the obvious choice for this technology.
In that scenario, it's literally the best possible fit.
That you don't have any viable better alternative at hand may be further evidence of that? (and I don't mean from a standards basis 'well, it's locked into financial rules now, so gov't intervention'). I mean, what else was going to work considering all the factors involved? What else could work better, considering the factors involved?
JavaCard is in fact so widely used and implemented (SIM cards, bank cards, health cards, passports, etc.) that it probably has literally 10's of billions of devices manufactured using it (3.5bln claimed as of 2010 - https://www.oracle.com/technical-resources/articles/javase/j...), in essentially every high value target rich environment niche you can think of, and at extremely low costs. Literally sub-cent per-item.
And with very high environmental stresses (like debit cards getting sat on, left in hot cars, run over, dropped in puddles, jammed into random dirty readers over and over again, etc.), those devices keep working.
And everyone from random countries gov'ts to random financial firms to telcos have managed to implement what they need in it without too much difficulty, and a minimum number of security issues. Which is frankly astonishing if you've ever dealt with folks like that.
So love or hate Java, or JavaCard from a stylistic perspective - any perceived complexity for implementing a Java VM in hardware has had no practical economic effect, or slowed down implementation meaningfully.
It's fit for purpose.
Probably also ugly and feels gross using them sometimes, but a lot of fit for purpose stuff is until you've experienced the alternatives. Hopefully you never have to fix a sewage lift station pump, or clear a clogged sewer line, or clean out a transmission after it's burned out.
Each of these has literally hundreds of years of specialized knowledge and expertise behind their often boring looking facades. They're all amazingly complex if you learn about them. And they're all better than throwing sewage in the street, or carrying everything on horseback. And they're beautiful in their own way when you appreciate why they are how they are.
Even if they're not shiny and flashy, there is beauty in them, because they work well.
And they're still amazing engineering marvels, necessary for our lives as we know them and based on the actual engineering principals involved and the problem domain.
So we fundamentally disagree.
> has had no practical economic effect, or slowed down implementation meaningfully.
This goes back to my "to prove me wrong you have to show me alternate universes where other options had that investment made under the same circumstances".
This is an assertion that can only be disputed with an alternate universe.
Notice the difference?
To propose alternative solutions based on practical economic effect just requires a reasonable degree of comparison to projects of similar scope at other times, reports of difficulty from various vendors, and comparisons of end price for this solution, end price for other solutions of similar scope, to the overall scope of the solution and value it brings. None of which requires perfect alternative universe A/B testing to come to some reasonable analysis.
If another solution could be done for half the price (say $0.005 per unit, instead of $0.01 per unit) but the perceived value for vendors is $1/unit - then it's hard to say there is any practical economic effect going either way. Neither solution would block profitability or value. That said, they could easily be compared and better/worse solutions could also be determined or tradeoffs analyzed based on that data, also without perfect alternate universe A/B testing to come to some reasonable analysis. Industry does this all the time at scale, including projected costs of implementation of various solutions.
If the solution was rolled out within a timeframe considered useful/expected for this kind of solution, then it also didn't slow down implementation meaningfully - as in it didn't block it, or add serious delay. If there is another solution which could have been done in half the time, that's cool. But it wasn't required. Identifying such an alternative, if one exists, could be done if you have any data, without having to do an alternative universe A/B test. Though since the proof is in the pudding, to REALLY be sure maybe it would. But that's hardly what I've been referring to or asking for, clearly.
Doesn't mean they wouldn't have been better solutions, and proposed them as alternatives can easily be done without parallel universes! In fact, chances are they have already been implemented somewhere in another niche, so there is adequate data to do so.
No.
> If the solution was rolled out within a timeframe considered useful/expected for this kind of solution, then it also didn't slow down implementation meaningfully - as in it didn't block it, or add serious delay.
Also no.
JVM wastes cycles on things like classes, which is not necessary at all. Going forward, Rust has already proven that you can do things at compile time to guarantee things like memory safety.
ART also compiles to native since Android 5.
Android 14 is around the corner, time to keep up with the times.
Does it always make sense? No. But clearly a bunch of folks have found it valuable in many niches.
32 bit 4mhz processor with ~64kb of nvram all running off of an induction charge!
https://source.android.com/docs/core/connect/esim-overview?h...
I think the e-sim idea is probably a net security benefit imo, apple has leveraged the carriers out of a fairly dangerous tool.
https://www.youtube.com/watch?v=31D94QOo2gY
Recently I was thinking about what a program written to take advantage of Optane's persistent-memory model would have looked like, if you use it like RAM it's forever gonna be slow shitty RAM. Javacard seems to be the closest hit for that, in some ways. Maybe some of the higher-tier javacards have GC.
I guess at that point it's basically a JVM application state snapshot, which is the same thing, so maybe not any better.
When I think about a "Java OS", I imagine a JVM running in kernel mode, providing minimal OS functionality (scheduler, access to hardware I/O ports) and there not being any kind of userspace.
While Android isn't proper Java, converting JVM bytecodes into a better format for embedded deployment is quite common on embedded world.
PTC, Aicas, Gemalto, microEJ, WebSphere Real Time, Aonix,....
Current Java OS as per your definition, would be PTC and Aicas real time JVMs for bare metal deployments in embedded scenarios.
- SavageJE
- microEJ
- PTC and Aonix bare metal Java runtimes
- SunSPOT mit SquawkVM
[0] - https://en.wikipedia.org/wiki/Singularity_(operating_system)
[1] - https://en.wikipedia.org/wiki/Midori_(operating_system)
Seems like a reasonable goal.
The reason to do this, beyond the inherently neat Inception factor, is that JVMs are a PITA to work on because they're normally written in languages like C++ or Rust which optimize for performance and manual control over productivity. That makes it hard to experiment with new JVM features or changed semantics. If you could write a JVM in a high level very productive language like Java (or Kotlin or Scala) then the productivity of people writing and experimenting with JVMs would go up. It would also make it feasible for "ordinary" Java devs to actually fork the JVM and modify it to better suit their app, at least in some cases.
There's also something conceptually cleaner about having a language and its runtime implemented purely in itself. As long as you don't mind the circularity, that is.
Espresso for example has hot-swap features HotSpot doesn't have, so you can modify your program as it's running in more flexible ways than what regular Java allows.
It all adds up.
But I don't have enough experience with Rust to really have formed an opinion on that yet.
But IMO it's very worth it. What do you get out of wrestling with the unholy, fractured C/C++ ecosystem? The privilege of being able to _use other people's code_. You get that without much hassle in Rust.
The borrow checker slowed me down a lot at the start but these days it's not a big deal. Eventually you internalize the easiest ways to make it happy (clone and pass references as needed) and can reserve tricky optimizations for hot paths where they actually matter.
I would say that today, most of the times I sit down to write some Rust I'm nearly as productive as in higher-level languages... but then ~20% of the time I get bogged down in complicated type signature stuff (especially with async code, ugh) and it's a time suck.
fn execute_instruction( &mut self, vm: &mut Vm<'a>, call_stack: &mut CallStack<'a>, instruction: Instruction, ) -> Result<InstructionCompleted<'a>, MethodCallFailed<'a>>
When I try to add a lifetime to the `Err` variant of a `Result` and that lifetime is invariant (which it is due to `vm` and `call_stack`) it usually means that I can't use the question mark operator or have early returns in the code[1]. This makes error handling more verbose and less readable. Is that your experience as well?
[1] https://users.rust-lang.org/t/nll-and-early-return-not-allow...
In that case I don't understand what the point of 'a is on VM and CallStack. You can create[1][2] those with any unbounded lifetime (including 'static[3]), which means it is not constraining anything. What is the lifetime 'a doing here? Why not remove it?
[1] https://github.com/andreabergia/rjvm/blob/be9c54066c64a82879...
[2] https://github.com/andreabergia/rjvm/blob/be9c54066c64a82879...
[3] https://github.com/andreabergia/rjvm/blob/be9c54066c64a82879...
I also struggled with a got a ton of errors from the borrow checker initially, and I fixed many of those with a lot of explicit lifetimes, but it's not impossible that in some places they are unnecessary.
That's not being expressed in the type system. The lifetime 'a is unbounded (meaning you can make it anything you want, including 'static) so anything that shares 'a can outlive the vm without rust complaining. it would be no different then if you removed 'a completely. If you wanted to ensure anything couldn't outlive the vm you could tie the lifetime to a reference to the vm, but then the vm can't hold those values (it would be a self-referential lifetime).
If they're interested in bolting on a GC, it couldn't hurt to look at MMtk. (https://www.mmtk.io/) Some high quality collection algorithms, written to be pluggable to various VMs, and written in Rust.
Haven't read, but I bet it's likely related to expectations around the x86_64 memory model & atomics. In the long run I see no reason why it couldn't be made portable, but I imagine the authors efforts are elsewhere for now.
If you’re looking for a job then ping me on Twitter, Mastodon or my work email, I’m sure you can figure them out from my user id here.
I've been a professional software developer for almost 10 years, and I _know_ I'm competent (and not an impostor) as demonstrated by my current position and ability to ship things.
However, lately after viewing developer blogs I become overwhelmed that I actually don't know enough and am not a "real" developer. I seem to have formed a notion of an ideal developer in my head and I compare myself against this imagined construct which leads to these feelings. I admire how these people have so much deep knowledge and can express themselves so clearly and concisely, then wonder why I am not like that.
I barely have the energy after work after taking care of my family to do anything further, and I know programming isn't everything but I do have a desire to learn more and improve myself.
I recognize this isn't healthy nor is it rational, but it's just a feeling I can't shake lately.
In short: recognizing your insecurities is the first step. The next step is figuring out what's important to you, shedding impossible to achieve and irrational ambitions, prioritizing your goals in life, and articulating concrete steps to further them.
I've been deep-diving into sockets recently. 2 weeks ago I had only a high-level understanding of sockets (learned from casually reading manpages, docs, blog posts, etc.). I decided to read as much as possible because I wanted to understand networking fundamentals, and after a week I learned enough to write some sockets code in Python and C. I know Python quite well, so reviewing the ``sockets'' library made more sense after my deep dive.
If you want to get better at technology A using language X, I suggest either reading/watching as much as you can about tech A, and build stuff with it in language Y. Then you can circle back to learning language X and you've already mastered much of the concepts around technology A.
e: spelling
A byte code interpret is a stack, some way to represent functions on that stack, and then a loop to interpret beach byte code and move the program counter.
I did have a bit of experience with VMs before, I wrote many years ago a short series of posts about it on my blog, and at my previous job I dabbled a bit in JVM byte code to solve one very unusual problem we had for a customer. I also read the _amazing_ https://craftinginterpreters.com/ years ago and that gave me some ideas.
But this project was definitely big and complex. It took me a lot of time, and it got abandoned a couple of times, like many of my side projects. But I'm happy I finished it. :-)
If it's zero (and no judgement from me if it is; plenty of other things to focus on), then it shouldn't be surprising that someone for whom that number is (speculatively) 10-20 hours per week on average for years has impressive side projects.
The magic will fade away quickly.
:-)
osdev.org, sandpile.org, RBIL, and freevga. The biggest PITA is hardware support. There are many good vintage hardcopy books with recipes for things like reliable port IO and undocumented hardware tricks.
- Intel® 64 and IA-32 Architectures Software Developer’s Manual Combined Volumes: 1, 2A, 2B, 2C, 2D, 3A, 3B, 3C, 3D, and 4
- Microsoft MS-DOS Programmer's Reference (also includes real-mode BIOS calls)
- PC Interrupts
- Undocumented PC
- PC Intern
- Programmer's Guide To The EGA, VGA, And Super VGA Cards
- Graphics Programming Black Book Special Edition
Also, it's worth toying with advances in OS dev past the era of monolithic, microkernel, and hybrid.
1. Capability-based like seL4. It has a number of inherent performance and security advantages including capabilities and excellent IPC.
2. POSIX compatibility layer. Even embedded OSes without the concept of threads or processes can implement POSIX.
3. Hypervisor. They're much easier to add with intel's VT-[xd]. Failing that, fall back to emulation. Translational emulation is very performant.
4. Get good at generalizing interrupt handlers, making them fast, avoiding race conditions, and using lock-free patterns.
Also:
5. Rewriting or trapping unsupported instructions including x87 and MMX.
6. The failure of pure microkernel was the added complexity and management of sequencing multiple resources in a transactional manner. There are great theoretical security and operational advantages in microkernel architectures but they never caught on widely in a pure form.
Do you have resources about this? I can't quite fathom how it would work, but then again I have no expertise.
> 3. Hypervisor. They're much easier to add with intel's VT-[xd]. Failing that, fall back to emulation. Translational emulation is very performant.
Speaking of which, QubesOS was on HN recently. The essence is having many VMs to minimize cross-app attack surface and privilege escalation.
> 6. The failure of pure microkernel was the added complexity and management of sequencing multiple resources in a transactional manner. There are great theoretical security and operational advantages in microkernel architectures but they never caught on widely in a pure form.
I recently saw an interesting brief article[0] about how memory-safe languages can supercede the compartmentalization that microkernels provide. I'm reminded of Theseus[1], written in rUsT, which happens to reflect the sentiment. I've actually been putting off rereading the Theseus USENIX paper. Again, I'm not nearly qualified to answer whether the security of this is comparable to the best of that of microkernels, barring formal verification. Still, I think it should be explored more.
[0] https://catern.com/microkernels.html (Write modules, not microkernels)
[1] https://www.usenix.org/conference/osdi20/presentation/boosOf course theres a crate for all that, but thats not the point of making an OS.
A lot of the foundation and ground-floor mechanics are pretty interesting though.
Started working on something very similar a few years back and gave up pretty soon for some stupid reason. Maybe I should try again, am getting better at getting stuff done.
I think of them as retirement projects before I retire. When I actually retire I'll maybe finish them.
(That said, I have in the past tried taking jobs that were adjacent to my "research" interests, and found the joy of building these things from scratch is much better than fiddling with the levers on the side of someone else's thing they built from scratch years ago. I like working on and improving production systems, but if they intersect too closely to my personal interests, it can be demoralizing.)
I know the feeling - this project, like most of my other side projects, got abandoned a couple of times. But I was really curious about implementing a GC and, for once, I managed to finish something. I'm glad I did! :-)
This is a typical way to present a project that ends up replacing the established/existing implementation.
Recent posts:
https://news.ycombinator.com/item?id=36735344 - 6 days ago
https://news.ycombinator.com/item?id=36717967 - 7 days ago
https://news.ycombinator.com/item?id=36710803 - 7 days ago (OP)
Btw. nice project!