As pointed out by other comments, this concept of a "meta-circular VM" isn't new. It's been done before by two other projects, Maxine and Jikes. What's different about SubstrateVM is that this is a production tool rather than a research project, and it's not just a JVM, it's also a way to pre-initialize the app. Therefore programs compiled with native-image can start as fast as programs written in C. Actually, slightly faster in some cases. You may wonder how that's possible given that Java apps normally start slowly, but it's because there's no JIT compilation and the state of the heap is snapshotted, with classes pre-initialized including the JVM itself. So the program can literally just start executing at main() in machine code with no VM startup overhead, because it's done already.
The downside is that snapshotting and AOT consume a lot of disk space.
There are some questions below asking how this works. It sounds initially "impossible", like a lot of stuff GraalVM/Truffle does, but it's quite easy to understand really.
You start with a bytecode compiler written in Java. This is a normal program written in the normal way, because a compiler is ultimately just a function that converts one stream of bytes to another. Then you write the runtime and GC in Java too, and compile that as well. This code is a bit special. It's still syntactically Java, but, some classes and methods are given special meanings and some extra rules apply. They aren't compiled in the same way as normal Java code. For example you can write code like this:
UnsignedWord value = Pointer.readUnsignedWord(address)
This doesn't allocate an object or call a static method. Instead it will be compiled down to a single mov instruction. Likewise for writing to memory - there are magic methods that are taken to mean "emit this assembly" instead of doing normal method calls.Several other tricks are required. GC code can't allocate because it would mess up the heap it's working with, so you can use annotations to mark methods as "never access the heap". But then, GC code is written in Java and Java must allocate for almost anything non trivial, so how does that work? The answer is, the GC code is initialized at build time and all the objects it needs are snapshotted into the default heap that's mapped into memory at startup.
There are lots of other tricks, mostly annotations that control the compiler so that e.g. methods are guaranteed to be inlined and removed, objects are guaranteed to be stack allocated. This isn't available to normal Java but when you control the compiler it's not a problem. The advantage of this Java-superset (or subset) is that you can use all the normal tools that understand source code, like IntelliJ, JavaDoc etc.