You’re totally right that benchmarking and profiling is hard even for native code. I think this post fetishizes whether or not a piece of code got JITed a little too much. Maybe the author had a bad time with microbenchmarks. There’s this anti pattern in the JS world to extract a small code sample into a loop and see how fast it goes - something that C perf hackers usually know not to do. That tactic proves especially misleading in a JIT since JITs salivate at the sight of loops.
I know that compilers are smart, but the API doesn't make things easier at all
SomeClass x = new SomeClass();
won't make anything slower (ignoring auto for now)
Now, I read Java code and I can't help but wonder how things are harder
For example, setters and getters, what would be a simple memory write becomes a function call. Not complaining when you actually need it.
Reading the examples here (and those are not too bad) https://developer.android.com/reference/java/net/HttpURLConn... it seems you have to actually fight the API to get anything done
Why do you need to cast the return of a url.openConnection() to a HttpURLConnection? (I mean, how many connection types exist?)
Two? https://developer.android.com/reference/java/net/URLConnecti...
https://docs.oracle.com/en/java/javase/11/docs/api/java.net....
In general having some higher level, well optimized helpers can certainly reduce verbosity and increase speed. That said, some types of verbosity just make the programmer write what the compiler would translate to anyway. Or can end up being unnecessary - e.g. a strongly typed language with no inference certain slows the programmer with no effect on the program (though maybe an effect on the type checker).
rep movsb
will do the trick. This assumes that the source register points to source, destination register points to destination and counter register is set to number of bytes to copy.- A verbose SIMD copy loop is often faster even on modern Intel CPUs that have the more modern rep implementation.
- A simple byte copy loop will beat both SIMD and rep when you're copying smaller amounts at a time.
Still, the JVM JIT is faster than JS due to reasons explained in the sibling comment of the parent one.
You're thinking about generics. .class files preserve a whole bunch of type information (I'm building a .class decompiler in my free time, and I'm looking at that very same data in my debugger ATM).
Even generic types are available in the Java .class file and are accessible from the reflection API. Spring for example uses this quite heavily.
Type information is present in fields and class inheritance
For example, a class like this
`class Foo implements Bar<String>`
Retains the fact that the generic type is a String.
That information is completely lost at method invocation. So a method that takes a `Bar<String>` ultimately compiles to a method that takes a `Bar` and knows nothing of the String.
To get that generic information down you have to engage in some fun tricks using either the class or field method I mentioned earlier. (Usually you do this with a second type parameter where it matters).
In JS you can only get the types by profiling.
Also it’s not really true that genetics are erased. There’s that horrible thing javac does so that the VM can support reflection for generics. I get your point though.
Overall I think C# got this one right. Generics are right there in the binaries.
bUt ActuALLy, I think the rightest right thing is to do specialization up to representation at link time (or compile time), when the whole program is available, ala MLton. Virgil does this. Of course this is not possible in a dynamic code loading environment but I only have so many fucks to give in this life.
I agree C# got this right.
I agree that doing specialization up to link time is ideal from a certain standpoint.
e.g.
class A<Y extends X> {
// in this scope, Y is known to be of at least type X,
// so, we can call methods on expressions of type Y
// that belong to type X (and not Object)
Y m() { ... }
}
Type arguments are omitted from usage sites. The technique is literally called erasure in papers and documentation. a = new A<Foo>();
f = a.m(); // should return a Foo, in bytecode returns Y
// and compiler inserts a cast from Y -> Foo
Generic code is slower in Java because of these extra casts. To get back (most of) the performance, the JVM has to inline enough methods to be able to track the types from start to finish. It can't always.I don't know that the java JIT has type feedback. However regardless Java being statically typed means Java code is structurally closer to what a JIT wants and can analyse, it's much more difficult (and thus rarer) to fuck around with types and objects generated on the fly at runtime for instance, you're not going to add new attributes to an instance whereas that's just tuesday in javascript.
To whet your appetite, what happens if you redefine Math.round? Java prohibits this for obvious reasons, but a JS joker may write:
Math.round = () => Infinity;
JS engines really do inline Math.round, and must be prepared for this nonsense.It gets worse. Maybe you check for a property which is usually not found:
if (!val.uuid)
val.uuid = uuidgen();
Hidden classes can almost reduce this to a pointer comparison. But what if someone adds a uuid property on Object.prototype? Every check is busted! v8 handles this with "validity cells", and it's ugly, requiring that every object know about every object which "inherits" from it.Now if you are a monster you may choose to write:
Object.defineProperty(Array.prototype, "42", {value: "lol"});
console.log([][42]); // yes it prints "lol"
Every array gets a default value for 42. Think about how you would JIT numeric code in such an environment...There's also the "frame evaluation API" PEP [1], whose purpose is to allow pluggable evaluators in CPython without forking the entire interpreter, like Unladen Swallow had to.
Not as straightforward as with JS, but Java JIT still have to account for those things.
May be it will assume some things about standard classes for optimization, but for user classes that will be true anyway.
Yeah, that is what they meant, but it is a little misleading. Javascript's performance is usually within a single digit multiple of java's, whereas python is often significantly slower. Javascript is somewhere between Java and an abacus too.
(Broken ecosystem = lots of things don’t work with lots of other things; package managers are another example; Mypy another)
E.g. pg8000 looks like it's being maintained?
Similarly, Postgres doesn't work for Pypy unless you're willing to use libraries that are unmaintained or only very lightly maintained, and even then who knows what the performance tradeoffs will be. Those are pretty substantial caveats for a production use case.
I haven't used Poetry, but I've heard good things. I'm very cautious though, because I've also heard high praise about pip/virtualenv, pip-tools, pipenv, etc and every single one had spectacular problems on the happy path (30 minutes to resolve dependencies with pipenv is my favorite).
Sure that's true, but performance is about measurement.
Poetry has some missing pieces still (I can list them if you're interested but I'm on my phone) but it's the first time I've felt ok recommending a packaging tool to beginners. The experience is way better than the alternatives like pipenv.
I have a very low degree of confidence that a pure Python solution (even JIT optimized) is going to compete with a native solution, not to mention that by all appearances it isn't sufficiently well-supported to meet our criteria. Measuring the performance for me feels like a waste of time; it's easier just to forego Pypy until the ecosystem matures.
> Poetry has some missing pieces still (I can list them if you're interested but I'm on my phone) but it's the first time I've felt ok recommending a packaging tool to beginners. The experience is way better than the alternatives like pipenv.
This is good to know. I'll have to give it a try.
You are correct that JS is not between Python and Java. Python is faster than Java which is faster than JS. Though some people seem to think calling APIs written in C is "not Python" but if the ecosystem provides the library and I call it from Python then it's Python enough for me!
Get used to this approach since it crops up everywhere. If you target a gpu, you try to keep the work in the gpu. When we move to use bpf and io_uring things will be a race to the next kernel program. If you target the CPU then your code is a race to C or a race to the function that calls the relevant cpu instruction that will make your code fly.