I understand that java is very fast, faster than most people believe, but does it beat hand-optimized assembly? Fortran? C?
Why Java?
I understand that java is very fast, faster than most people believe, but does it beat hand-optimized assembly? Fortran? C?
Why Java?
When an HFT system makes money, some other person or system is loosing money i.e. it is a zero-sum game. And that person isn't going to sit idle; they are going to analyze their mistake and up their game. The industry parlance for this is, I think, that the "strategy dissipates into the market". This causes HFT systems to be always in active development.
Spotting a strategy and putting together a system to take advantage of it, and spotting where your system is losing money and fixing it - both these need to be done with quick turn around time.
Most strategies can be thought of as picking up nickels from in front of a on-coming train. Your best case "win" is a tiny amount of money; but your worst case "loss" is much much more money. The trick in HFT - the HF part - is to make these tiny wins over and over. But much actual money is involved, with very bad consequences should the program run into inopportune bugs.
All of these - constant development, time to market pressure, high reliability - give the impression that Java is a better language to code this HFT systems in than C or Fortran.
...Now I've never used Fortran myself.... but none of those arguments seem to stack up.
The majority of the fast paths have performance characteristics that Java is suitable for as long as it's actually executing your program. hence the focus on GC, networking and concurrency primitives. And mostly Java is chosen where the bulk of the rest of the codebase is Java.
Where you need ridiculously fast, then you have to drop down to what suits. Be that C, ASM or FPGAs. But these have costs wrt dev time and/or ease of establishing correctness. So they tend to be used judiciously.
Java could very well be the first programming language designed by adults and it shows. (That is, they solve the hard problems, not just pretend they don't matter)
That's going away as exchanges upgrade their gateways.
Take Android for instance, getting solid performance out of the JVM there is a huge bear because the language is just not suited to low latency operations. I won't claim that C is a perfect language but I hate it when I see people throwing around that you're committing some sin for each line of C code you write.
If you want cache friendly operations and predictable performance you're not going to find it in any JVM or language that has a GC.
For instance the subjective experience of using Android is that you can never close any app that you've opened other than by uninstalling it. Even after you turn your machine off and turn it back on you still see windows for every f--king app.
Then they run a bunch of articles about what an idiot you are if you try to close these because it won't save your phone's battery.
Well I admit I do have some cognitive limitations and it is --hard-- to scroll through 30 apps just to switch from (say) the web browser to the PDF viewer, but I guess Google thinks it is great this way because you always have a Google Plus window open.
Then there are all the articles about the fancy power management they are going to have someday that doesn't face up to the fact that an android device may or may not charge if you plug it into a charger, might turn itself off when it is running, and that the most reliable way to turn it on is to do a hard reset... And this isn't any old piece of junk, this is a nexus device.
FWIW Force Stop from the settings->app section will stop an app unless it's forced to be sticky because another service depends on it.
That's not entirely accurate. Many of the apps can and will stay in memory (it doesn't suspend all of them). Not that that's necessarily a bad thing.
But yeah, Android doesn't kill apps outright when you switch away, that would be a pretty poor experience to restart each time you opened a link in Twitter and came back from it for instance.
Oh I know I was just being specific since you said they were "just" screenshots. Wanted to make sure it was clear some of those apps may in fact still be in memory. ️
I love android but it feels like every iteration has plenty of huge bugs never really addressed. iOS always seems more reliable but far more limited with some of the things I want to do with a phone.
You should not close apps because you may interfere with tracking your location and other data collection.
Data collection is necessary to enrich your "user experience". When we know what you want we can fulfill your every wish! Just say "OK, Google". It will be great!
The developers behind this crap do not hear from satisfied users. Because the truth is no one really cares about this stuff. They care about things like reception and battery life...
except for some nerds like the ones who comment on HN who can easily point out all the stupidity of these "business" models.
When users are like puppets on a string, helplessly dependent. When there are no alternatives, no competition. Is that a business? And I suppose shooting fish in a barrel is game of skill.
Oddly enough, Googlers read HN comments and frequently defend the company, speaking only for themselves of course. Why should they care that anyone sees Android for what it really is? Whining nerds do not count, right? So why pay attention to what they think?
If developers of Android and iOS had any respect, if they had a conscience, then they would not be usurping people's computing resources for their own ends. Sure, users will be oblivious to what is going on and they will not complain. That does not mean it's OK to do these things.
(Edit: Probably the repeated pejorative usage of 'nerd', I guess.)
> You should not close apps because you may interfere with tracking your location and other data collection.
That's bullshit reason - you'd write a background service, not try to stop users from closing the app. For any single application I'm using I'm happy for it to be cached rather than restarting every time. This is just arguing against a cache layer, because <made up reason>.
> A started service must manage its own lifecycle. That is, the system does not stop or destroy the service unless it must recover system memory and the service continues to run after onStartCommand() returns.
A started service shouldn't be killed randomly - after all, that's how the downloads are handled - you don't see those disappearing without a reason.
The issue is not simply privacy, it's control of the hardware.
The real question, I guess, is "do you have the knack"?
Is there a reason the user should explicitly terminate an app rather than just stop using it? Why not have a GC button, too? "Closing" and "file system" are meaningless to anyone who hasn't got a grasp of implementation. Don't show that stuff to the user.
I don't know you and it's always dangerous to argue with people on HN due to who might be reading but I'd very strongly disagree with your above quote.
Predictable performance can be obtained on a very low latency, I've seen it:)
Cache friendly can be tougher and I won't argue the point as I think it will end up being a "no true Scotsman" argument where I claim Java can be cache efficient and you counter that C can do better so therefor Java isn't truly cache efficient
That said the amount of restrictions you have makes it really painful(needing to use pre-allocated wrappers if you want to read more than intrinsic types, not being able to grow/shrink anything). There's sun.misc.Unsafe, but that isn't available on every JVM implementation.
A GC is going to by nature introduce pauses if since at some point it's going to need to do work on the memory you hold. There's a set of tunable tradeoffs but it will impact latency.
I guess my meta point is if throughput and latency are paramount to you then you should really consider a language suited for it. For me today that would be Rust or C++ depending on how risk-adverse the organization you're working with is.
There's also talk about the possible removal of sun.misc.Unsafe in Java 9¹.
¹ http://blog.dripstat.com/removal-of-sun-misc-unsafe-a-disast...
The Oracle people are right about this one. Rip it out and sterilize the wound with fire.
The JVM is in HFT in because of sheer inertia; it has been worked on for so many years to shoehorn it in, and some smart people have made it work. The tooling is mature, and there are alternative syntaxes close to it like Kotlin that remove a lot of the noise for code review and maintenance.
The JVM is the reason I stayed away from Clojure, even though I like Lisp/Scheme syntax and the semantics. Where is Clojure-clr nowadays?
Umm... here
https://blog.twitch.tv/gos-march-to-low-latency-gc-a6fa96f06...
It sort of looks like language developers are finally getting the point that while you can easily deal with a throughput problem by just throwing faster processors, more processors, and or memory at at it, working around long gc pauses is just pain all around.
Take Android for instance
Android <> the JVM. There was a recent, fairly well publicised legal battle around this fact.Complete hyperbole. Nothing stops you from using a safe string library in C.
> Also until pretty recently, C and C++ did not have a sane memory model
Technically true, but in practice it was rarely relevant.
Nothing technically stops you, it's just really hard. When even djb has suffered a string-processing-related bug¹ you can't reasonably expect to get string processing right all the time. Everyone thinks they can safely handle strings in C. It's like programmer chunibyo² or something.
--
¹ http://article.gmane.org/gmane.network.djbdns/13864
² “8th-grader syndrome”; the sort of overly high opinions middle schoolers have of themselves
Do you really want a list of string-heavy s/w thats developed in C/C++ and in very widespread use?
Should I start with:
- Most interpreted languages...
- Web servers....
Correctness, maintenance, extensibility, developer productivity, tooling, deployment, and recruiting are all also important.
A common pattern in Java based systems is fan-out based processing. One immutable message is created, sent to multiple processors, who then combine signals to make a decision. Tracking the memory here would slow you down. But usually the JVM handles this just fine, and kills those objects in first generation.
Also, Java has a fantastic standard library, probably the best concurrency system, and the JIT can actually make code faster than C++.
It doesn't work for extreme low latency people - those guys write custom code for FPGA - but there is a healthy chunk of HFT for whom Java is a great choice.
How? Demonstrate with real examples
A JIT compiler can sometimes beat an AOT compiler because it has more information.
For example, it is entirely feasible for a JIT to heavily optimise a fast path even if the optimised code wouldn't be correct for all cases that the source could be called for. If the JIT detects an uncommon case it can just fall back to the interpreted code.
An AOT compiler will forego optimisations if it can't be sure that it will produce correct code. For example, C++ was generally considered slower than Fortran until restrict was added as the compiler had to be more conservative. However, restrict is, well, restrictive. Conceptually, a JIT could work around this if it had a function that didn't have restrict arguments but was usually called on distinct memory. It could hold an optimised path that assumed restrict but fall back if it detected otherwise.
Now, some of this benefit can be had in an AOT compiler with profile guided optimisations. But usually, the AOT compiler will still tend to the conservative to balance aggression with code bloat.
I also think that the JVM is the only JIT that has had close to same number of resources pushed at it as some of the Fortran/C/C++ compilers. But to differing markets the JVM focused a lot less on easy to benchmark numeric and actually looks at other types of code.
Then of course there is the hybrid JVM/LLVM using Graal/Sulong [1]. Which hopes to better than JIT or AOT apart.
ParentClass g = ....might be a subclass...
for (int i=0;i<100000;i++) {
g.func(i);
}
The JIT can optimize away the vtable lookups to find func, and sometimes inline the code. Ok, maybe you could do this in C++ too. ParentClass[] g = ....might be a subclass...
for (int i=0;i<100000;i++) {
g[i].func(i);
}
Suppose 99% of g's are the same class. The JIT can optimize away most of the virtual function lookups (particularly if this is one of your hotspots). I.e. the code becomes: ParentClass[] g = ....might be a subclass...
for (int i=0;i<100000;i++) {
if (g.class == COMMONCLASS) {
inlined_func(i);
else {
g[i].func(i);
}
}
In the common case, a simple pointer comparison + inlined code is a lot faster than a function dereferene.While C++ code does use virtuals, it's nowhere near the amount as Java - there are language constructs to avoid that and move the dispatch selection to compile time.
This is perhaps not fully "ahead of time", granted, but it's extremely easy to deploy and highly effective, and entirely accessible to C++.
The second is more dubious... sure, a JIT compiler has information if it spends RAM and cycles on collecting that, but it also has to run quickly and fit in the runtime environment, and an AOT compiler can run arbitrarily slowly, use a whole rackful of servers, and can use PGO without incurring any profile collection costs at (normal) runtime.
This is nontrivial because lots of optimizations depend on class load ordering and runtime profile information.
Meh. A lot of systems use specialized memory pools for such use cases: - allocate memory pages from OS - allocate objects via this pool - release the full pool when done (i.e. just unmap t page, which essentially has no cost at all)
If the task at hand is known to have an upper bound of same sized objects, this basically reduces management overhead to maintaining a single pointer (and you can work with guard pages to just catch the segfault when trying to access out-of-bound memory, which is not that far fetched because it is actually what some JIT'd Java code does to optimize out null checks [0]).
There's also the issue that even without GC running you pay the cost of card marking (GC store barriers) on every reference write. There's unpredictability due to deoptimizations occurring due to type/branch profile changes, safepoints occurring due to housekeeping, etc.
It's unclear whether that style of Java coding is actually a net win over using languages with better performance model.
As I said, the JVM is an acceptable platform for the slower HFT. That's the kind where a clever predictive strategy matters (maybe with lead time of seconds) and you'll get more money from accurately predicting the future than from shaving off 250us.
Make no mistake - you'll still make money shaving off 250us, but not so much that you want to be bogged down structuring your code the C++ "if we structure it right we won't leak things" way.
I know you were throwing 250us out there as a pseudo example, but that's actually a very long time even outside of UHFT/MM.
Also don't forget that your trading daemons will be under a fire hose consuming marketdata, so beyond being able to tick-to-trade quickly, you need to be able to consume that stream without building up a substantial backlog (or worse, OOM or enter permanent gapping).
If you make sure that "almost all" allocations are short-lived, GC is very fast. Allocation is bumping a pointer and cleanup is O(number of new, live objects). It's considerably faster than malloc/free for general-case allocation.
Also don't forget that when GC runs it trashes your d/i-caches; temporaries/garbage allocs reduce your d-cache efficacy; GC must suspend and resume the java threads, which is trips to the kernel scheduler; there are some pathologies with Java threads reaching/detecting safepoints.
GC store barriers (aka card marking) don't have anything to do with thread contention (apart from one thing, which I'll note later). This is a commonly used technique to record old->young gen references, and serves as a way to reduce the set of roots when doing young GC only (i.e. you don't need to scan the entire heap). So this isn't about thread contention, per say -- with the exception that you can get false sharing due to an implementation detail, such as in Oracle's Hotspot.
The card table is an array of bytes. Each heap object's address can be mapped into a byte in this array. Whenever a reference is assigned, Hotspot needs to mark the byte covering the referrer as dirty. The false sharing comes about when different threads end up executing stores where the objects requiring a mark end up mapping to bytes that are on the same cacheline - fairly nasty if you hit this problem as it's completely opaque. So Hotspot has a XX:+UseCondCardMark flag that tries to neutralize this by first checking if the card is already dirty, and if so, skips the mark; as you can imagine, this inserts an extra compare and branch into the existing card marking code - no free lunch.
"Performance-critical code" can even go in that space in an environment where developer cycles and program safety are things that matter, which is definitely the case in HFT.
Also, what's an (non-toy) environment where developer productivity and safety/correctness don't matter? I always find that statement bizarre when talking about production systems.
I went wow the first time I edited a program during a debugging session and Eclipse just recompiled the class and reloaded it into the program.
Yes, there details of getting this right (transferring the state), but nevertheless.. It'd be one heck of a job implementing this _correctly_ in a MT scenario for C and C++. JVM just does "magic" here.
I think every major debugger can do this for C and C++ now
They could probably do a lot better with, say, OCaml or Rust. Jane Street uses OCaml for, presumably, this reason.