My process for writing fast code on the JVM:
1) Measure. Set up a benchmark so you know whether you're on the happy path or not, or whether you've fallen off a performance cliff, for whatever reason. Make this part of your testing suite with some kind of notification for regressions.
2) Start small, ideally with a do-nothing loop over the input. This gives you a baseline; you can't go faster than a do-nothing loop (presuming you can't skip part of your input, which is an algorithm problem, not specific to JVM optimization).
3) Incrementally build up your algorithm and pay attention to when it falls off the performance cliff, using your benchmark from (1). If and when you fall off the performance cliff, that's when you start bringing in tricks like avoiding new, ensuring call sites are monomorphic / bimorphic, avoiding boxing, reducing pointer indirections and other cache friendly code, etc.
Another trick to consider is playing around with inlining, but not in the way you might think: try pushing infrequently executed code (conditionals) one level deeper in the call stack (i.e. making the body of an if-block into a method and calling it). The idea here is to encourage inlining of method doing the calling. Inlining is where the JVM gets to specialize your code to the specific situation at hand, but the JVM is reluctant to inline big methods because it has a time budget. So you need to help it.