Raw String Literals Removed From Java 12 as Feature Set Frozen
infoq.com
infoq.com
Unicode escapes are processed in Java not just inside string literals, but everywhere in the source code. So, for example, the following program prints "Hello, world!" even though that line of code seems to be commented (\u000a is new line, so it ends the comment):
public class Test {
public static void main(String[] args) {
// \u000a System.out.println("Hello, world!");
}
}
Moreover, a \u000a inside a string literal is the same as an actual newline, so the compiler doesn't accept it: Test2.java:3: error: unclosed string literal
String s = "\u000a";
^
But now with the raw string literal proposal, JEP 326[1], Unicode escape processing is disabled inside raw string literals, and \u0060 escapes (backticks) aren't considered backticks for the purposes of starting raw string literals.So, with this proposal, Unicode escapes are in a worst-of-two-worlds middle way:
1) They can't be handled uniformly at a low level anymore, so a Java parser can't naively convert escapes while reading the source file, but
2) They must still be naively interpreted in unexpected places like comments and normal string literals, as shown in my examples.
What a mess.
Why did they do it like this
Just why
// \u000a no commenthttps://en.wikipedia.org/wiki/Digraphs_and_trigraphs
For example C has escape sequences for characters as basic as # and [.
For example, in C, I can write:
int a<:10:>;
Or I can write: int a[10];
They are the same, because <: is exactly the same as [, and :> is exactly the same as ] (except… mercifully, digraphs are not interpreted in strings; trigraphs are another story). However, in a string, the escape sequence \" means something different from ". One is a character in the string, the other ends the string.Digraphs and trigraphs exist because not all keyboards have the corresponding characters, like [ and ]. For example, look at a Swedish keyboard. (You can type [ and ], but not as easily, and maybe not at all on some older keyboards.)
But those in C++ are deprecated, yay!
This is the most unusual part and just leaves me asking why!?!? Was it an error of omission, or a deliberate design decision, to essentially preprocess the whole file blindly with Unicode escape parsing instead of putting it only inside the string literal (and character constant) parser like just about every other language that has a similar escape system? The fact that \n behaves differently (and in the sane manner that those are accustomed to from C/C++) is the most bizarre part --- Unicode escaped could've been handled by the same parser, but they didn't.
Sibling comments mention trigraphs. Those are different, starting with ?? instead of \, so their "applicable to the whole file" behaviour is more reasonable. But \-escapes are, outside of preprocessor line continuations (where it's literally \ followed by a newline), customarily only interpreted inside character and string constants.
> The Java programming language specifies a standard way of transforming a program written in Unicode into ASCII that changes a program into a form that can be processed by ASCII-based tools. The transformation involves converting any Unicode escapes in the source text of the program to ASCII by adding an extra u - for example, \uxxxx becomes \uuxxxx - while simultaneously converting non-ASCII characters in the source text to Unicode escapes containing a single u each.
The Java language grammar is defined in terms of Unicode, but it was designed in the early days before Unicode was ubiquitous. In particular, this was long before UTF-8 took its current place as the de facto standard character encoding.
The designers of Java wanted to support arbitrary Unicode characters in source code -- not just in string literals, but in identifiers as well. But they also wanted to preserve interoperability between systems using different character encodings.
It's certainly confusing that \u behaves differently from other backslash escape sequences, but I guess they figured that was less bad than introducing an entirely new escape character.
[1]: https://docs.oracle.com/javase/specs/jls/se11/html/jls-3.htm...
In my experience (although not with Java; maybe its users would make more use of the feature, but I doubt it) even developers whose native language is not English but something very different like Chinese or Japanese will continue to use ASCII-only identifiers, and only write things like string constants or comments in their native language (which may mean using high bytes and a non-UTF8 multibyte encoding.) What the identifiers mean may be Romanised foreign words, but they're still written with [A-Za-z0-9_].
At the time, the Java team really didn't have any way of knowing for sure whether or not non-ASCII source code would catch on, because most existing languages didn't support it.
That's bad! I recently noticed "my" C++ compiler allows you to write something like int mäin(){}. I didn't try it out but I guess it breaks with different file encodings.
Please stay with ASCII!
int numLetters = switch (day) {
case MONDAY, FRIDAY, SUNDAY -> 6;
case TUESDAY -> 7;
case THURSDAY, SATURDAY -> 8;
case WEDNESDAY -> 9;
};
Switches can return values, and can have multiple cases on the same line.I'm sure in 2 years when my organization gets around to switching to Java 12, I'll enjoy using it.
Interesting to see the faster pace releases to get feedback on features, but many major libraries are just getting Java 11 support.
Post Java 11, I doubt you will see the same problems you did with going from 8 to 9 or 11
The cases of a switch expression must be exhaustive; for any possible value there must be a matching switch label. In practice this normally means simply that a default clause is required; however, in the case of an enum switch expression that covers all known cases (and eventually, switch expressions over sealed types), a default clause can be inserted by the compiler that indicates that the enum definition has changed between compile-time and runtime. (This is what developers do by hand today, but having the compiler insert it is both less intrusive and likely to have a more descriptive error message than the ones written by hand.)
This is pretty much the behavior you'd expect if you're familiar with languages that support pattern-matching.
It’s also the first part of a much larger feature adding pattern matching to the Java language. Getting all the ducks in a row for this sort of thing is tricky because it touches so many other ares, but switch expressions can be delivered on their own quite nicely.
Adding patterns means looking at variable scopes (where is a variable bound in a pattern in scope?), record classes (because destructuring patterns should not be implemented by hand everywhere), and even constants and raw string literals (because being able to switch on regexps would be nice, and you need a good way to write that).
This is not an exaggeration. I was writing code in Java to access SQL databases in 1999 and needed it to express long SQL strings.
One would think they could move a little faster.
I like my SQL right next to the code that sets parameters, processes results, etc - mismatches are apparent (and also validated by IntelliJ).
If that’s right, it’s exactly how the new release cadence is supposed to work—there’s no longer such a hard decision between delaying the release, leaving a feature to languish for several years, or releasing a feature without time to validate it.
- import com.package.foo.Bar -> use Bar.methodOfBar() in code
- import static com.package.foo.Bar -> use methodOfBar() in code
What does "import com.package.foo.bar as bar" add? Other than some kind of aliasing, as in you import Bar, but alias it as Baz, so then you use Baz.methodOfBar() ..?
import java.util as jvutil;
import com.betterlist as better;
class Example {
jvutil.List list = new better.List();
}Not really. VS Code transparently handles aliasing of imports in JS. You still never need to look at them manually AND you get the benefit of cleaner naming.
As with all code readability, I assume to write and use import aliases, that's very easy. The person who's gonna read it after you, they've got a problem. And that's why I'd personally be strongly against this in my code review.
Is that java.awt.List or java.util.List?
import java.awt.List as awtList
import java.util.List as utilList
Granted, there's not a lot of scenarios where you have this sort of collision, but it does happen and using aliasing makes the code easier to read.
import com.package.foo.bar.Something as barSomething
according to the documentation:
https://kotlinlang.org/docs/reference/packages.htmlAlthough you'd have to do it for each class, so importing the whole package with an alias would be nice.
- https://openjdk.java.net/projects/jdk9/
- https://openjdk.java.net/projects/jdk/10/
At the very least, have a default, optional string substitution feature. It ain't that hard and it's something users of raw strings will NEED anyways. What good is it to have snippets of raw strings that then need to be parsed again and split and re-glued again?
JavaScript backticks are much more useful out of the box. But that isn't the Java way... In Java, everyone gets to write their own string substitution library, littering the heap with completely unnecessary substrings.
After the disasters with generics and enums, do it right, for once, please?
Are you referring to Clojure or Ceylon? Or both, or neither?
Pattern matching is such a powerful feature. I wish more languages (like python) would realize it’s value and add it.
I've used Scala before.
There are some cases, like compilers where I could see matching and extracting complex and nested patterns to be useful.
But in most cases monadic operations/functors plus if/else seems pretty much as good.
I probably use monadic operations (map, fold, etc) at least 5-10 times a day in my job (I’m a scala dev), and match expressions at least 3-5 times a day.
So pretty useful I’d say. Though then again I’m dealing with some pretty baroque case classes that contain dozens of values of financial information.
Scalaz and cats are perhaps the most common dependencies in the Scala ecosystem.
https://cr.openjdk.java.net/~briangoetz/amber/pattern-match....