C# Raw String Literal Proposal
github.com
github.com
This is the same rule that Oil has; I think it came from the Julia language (or at least that's where I got it from)
Oil Has Multi-line Commands and String Literals http://www.oilshell.org/blog/2021/09/multiline.html
var xml = """<element attr="content">
""" <body>
""" </body>
"""</element>
""";
This would give you the string: <element attr="content">
<body>
</body>
</element>(no newline)
If you wanted a newline at the end, you'd do this: var xml = """<element attr="content">
""" <body>
""" </body>
"""</element>
"""
""";
Basically the end delimiter of the string would be the last """. You could concatenate two strings like so: var xml = """<element attr="content">
""" <body>
""" </body>
"""</element>
"""
""" // this string ended on this line
+
"""<element attr="content">
""" <body>
""" </body>
"""</element>
"""
"""; // this string ended on this line
This could use the same logic for using at least three quotes as the indicator that it's a multiline string.Please, tear this apart and offer improvements.
Edit: this is conceptually similar to Zig's multiline literal: https://ziglang.org/documentation/master/#Multiline-String-L...
Edit: To be specific, the IDE could handle formatting when pasting into a line beginning with """. Or offer a "paste as cool new multiline string syntax" feature.
But off the top of my head, mainly just that there's a clear visual indicator of the start of lines of text, rather than counting/lining up leading whitespace. In the first example, the strings are all tabbed evenly for the sake of looking "pretty" in the code, but the following would generate the same string, since each line begins after the """:
var xml = """<element attr="content">
""" <body>
""" </body>
"""</element>
""";A few points.
> but the following would generate the same string, since each line begins after the """:
That's not a virtue here. The point is to be able to write clear literals that never need escapes and which allow for easy grokking of what the content actually is.
All current string forms in C# require some amount of manual (or tooling) help to fix them up to be legal. That's not the case with this literal. The content can always work as-is without having to touch it at all.
In my syntax, the IDE would ideally treat the """ block virtually like a <textarea>.
Note: this feature is entirely optional. You can absolutely not have leading whitespace trimming at all. Indeed, this is a requirement of the proposal as we have to make it possible to actually represent text that has leading whitespace :)
https://www.eclipse.org/xtend/documentation/203_xtend_expres...
It has this idea called whitespace preprocessing, it has several rules but one of them guarantees that your first 2 examples would work as you intend them to.
But also how common the theme is that once people use it a bit they are very pleasantly surprised.
This seems particularly endemic in the javascript crowd and its almost as if they have been brainwashed by too many kool aid blogs and don't have broad experience of having tried other technologies, but yet they are so self-assured.
I really don't understand why C# isn't used so much more widely in startups, and can only think it's a fashion and misplaced virtue signaling? Genuinely interested in opinions on this, what do you think?
I've built 3 adtech platforms with .NET and the productivity and performance has let us outcompete much bigger companies with a smaller team so it's fortunate for those who do know about it.
var xml = """
<element attr="content">
<body>
</body>
</element>
""";
And then they say that xml gets this... <element attr="content">
<body>
</body>
</element>
But they don't explicitly say if the new lines after and before the """ 's are considered part of the literal string or not.Are they?
[1]: https://cr.openjdk.java.net/~jlaskey/Strings/TextBlocksGuide...
[2]: https://cr.openjdk.java.net/~jlaskey/Strings/TextBlocksGuide...
jshell> var s = """
...> hello world
...> """;
s ==> "hello world\n"
You have to escape it: jshell> var s = """
...> hello world\
...> """;
s ==> "hello world"Later on,
> In the case of multi_line_raw_string_literal the initial whitespace* new_line and the final new_line whitespace* is not part of the value of the string.
``` multi_line_raw_string_literal : raw_string_literal_delimiter whitespace* new_line (raw_content | new_line)* new_line whitespace* raw_string_literal_delimiter ; ```
Which I think says that the opening and closing new lines (after and before the """'s) are NOT part of the content of the string literal, but new lines between them can be part of the string literal.
Another comment here suggested just make a trivial way to reference an embedded text file resource, and that is already very possible and not hard already, as well as use of string Resources.
.NET does allow you embed files directly into your project and read in those files but it's a lot of confusing boilerplate.
If .NET provided a really easy way to take a file from the project and create a compile-time string from it then I think that would be significantly more useful than this proposal.
C# is beginning to suffer from over-complexity at the language layer, and this solution of making it trivial to reference embedded text files avoid adding more complexity.
That said, it's actually pretty easy to reference embedded text files already, not trivial but easy enough.
I think "a lot" is a bit of overstatement considering the amount of power you have here. My implementations around embedded resources usually look like this same set of 5 lines:
var assembly = Assembly.GetExecutingAssembly();
using var stream = assembly.GetManifestResourceStream("My.Embedded.Resource.png");
using var memoryStream = new MemoryStream();
stream.CopyTo(memoryStream);
var myPngBytes = memoryStream.ToArray();There are so many different ways to do strings in C#. Adding features like this just makes the language harder to learn, and the compiler harder to implement.
At this point, it's probably better to adjust the compiler to make it easier to turn a text file into a hardcoded string. The embedded resource approach works, but it could be significantly smoother.
Or, maybe the compiler needs some form of a plugin architecture so people who want obscure features can figure out how to add them?
To your question of "why?", we tried to cover the reasoning in teh proposal. But, the core reason is that today people do use strings a ton. And in many cases it's unpleasant to do so because you always end up with reasons that you need to escape the content. This escaping serves to satisfy the compiler, but really doesn't buy value to teh user the majority of the time. The idea here is that you can just use a raw-string and say: here's the content, exactly as i want it.
var v = """"""
contents"""""
""""""
lol.I like the proposal to have these sort of raw strings, where indentations are removed, but can't they use a symbol before the string like they do with interpolation `$` or literals `@`?
I know it says design decision to go with 1 more " than the longest sequence of " in the string, but why ?
I'd also much prefer some kind of prefix, maybe double "at", e.g.
``` var myString = @@"blah blah blah blah " ```
This feels a lot more natural to me.
In practice, using """ will be sufficient almost all the time.
> Provide a mechanism that will allow all string values to be provided by the user without the need for any escape-sequences whatsoever.
###"There can be "stuff without escapes" #" "# "###
C# already uses @ for raw string prefixes so they could extend it's usage for multiple prefixes+suffixes.""" Is fine though.
R"SQL(my string without the sequence S Q L goes here)SQL"
You can use any extra delimiter you want. The concatenation rules make it easy for you to easily insert source code line breaks and indentation without literal string breaks or indentation.We looked into this. However, there didn't seem to be any benefit to this above just the N-quote version (which fits into how C# does strings everywhere else). In the above case, the `SQL(` and `)SQL` tokens are just akin to N-quotes. Since there's no additional benefit, we went with the simpler approach that solves all these needs, but will look the same across all codebases.
Hi! I'm the language designer here :)
I know it says design decision to go with 1 more " than the longest sequence of " in the string, but why ?
Because if we use a symbol before the string, then there needs to be some mechanism to escape it within the string. e.g. if you use `@` literals, you still need to escape quotes within the string literal. The point of this feature (which we try to spell out in the spec) is so that you can have content without the need to escape anything at all.
var s = """This is a
multiline string""";
Requiring the start and especially end quotes to be on a separate line makes it take a lot of vertical space. But OTOH, that is consistent with the default coding style in C# which is vertically verbose (with {} on lines by themselves).Because it is already in the language.
var xml = @"
<element attr=""content"">
<body>
</body>
</element>";
This proposal mostly seems to be about some edge case where @" " syntax isn't good enough. But really, this whole thing is an improvement to an anti-pattern, and you should instead be looking into not needing multi-line block specific string literals in your code (e.g. putting templates in their own files/resources). var xml = @"
<element attr=""content"">
<body>
</body>
</element>";If you want to store XML literals, then by all means do so, but within the code itself is inappropriate. Even the existing @" " syntax is a code-smell, the new syntax doesn't address why that is (e.g. validation/colorization/etc don't work for string literals containing arbitrary other languages).
.Net already has constructs to allow the dynamic creation of XML blocks (and JSON) without resorting to string-comcat shenanigans.
Times have changed, xml/web is becoming ubitiqos... espically in LOB apps where u can use WebView based frameworks to create cross-platform apps that share 99% of code.
There's plenty of stuff in a middle ground where a separate template file is a waste. Your example, now that you've corrected it shows that nicely. A separate template would be a waste here for these few bytes, and yet "verbatim strings" mean instead of this just being some actual XML you can copy-paste it has to be escaped / unescaped.
If your issue is that you don't think string literals should be a thing at all, C# is the wrong language for you. Try one of the early numeric langauges, or something modern like WUFFS that eschews strings entirely because they're too dangerous. Once you accept that literals should be a thing (notice these aren't interpolated, they're just literals) this is an obvious idea.
The bad "verbatim" syntax should go away in favour of a raw literal syntax such as the one proposed here.
```c# var s = """xml <Book><title/></Book> """; ```
and the like. Thanks!
var s = xml"""<Book><title /></Book>"""; $"A{new List<string>{$"B{"{C}"}D"}.First()}E"
... was valid C#.This means that a C# compiler can't start with a simple tokenizing loop. That compiler phase would have to keep track of state in a stack, recording what each } character means while its still looping through code character-by-character.
Now we're adding {{ and }} into the equation. Yay.
Not true, it just means that the parts between double quotes aren't bunched into a single token.
To figure out how the compiler makes sense of this code, try
https://roslynquoter.azurewebsites.net/
You'll see how it tokenizes the string.
A simple loop that would have worked with the 70s era of programming languages, would go through each character and once the boundary between two tokens has been identified, write out a token to a one-dimensional list. This would be a mostly stateless loop, tracking enough state for the current token in hand only. The next phase would go through the tokens and pair up brackets, etc.
A C# tokenizer can't do that. It needs to keep a stack of state. When it sees a '}', it needs to know if that's a "normal" brace or the } that resumes a interpolated string literal.
I was writing a tokenizer myself and I wanted to have something similar to string interpolation. I very quickly realized my simple loop that I would have written for my CS degree isn't going to cut it and I had to start over.
C# has never had a "simple tokenizer". Indeed, even the first language has complex lexical constructs that are part and parcel of the language. For example, our comments can store structured data in them (like xml).
> A simple loop that would have worked with the 70s era of programming languages
Yes. But 70s era compilers had to deal with things like not having enough memory to even store basic amounts of data. It also had to work in spaces where things like a 'stack' was just not tenable. We're literally 50 years from that point, and having a compiler do stuff like keeping a stack is not an issue anymore :)
Although you did deny me an excuse to not have to start over. If the C# team didn't bother, why should I? (Except you did bother.)
Hi, i'm the lang designer and feature implementor here :)
The complexity of lexing/parsing did not get worse here. We actually just lex/parse this stuff the same way that interpolated strings have always been lexed/parsed. This has been supported in the language for almost 10 years at this point :)
> Any line breaks within verbatim string literals are part of the resulting string. If the exact characters used to form line breaks are semantically relevant to an application, any tools that translate line breaks in source code to different formats (between "\n" and "\r\n", for example) will change application behavior.
https://docs.microsoft.com/en-us/dotnet/csharp/language-refe...
We absolutely do not normalize newlines as that would defeat the purpose of raw literals. The point here is that your content is not interpreted as that's the pain area that people are hitting today. How you write your literal is what you get at the end of the day.
Note: if the content needs to be `\n` then just use that actual newline in teh code. WRT to the file line endings and whatnot, my recommendation is that you never use tools that arbitrarily change that behind your back as it does already have impact today in C#. For example, that will break standard `@""` strings today.
If your line endings are important, then your tools should be setup to respect what you wrote and not change them. All editors can be setup this way, as can git. And that would absolutely be my recommendation on how you should structure things for your code if newlines are relevant.
let string = "foo\
bar";
Results in `string` having the contents "foobar". That would allow the user to specify the exact newline they need without depending on invisible characters remaining unchanged. "*n [stringA] "*n
"*n [stringB] '*n
'*n [stringC] "*n
'*n [stringD] '*n
Each string may contain runs of contiguous single or double quotes of length less than n, and furthermore:stringA may start and/or end with a single quote
stringB may start with a single quote and/or end with a double quote
stringC may start with a double quote and/or end with a single quote
stringD may start and/or end with a double quote
"""""""
""""""""
Are those strings containing `"` and `""` or are they empty strings? Is the first case an error because the starting and ending quote counts do not match? If the number of quote chars is even, do the contents alternate between `"` and empty as the number of surrounding quotes increases?> A single_line_raw_string_literal cannot represent a string value that starts or ends with a quote (") though an augmentation to this proposal is provided in the Drawbacks section that shows how that could be supported.
so I'd assume the odd count would lead to an error (single trailing "?) rather than a string containing ".
E: I'd assume it doesn't allow empty single line strings because otherwise how do you tell the difference between that and the start of a multi line one?
string longString =
`This
` allows
` differentiation
` of
`formatting
` indention
` from
` leading
` string
` spaces
` using
` back-ticks (\`)That would violate a core goal of the feature which is that the content itself doesn't need escaping. This sort of approach would require all users to have tooling that would make that pleasant, instead of providing a feature that was easy to use across any editor.
Thanks!
Other than that, ++ for any mechanism to quote to arbitrary depth. I have imagined
[abcfoo[ ...anything but ]abcfoo]... ]abcfoo]
as another approach.
But I don't like the indented form, where nested triple-quotes are ignored. Whitespace formatting is fine when I'm working with Python, but I really don't want to mix that paradigm when I'm working with curly-brace languages.
Maybe a source generator would be applicable here? Not as nice has having it built into the language of course, but it would at least eliminate these runtime costs.
I've been considering a safe String class that prevents some characters like CR,LF,\ that are seldom needed in business strings but used in system level things. Drawing a line between these two would increase security.
I like the idea but really I want the type to capture unsafe/semi-safe/safe. Now the trick is, how expressive can the idea of semi-safe be?
https://www.joelonsoftware.com/2005/05/11/making-wrong-code-...
var foo = "<foo>" + Environment.NewLine +
" <bar>" + Environment.NewLine +
" <quax>" + Environment.NewLine +
"</foo>";
This would fix stuff like this nicely, but it's horrible syntax I think.It doesn't for code, it does for strings.
If you have:
var foo = "<foo>
<bar>
<quax>
</foo>";
That string is actually <foo>
<bar>
<quax>
</foo>
In raw string form, as per the proposal, the string would be: <foo>
<bar>
<quax>
</foo>
It allows you to write strings nicely inline, and the indentation in the string itself doesn't matter.This is just some syntactic sugar for strings that contain escape codes. It's still just a 'string'.
Admittedly having one string type is still more than C, or indeed C++ bother with but we might notice that those languages have a pretty terrible relationship with strings and suspect that's not a coincidence.
https://github.com/dotnet/csharplang/blob/main/proposals/utf...
new [] {new [] {1, 2}, new [] {3, 4}};
Something like:
@[ @[1, 2], @[3, 4] ]
Thanks!
When a language gets its fourth or fifth string literal syntax, its process is probably broken.
cat <<-EOF
content
not
indented
EOFAnd the code sample you show differs only in that closing quote isn't indented. You don't explain why and how that change would affect the generated string.
Each line in the literal will have leading whitespace trimmed off, up to where the closing quotes are.
(What happens if the closing quotes pass some of the text?)
That's an error. Called out here: https://github.com/dotnet/csharplang/blob/main/proposals/raw...
Hi there. This is explained in the spec in a few places. In the examples section it explicitly states:
> To make the text easy to read and allow for indentation that developers like in code, these string literals will naturally remove the indentation specified on the last line when producing the final literal value.
> If the indentation behavior is not desired, it is also trivial to disable like so:
I thought that was clear as the prior explanation says that we remove the indentation from teh last line. And then i show how you can disable it. Specifically, as you noted because the closing quote line is no longer indented. Cheers!