Stop Requiring CRLF Line Endings
fossil-scm.org
fossil-scm.org
The most up-to-date version of HTTP/1.1 spec is RFC 9112, which says:
> Although the line terminator for the start-line and fields is the sequence CRLF, a recipient MAY recognize a single LF as a line terminator and ignore any preceding CR.
"MAY", of course, is different from "MUST" or "SHOULD", so I feel like the author's claim that implementations rejecting bare NLs are broken is at odds with the specification.
> This flexibility regarding line breaks applies only to text media in the entity-body; a bare CR or LF MUST NOT be substituted for CRLF within any of the HTTP control structures (such as header fields and multipart boundaries).
They aren't going to allow sending LF until at least one bump to a higher protocol version where every server MUST accept it.
In the spaces world, you have the battle over how large indentation should be and everyone has to live with what someone else chose.
Consider this snippet, where tabs (⇥) equal four spaces (·) in my editor:
if (someCondition) {
⇥ Foo foo = SomeClass::SomeFunctionCall(myParameter1,
⇥ ⇥ ⇥ ⇥ ⇥ ⇥ ⇥ ⇥ ⇥ ··myParameter2,
⇥ ⇥ ⇥ ⇥ ⇥ ⇥ ⇥ ⇥ ⇥ ··myParameterN);
}
Looks good.Now you open this snippet in your IDE, where tabs are set to two spaces:
if (someCondition) {
⇥ Foo foo = SomeClass::SomeFunctionCall(myParameter1,
⇥ ⇥ ⇥ ⇥ ⇥ ⇥ ⇥ ⇥ ⇥ ··myParameter2,
⇥ ⇥ ⇥ ⇥ ⇥ ⇥ ⇥ ⇥ ⇥ ··myParameterN);
}
This matters more when the presentation is important, like I might write ASCII diagrams in my code, are have a large matrix with elements aligned, or similar. if (someCondition) {
⇥ Foo foo = SomeClass::SomeFunctionCall(myParameter1,
⇥ ······································myParameter2,
⇥ ······································myParameterN);
} blah(
blip,
blop,
bloop
);
And anyway, 'you see the code the exact same way I saw it when writing it' has the following problems:1. It doesn't help because you were writing the code, I have to read it.
2. It doesn't help because the way my brain works most likely isn't 100% interchangeable with how yours does.
3. It may even not be true because I may be using a different editor, a different editor theme, a different font, or, by golly, an entirely different monitor resolution.
That being said, I will agree with my sibling comments which recommend the combination of tabs for indentation and spaces for alignment.
I'm following a junior who is working on a project where everything, including braces, is indented like that.
Some lines overflow on a 2k screen.
So stupid... And it requires much more effort than just indenting logically depending on the depth of the current block. You have to wonder why people do this to themselves and to their colleagues.
Please let go of that childish esthetic to align certain stuff with spaces.
That's not a good reason, there are way too many writers with poor design sensibilities, so flexibility is better.
And in your example you'd use elastic tabstops so that a single extra tab would visually expand to align arguments for you if you like it, but not for me if I don't
Why should you have control over what indentation width I choose to use?
There are a bunch of reasons why I might want a narrower or wider tab width. Most obviously, screen real estate - I might be coding on a tiny device.
When you do indentation and alignment properly (tab for indentation, space for alignment, as others have said), it's formatted the way the writer wants, while also respecting the reader's preference for tab width.
Spaces for indentation is incorrect, pure and simple. And the war is not "won".
I routinely do this while drafting text and tables, and isn't this akin to what multi-cursor input achieves in many text editors?
1
2
3
Since after typing the digit the writing head is behind that character. You have to add a backspace to go back one column. 1NLBS2NLBS3 which then pushes the Cursor back to the right column.Once you add more than one character per row its all over anyways.
The reason is because I have just been typing into a table in a WYSIWYG editor which undoes the problem for me, returning to the beginning of each cell.So I tricked myself, because that is of course not what's being discussed in the article.
Ah, but that can be solved by returning the carriage to the start of the line :)
Thus giving you the perfectly valid control sequence: LFCR!
As the references show, this is already a big source of vulnerabilities - trying to push for a change in standards would likely make the situation much worse. At the very least, old unmaintained servers will not change their behavior.
I think we should accept that this ship has sailed and leave existing protocols alone. Mandate LF and disallow CRLF in new protocols, that's fine, but I don't think we should open this particular Pandora's Box.
[1] Simple example that doesn't use CRLF/LF disagreement: https://portswigger.net/web-security/request-smuggling
[2] Complex example that uses CRLF/LF disagreement: https://portswigger.net/web-security/request-smuggling/advan... (see heading 'Request smuggling via CRLF injection')
[3] Random report on HackerOne I found where allowing LF created a vulnerability in NodeJS: https://hackerone.com/reports/2001873
[4] https://sec-consult.com/blog/detail/smtp-smuggling-spoofing-...
> CR by itself is occasionally useful for when you want to overwrite a line of text you have just written. LF, on the other hand, is completely useless. Nobody ever wants to be in the middle of a line, then move down to the next line and continue writing in the next column from where you left off. No real-world program ever wants to do that.
Just for curiosity, isn't that what Whisper does, the data file format for Graphite? I remember reading about it in https://aosabook.org/en/v1/graphite.html
I realise this isn't core to the author's point, and I mean the question to learn something rather than trying to be a pedant. I think it uses python byte offset semantics rather than CR, but for all I know the python implementation uses CR under the hood.
The article overall is a fun read! I love the idea of trying to identify simple pieces of legacy ideas and clear them out.
IMO the day where we don't program computers anymore will come before the day everyone gets to forget what /r means.
There may be some situations where it is helpful to have them separate, although for most uses (including most internet protocols) it isn't.
But, I agree with the part about Windows that O_TEXT should not be used; you should use O_BINARY.
The NL operation is defined differently for different OSs. Yes, CRLF is a relic of the teletype, but in practice, it's really the Windows convention inherited from DOS. L/unix has always been LF. On old Mac operating systems it was CR alone.
Unicode does in fact have a NEXT LINE character, 0x85. But it's not what the article is talking about.