The curious case of the missing period
tjaart.substack.com
tjaart.substack.com
If I wanted to root cause this, the real problem is right there. Implementing protocols correctly is hard and bugs like in the post are common. A properly implemented SMTP client library, like one you would pull off the shelf, would accept text and encode it properly per the SMTP protocol, regardless of where the periods were in the input. The templating layer shouldn't be worrying about SMTP.
That's why it's a best practice to specify protocols at a very high level (e.g. using cap'n'proto) instead of expecting every random sleep-deprived SDE2 to correctly implement a network exchange in terms of read() and write().
Engineers get better faster because they leverage better tools and build tools to overcome their own shortcomings and leverage their strengths, not by constantly being beat into shape by unforgiving systems.
To be fair, what you and OP said is not an uncommon mentality. It's even shared in a way by Torvalds:
> [easier to do development with a debugger] And quite frankly, I don't care. I don't think kernel development should be "easy". I do not condone single-stepping through code to find the bug. I do not think that extra visibility into the system is necessarily a good thing.
> Quite frankly, I'd rather weed out the people who don't start being careful early rather than late. That sounds callous, and by God, it _is_ callous. But it's not the kind of "if you can't stand the heat, get out the the kitchen" kind of remark that some people take it for. No, it's something much more deeper: I'd rather not work with people who aren't careful. It's darwinism in software development. It's a cold, callous argument that says that there are two kinds of people, and I'd rather not work with the second kind. Live with it.
He has similar views about unit tests btw.
I personally would prefer to work with people who are smart & understand systems and have machines take care of subtle details rather than needing to always be 100% careful at all times. No one writing an SMTP parser is at Torvald's level.
I'm not arguing that this excuses you from being careful or failing to understand things. I'm saying that defensively covering your flank against common classes of mistakes leads to better software than the alternative.
This is a point I agree with and the fact I see it mentioned so rarely, that standards are split across multiple RFC's makes me suspect that people don't mention it because they don't know because they never read them in the first place, and rather try to follow the implementation of some existing program.
This makes me wonder: How could the IETF's approach to standardisation be improved? I'm not sure how to fix this problem without overhauling everything.
SMTP is an example of an unnecessarily complex design, and the implementation bugs reflect it. SMTP shouldn't be hard for someone to correctly implement by themselves (even though I agree that people shouldn't be re-inventing the wheel).
If it wasn’t a period, it would be something else & you’d have to handle that instead.
That's an incredibly reductionistic view of the world that's utterly useless for anything (including actually engineering systems) except pedantry. It's obvious that the level at which you include control information is meaningful and significantly affects the design of the protocol, as we see in the submission. Directly embedding the control information into the message body does not lead to a design that is easy to implement.
> If it wasn’t a period, it would be something else & you’d have to handle that instead.
Yes, and there are many other design choices that'd be significantly easier to handle.
It's very reductionistic, because it intentionally ignores meaningful detail, and it's pedantic because it's making a meaningless distinction.
> It's a reminder that there is no magic.
This is irrelevant. Nobody is claiming that there's any magic. I'm pointing out the true fact that details about the abstraction layers matter.
In this case, the abstraction layer was poorly-designed.
Good abstraction layer: length prefix, or JSON encoding.
Bad abstraction layer: "the body of the email is mostly plain text, except when there's a line that only contains a single period".
There are very, very few problems to which the latter is a good solution. It is a bad engineering decision, and it also obfuscates the fact that there even is an abstraction layer unless you carefully read the spec.
-------------
In fact, the underlying problem goes deeper than that - the design of SMTP is intrinsically flawed because it's a text-based ad-hoc protocol that has in-band signaling.
There are very few good reasons to use a text-based data interchange format. One of them is to make the format self-documenting, such that people can easily read and write it without consulting the spec.
If the spec is complex enough that you get these ridiculous footguns, then it shouldn't be text-based in the first place. Instead, it should be binary - then you have to either read the spec or use someone else's implementation.
Failing that, use a standardized structured format like XML or JSON.
But there's no excuse for the brain-dead approach that SMTP took. They didn't even use length prefixing,
MTP had one concern which was to get mail over to a host that stood a better chance of delivering it, where the total host pool was maybe a hundred nodes?
I speculate that Postel and Sluizer were aware of alternatives and rejected them in favor of things that were easily implemented on highly diverse, low powered hardware. Not everyone had IBM-grade budgets after all.
Alternative implementations of mail that did follow the kinds of precepts that you suggest existed at one time. X.400 is the obvious example. If I recall correctly, it did have rigorous protocol spec definitions, message length tags for every entity sent on the wire, bounds and limits on each PDU, the whole hog. It was also crushed by SMTP, and this was in the era when you needed to understand sendmail and its notoriously arcane config to do anything. So sometimes the technically worse solution just wins, and we are stuck with it.
JSON needs to escape backslashes, SMTP needs to escape newline followed by period. If you're already accepted doing escaping, what's the issue?
> Bad abstraction layer: (...)
In this context, it shouldn't matter. Sure, "mostly plaintext except some characters in some special positions..." is considered bad in modern engineering practice, however it's not fundamentally different or more difficult that printf and family. You wouldn't start calling printf without at least skimming the docs for the format string language, would you?
> It is a bad engineering decision, and it also obfuscates the fact that there even is an abstraction layer unless you carefully read the spec.
There's the rub: you should have read the spec. You should always read the spec, at least if you're doing something serious like production-grade software. With a binary or JSON-based protocol, you wouldn't look at few messages and assume you understand the encoding. I suppose we can blame SMTP for design that didn't account for human nature: it looks simple enough to fool people into thinking they don't need to read the manual.
> There are very few good reasons to use a text-based data interchange format.
If you mean text without obvious and well-defined structure, then I completely agree.
> One of them is to make the format self-documenting, such that people can easily read and write it without consulting the spec.
"Self-documenting" is IMHO a fundamentally flawed idea, and expecting people to read and write code/markup without consulting the spec is a fool's errand.
> it should be binary - then you have to either read the spec or use someone else's implementation.
That's mitigating (and promoting) bad engineering practice with protocol design; see above. I'm not a fan of this, nor the more general attitude of making tools "intuitive". I'd rather promote the practice of reading the goddamn manual.
> But there's no excuse for the brain-dead approach that SMTP took. They didn't even use length prefixing,
The protocol predates both JSON and XML by several decades. It was created in times when C was roaming the world; length prefixing got unpopular then, and only recently seems to en vogue.
Exactly! This is an even better phrasing of my point.
"The first two bytes represent the string length, in big-endian, followed by that many bytes presenting the string text."
and an in-band signalling protocol:
"The string is ended by a period and a newline."
In the second one, you're indicating the end of the string from within the string. It looks simpler, but that's where accidents happen. Now you have to guarantee that the text never contains that control sequence, and you need an escaping method to represent the control sequence as part of the text.
You always know what the next byte means because either you did a prefixed length, your protocol has stringent escaping rules, or you chose an obvious and consistent terminator like null.
The terms harken back from the day of circuit switched networks but now that we have heavily transitioned to packets, bands are an artificial construct on top of packets and applying the term isn’t very clear cut.
The main property of in-band data in the circuit-switched network days is that you could inject commands into your data stream. If we apply that criteria that to a modern protocol, even if you mix metadata and data in the same “band,” if your data can never be interpreted as commands then “out of band” makes an apt description.
In this case, somewhere the protocol abstraction layer got broken, and the message text ended up being treated as already serialized. It's not a problem with the protocol per se, but with bad implementation of its API (or no implementation at all, just printf-ing into the wire format).
When we’re talking about whether someone can inject data into the link, we’re talking about the end user and not the software. If we’re talking protocol design, then you wouldn’t want regular data to be able to inject commands by simply existing.
It shouldn't, unless you're bypassing the actual protocol serialization layer (or hitting a bug in the implementation). Which is what's the case here. Protocol design can't address the case of users just writing out some bytes and declaring it's a valid protocol message.
I’m replying to a post where someone said most protocols have in-band signaling and therefore this problem is unavoidable.
I understand that constraint, and it seems reasonable - but in that case, why not use a length prefix? That should be even more efficient than having to scan for a line containing a single period and nothing else.
Hence such horrors as MIME delimiters:
-----=_Part_2827761_1947716067.1716352583971
Content-Type: multipart/alternative;
boundary="----=_Part_2827760_1372041171.1716352583736"
We still have the mess that is the required and standardized behavior of HTML5 parsers faced with bad data.And, even though you might not see a way to call into the unused code, an attacker might find a way (XZ Utils).
But for SMTP libraries, that's often part of stdlib (Ruby, Python, PHP, ...).
There is a multitude of classes of errors and security vulnerabilities, including "SQL injection", XSS, and similar, that are all caused by the same mistake that this case of missing period was[0]: gluing strings together. For example, with SQL queries, the operation of binding values to a query template should happen in "SQL space", not in untyped string space. "SELECT * FROM foo WHERE foo.bar = " + $userData; is doing the dumb thing and writing directly to SQL's serialized format. In correct code (and correct thinking), "SELECT * FROM..." bit is not a string, it just looks like one. Same with HTML templating[1] - work with the document tree instead of its string representation, and you'll avoid dumb vulnerabilities.
So, if you want to avoid missing dots in your e-mails, don't inject unstructured text into the middle of SMTP pipeline. Respect the abstraction level at which you work.
See also: langsec.
--
[0] - And therefore should be considered as a single class of errors, IMO.
[1] - Templating systems themselves are thus a mistake belonging to this class, too - they're all about gluing string representations together, where the correct way is to work at the level of language/data structures represented by the text.
[1] https://genshi.edgewall.org/wiki/Documentation/xml-templates...
[2] https://zope.readthedocs.io/en/latest/zopebook/AppendixC.htm...
This is not universally true.
JavaScript has an amazing feature called tagged template literals which let you tag a string with interpolations with a function that handles the literal and interpolation parts separately. This lets the tag function handle the literals as trusted developer written HTML or SQL, and the interpolations as untrusted user-provided values.
Lit's HTML template system[1] uses this to basically eliminate XSS (there are some HTML features like "javascript: " attributes that require special handling).
ex:
html`<h1>Hello, ${name}</h1>`
If `name` is a user-provided string, it can never insert a <script> or <img> tag, etc., because it's escaped.There are similar tags for SQL, GraphQL, etc. Java added a similar String Templates feature in 21.
Be careful with that "never". A curious and persistent person might discover a bug in the implementation, leading to something like the Log4Shell issue.
But it'd be similar with with other template systems. If the interpolation should allow any string, there's really no validation to be done.
The bulletproof way of doing this is working at the level of abstraction of your target language. With HTML, that would be a tree structure. For example, if your HTML generation looks more like:
["H1", "Hello, " + name]
and that is passed to code that actually builds up the tree and then serializes it down to HTML, then there is no way `name` could ever break the structure or inject anything.--
[0] - I skimmed the docs of Lit, it seems there are restrictions on where interpolation can be placed, but I don't think they're actually building up the tree expressed by the static parts.
Lit is not working at the serialized level, at all. It parses the templates independently of any values, and the values are inserted into the already parsed tree structure. There's is literally no way for values to be parsed as HTML.
In Germany, where I work, it is usual at the end of employement to ask for a letter of recommendation ("Zeugnis") that lists the tasks performed, and how good the employee was. It is an important document, as it will typcally be required when applying for jobs. Obviously, no employee would accept a document explicitly stating "this guy is a lazy bastard, do not hire him", so there is a "Zeugnissprache", a "secret code" to disguise this information as praise. One part of this code is that a missing period in the last sentence means "please ignore everything said here, this guy is horrible".
How do I know? I let a lawyer check my Zeugnis after my last employment, and (I assume out of lack of care, as all my performance reviews were positive) the last sentence was missing the period.
Gee, what could possibly go wrong?
This legend comes from the fact that HR people cannot be too explicit about the fact that you've been a pain in the ass (you could probably sue if it's too transparent), so if they have nothing positive to say they will commend your punctuality or something equally as mundane. It's not secret codes, it's like... "bless their heart", but in HR talk. Plausible deniability if you want to sue, I guess. "But it's a good thing, your honor! They were always on time!"
...sigh:
Secret codes as in "watermark-level omission of characters" are a myth. Lingo and jargon do however exist, and convey meaning in a particularly subtle way. They are shared and taught by culture, not by a secret handbook passed down from generation to generation. See also dogwhistling.
The goal is to protect the issuer, not to selflessly inform the recipient.
> You will be lucky to have this person work for you.
This isn't a veiled statement. It's outright dunking on the applicant.
I think there's some deeper issue with the language/culture here.
> you would be lucky to get this employee to work for you!
"In politics, a dog whistle is the use of coded or suggestive language..."
But in the specific german case, the code is not even that secret. This is a formal document with a very specific structure, and very standardized phrases. There is even specific software to generate the text out of performance ratings. Basically something like this:
- John was overal engaged: he is a lazy bastard
- John was engaged: he is OK
- john was very engaged: he is good
- john was always very and thoroughly engaged: he is very good
You just use the mail program from mailutils or whatever.
Just from a point of view of deliverability, developing bare bones SMTP interaction over a socket is a nonstarter. You can't just connect to random mail exchange hosts directly and send mail these days. A solution has to be capable of connecting to a specific SMTP forwarding host (e.g. provided by your ISP). For that, you need to implement connections over TLS, with authentication and all.
Also, a slightly ironic thing is that cron already knows how to send mail. The output of a cron job is mailed to the owner. Some crons let that mail address be overriden with a MAILTO variable in the crontab or some such thing.
In other words, every other platform expands until it can summarize emails.
It's not a good reason, no--definitely not--but it's a real reason.
You can require it and it gets provisioned.
But the other thing is: Don’t vendor your dependencies. Those libraries you use need to be updated regularly and timely, and absolutely not “only as necessary”. If updates lag behind or are avoided entirely, bugs like this can be huge problems even when the upstream code has been fixed, for people who thought that they should update only when they, themselves, see a problem or need.
I agree 100%.
Was very hard for the team. There was a shouting match over it, some hard feelings. The code was written in the spirit GP alludes to by an enthusiastic executive who wanted to help lighten the load. I should've rejected the PR but was intimidated to reject the exec's code (not an engineering reason for an engineering decision!). The exec was a good data scientist but not as strong a coder as me, and parsing binary files is one of my specialties.
Friends don't let friends parse using "indexOf()".
Another, more domain specific one: We have to talk to a GPS receiver that connects via RS-232 and outputs NMEA formatted data. This device, like all such devices, outputs a small subset of the standard. So parse just enough of it for that one device we have to get working, and ship it. Then someone attaches a different device that outputs a different subset of the standard and the software fails (sometimes gracefully, sometimes not).
It was a struggle to get the vendor to even acknowledge the problem, since it would only happen on 7 days out of the year, and usually 60 days apart. (There are only two consecutive months with 31 days, July and August.) It did eventually get fixed.
I seem to recall writing a workaround, but then it being stalled by our customer's Change Control Board.
The alternative seems worse: your own application's stability is now at risk against upstream changes that could break your code. Sure, you might not get a fix immediately, but I'd rather know I'm making a change because I need a fix than introducing instability and additional risk that I don't want to subject myself to. "If it ain't broke, don't fix it."
(For apps, a lock file will do it too.)
https://chromium.googlesource.com/chromium/src/+/HEAD/third_...
(And surely you should have tests to verify all your own functionality after upgrading a dependency?)
Relatedly, escaping somehow seems to be a foreign concept for a lot of programmers, who wouldn't ever see the above situation and ask themselves "but what if I want to send an email with a line containing a single dot?" yet another large group of them finds it perfectly logical and easy to understand.
I'd be surprised if any significant portion of software developers learned anything like that in the last 30 years, at least.
SMTP https://www.rfc-editor.org/rfc/rfc5321#section-4.5.2
Also in POP3 https://www.rfc-editor.org/rfc/rfc1939#page-8
As soon as I printed them, the error was clear. My version ran to two pages and the good implementation one page. I had not been careful to clear the buffer before sending the data (mbufs don't you know).
This still cracks me up.
... lots of text ... <br>We are happy to welcome you to our family.<br>
or whatever. But if you blindly split HTML into lines, it will break tags.the company is just blindly unaware to their minor problem.
While quoted-printable is supposed to have a max line limit of 78 characters including CRLF rather than 1000, email clients tend to be permissive.
The square brackets allow us to stick a port number on it: [ffff::0123:4567]:6301
As soon as I read the above, I knew the below would be the result.
> It seems one of our other teams haven't gotten around to patching this bug in their code.
I share your sentiment, thank you for reading.
[1] https://en.wikipedia.org/wiki/Wireless_Application_Protocol
[2] https://en.wikipedia.org/wiki/ColdFusion_Markup_Language
What the...
;-)
(-;
Maybe you're thinking of base64 encoding?
But this is SMTP, I have no doubt there are gateways out there that will reencode the mail and put the dot back in the wrong place, eg in the name of wrapping all URLs behind a phishing-warning-page
Other than the obvious moral that protocols should be implemented properly, the moral of the story is that all abstractions are leaky, and it will always be useful to understand the lower levels.
Of course if it was modern, different question.
Why is it implemented that way?
If a single period means end of mail then more than a period means it's mail data.
Why deleting the period in the first place? Couldn't they store one byte to check the next?
. Ends body
.. One period <--
... Two periods
.A A
..A .A
A A
Anyway that’s stupid and only helps if you compose your email right in a tcp session.there's a secondary issue here, why in the world would you auto split a monetary value across a numeric decimal indicator? why would you split lines at all for this use case?
> The maximum total length of a text line including the <CRLF> is 1000
And, they also state that the period disappeared because it was placed at the start of the next line when the split occurred.
From which one can deduce that they were doing the most basic "split" possible, splitting at the exact 1000 octet point, i.e. something like:
if (length(line)>1000) then:
line1=string_range(line,0,999)
line2=string_range(line,1000,end)
fi
And if the period in 27.00 ended up exactly at offset 1000 in "line" then it got 'split' into line 2 as the first character of line2.So a message intended to be sent by an SMTP client:
DATA
Hello customer,<br>[978 characters] 27.00
Was erroneously formated into:
DATA
Hello customer,<br>[978 characters] 27
.00
.
The period after 27 will be removed. And this is how the html will be rendered.
Hello customer,
[Lots of text] 2700
so splitting 27.00 on the . becomes 27 00, because the CRLF is significant to the client.
you would want to split at whitespace, not at any other character -- unless you had a 999+ string of non-whitespace of course.
perhaps the author didn't know or didn't realize or thought it insignificant to his point that in addition there was a quoted-printable encoding, in which case i believe the trailing/mandatory CRLF can be made non significant for client rendering. personally i still would have split on actual whitespace. (well, i wouldn't have written an smtp client in the first place.)
ironically titled..
Good luck having it handle any of the SMTP craziness that isn't on the short introduction to the protocol.
Or the usage example.
Not to mention all the weird bandaids on top of bandaids to try to get sender verification and tamper proof emails working. That alongside the complete lack of end to end encryption.
It's just an incredibly unpleasant tech stack from top to bottom, through and through. The amount of moving parts/pieces of running software needing to cooperate just right to even function as a simple outgoing-only mail server is too damn high.
In the future, I hope that email will be like fax, or at least treated like http in comparison to instant messaging (https).
- The SMTP client spec says that an additional period would be added here.
- The SMTP server spec says that it would remove this additional period, bringing us back to one period.
I don’t get how this led to there being no period at all. Am I missing something?