Plain text is still one of the best technologies we have
deadparrotbbs.com
deadparrotbbs.com
[1] https://archive.ph/FhG5L (the original either got deleted or login-walled, here is the archived version)
[2] https://hn.algolia.com/?q=always+bet+on+text
[3] https://news.ycombinator.com/item?id=26164001
[4] https://news.ycombinator.com/item?id=8451271
No Javascript, no CAPTCHA, no geo-blocking
https://web.archive.org/web/20141014043202if_/http://graydon...
No Javascript Algolia search
Text-only (JSON)
https://hn.algolia.com/api/v1/search?query=always+bet+on+tex...
For one, there are a few different ways to terminate lines. All major operating systems now tend to use just \n, but I have older files that use \r\n (Microsoft), \r (Macintosh) or \n\r (RiscOS).
There are also different opinions on how text files end in different operating systems. In POSIX, all lines are terminated by \n, even the last one. Microsoft software still tends to insist that the last line of a file is special case that doesn't need to be terminated even now that they have otherwise adopted POSIX style line endings. In Microsoft's view, it seems that the line ending sequence separates lines rather than terminate them. Files created according to this view don't play well with tools like cat(1) if your intent is to concatenate the lines of two files, but it seems other Unix clone tools have adapted to the possibility that the last line isn't terminated properly.
Finally there's the encoding problem. I don't know of a good tool that determines the original encoding based on heuristics and re-encodes to UTF-8 but if someone does I'd love to know. If I know that the input language is English for example it shouldn't be too hard to determine what encoding the funny byte used in contractions or the funny bytes used in quotes belong to. Still, in English most of the files that use 8-bit encodings remain quite readable if you just box out the invalid bytes.
Granted, I knew that all files are UTF-8, so I didn't need to have it "guess" what the encoding was without the BOM.
0x0 0x0 0x0 0x0 0x2 0x0 0x0 0x0 0x0 0x7 0x1 0x1 0x0 0x0 0x0 0x0 0x3 0x1 0x2
and try to figure out what that is all about.
[0] Windows newlines not withstanding.
Fortunately they’re easy to test for and most languages have standard libraries that make this painless.
And in practice, still a rather large number of them, going back to the 1970's - https://en.wikipedia.org/wiki/Extended_ASCII
As it is programs have to guess the following;
A) encoding. If ANSI which code page? If unicode which encoding? In the case of utf-16 big or little endian?
B) line endings? CR? LF? CRLF? I guess no-one uses LFCR... right?
C) number formats? 100.000 or 100,000? Date formats? is that mm-dd-yyyy or dd-mm-yyyy?
D) what human-language is it in?
E) CSV? Don't get me started...
Yes. Text is a lot easier to load, parse, make guesses about than say XLS. Yes all of the above things can be "guessed" to a greater or lesser extent.
Yes BOM exists to at least try solving the encoding question. Of course most files don't use them. Lots of tools don't support them. And they don't solve any of the other issues.
But sure, ASCII files in English with US date formats, and Windows line endings....no problems at all...
Basically it's a text-based language for embedding semantic metadata into some kind of underlying text stream. Is this interesting to you?
The line endings problem really isn't a problem. Pretty much every text editor out there can handle different line endings.
I don’t think number nor date formats are relevant here. For example you could have that same problem entering text into MS Word. That’s really more of an issue if you want to use text as a database rather than a document format, which isn’t really something that even plain text advocates would generally recommend.
As for human language detection, that’s a much easier problem to solve than decoding a proprietary binary blob.
CSV definitely has its warts. But it’s not like that’s the only plain text option for serialising data. (JSON, jsonlines, YAML, XML, etc). Or you could use the actual ASCII codes reserved for records, if you really wanted something that didn’t require quoting and escaping in plain text. It’s actually a pity nobody does this.
This is the eternal tradeoff of specificity vs. generality.
Most national standards now call for a space (preferably a half width space, Unicode: U+2009) as the thousands separator and then either comma or full stop as the decimal separator.
The half width space was standardised in SI in 1948. See, i.a., https://snl.no/tusenskille
“Windows” newlines are also the standard in many communication protocols like HTTP and SMTP. That’s not because of Windows or DOS, it’s because it was the standard for teletypes which needed bot CR and LF. It’s arguably systems like Unix that deviated from that standard.
I agree that beyond ASCII there is a slope from well-supported to less-supported and quirky to problematic areas of Unicode. For example, HN filters many Unicode text elements like combining characters (Zalgo text) and emojis, and there is no specification to point at what it supports.
Even within ASCII, most control characters don’t have a portable meaning. So it’s really just the printable subset of ASCII, and strictly speaking not even that, given that there are regional variants of ASCII, such as the Japanese one where backslash becomes the Yen sign.
Don't ask me what my point is
With plaintext you can hand the file off to anyone on any device--the caveat being absolutely no one wants to be handed a plaintext script. The software I used ten years ago to write plays is long since deprecated and those files are basically unopenable. Plaintext however remains.
Sharing aside, I imagine a plain text format would also be helpful for version control.
I built sharing and (sort of) versioning into the web app. Everything is saved in localStorage and on a MySQL server. You can Save or Save As, and every time you Save As it inserts a new database row and gives you a new, sharable URL.
I wanted an experience where you didn't have to login to use it, but ultimately, having a way to keep track of all those URLs would actually be a pretty decent versioning system--add diff checking between script versions and you end up with something pretty handy.
It's not baked-in to the weights of any model, but agents can write their own tools to work with arbitrary binary formats and get many of the benefits of off-the-shelf unix utilities.
I love that the Godot game engine has a textual scene description. That's a stark contrast to Unreal Engine's binary format for blueprints (visual scripting).
Symbols on paper is best.
Plain text in a file that can be opened by any computer comes close, but needs a tool.
Fancy file formats that need not only hardware but also special software are the worst.
However, there's an important tradeoff, which is the fancier formats can present information in ways mere symbols might struggle with, and can also use interactivity to improve understanding.
That's how Unicode started but nowadays Unicode expresses things that are only used because Unicode itself introduced them.
And this while many important things in the "have used text in history" category have been unfinished or not tackled at all.
The way I see it, emojis are practically single-handedly responsible for near-universal adoption of Unicode in the Anglosphere. Otherwise, it would be all too easy for lazy developers to just use ASCII and then shove their fingers in their ears about any other writing systems.
My point is that Unicode's own mission statement says: "The Unicode Consortium enables people around the world to use computers in any language."
I think it strayed from that. And I'd argue the mission was only ever fulfilled when seen through an anglospheric lens. You don't even have to summon Han Unification, the elephant in the room. Documents from Western Europe from a couple of decades ago can't be represented properly.
Example: you cannot properly represent a German telephone book today. Take "Müller" and the French surname "Haüy". Both contain the same Unicode character, but in "Müller" it is an umlaut and sorts like "ue" (Mueller), while in "Haüy" it is a trema and sorts like a plain u (Hauy). For all intents and purposes these are different letters. That they look similar (not the same, in proper typography) is a mere coincidence. Yet Unicode encodes them identically, so the information needed to sort correctly is simply not in the text anymore.
Han Unification is the other sore point, and a much bigger one.
I'd like Unicode to focus on its core mission, encoding the languages that exist, before it starts shaping new ones. For that I'd propose a grace period: a new character should have to be in significant active use for 10 to 20 years (not merely implemented on a popular device) before it is accepted into the standard.
Why say anything in 2000 bytes, which you can read in 10 seconds, when you can express it in 4MB, and watch it in 7 minutes?
Associating such meta data with text files wasn't simple in the beginning, but that changed in the 1990s. Nearly all filesystems since then allow meta data: POSIX compliant filesystems support extended attributes, mostly used for security meta data. Even ancient HFS (Apple Mac) supports meta data. NTFS supports it with alternative file streams.
While possible, no system I know of attempts to use any such meta data, probably due to a lack of standards. A pity.
Oh and now you could have LLMs write the Markdown, but chances are no human reads that, or even if you do, the formatting is going to be insane. I have to keep telling Claude to give me a .txt instead of .md when asking for a context dump. Maybe .md is popular for agent readmes because they optimize around its headings.
I'm assuming any tool in this regard would expect the user to write an EBNF grammar.
I found ANTLR to be nice but it's way too convoluted to use as a tool with Java dependencies and seems to be stagnating. And tree sitter just is too convoluted and requires the added C/C++ overhead to understand how to use it.
Another choice, though not necessarily taking in an EBNF, could be Racket [3].
[1]: https://lark-parser.readthedocs.io/
The fact that unicode maps the lower 7 bits to its own character set is a nice touch but none of the unicode sets are plain text.
Unicode are multibyte characters with variable byte length and endianess at play. If you read it wrong or guess the length wrong your results might be anything but useful.
> The fact that unicode maps the lower 7 bits to its own character set is a nice touch but none of the unicode sets are plain text.
That is only true for the mapping of Unicode character (code point to be exact) to UTF-8, which is an encoding of Unicode characters.
> Unicode are multibyte characters with variable byte length and endianess at play.
That is only true of the UTF-16 encoding. UTF-8 does not have endianess, UTF-32 does not have variable length per code point.
None of that is true for Unicode, because it is abstracted away from any byte representation. Furthermore, getting back to the start:
> ASCII (7-Bit) is the only widely understood charset there is. Everything beyond this point depends on the loaded charset.
ASCII also depends on how you try to decode your text. If you interpret two bytes as one character, ASCII will never be correct. There is nothing magical about one byte mapping to one character. I'd even argue that the only reason ASCII support is so universal internationally is because of UTF-8. Otherwise many countries would default to encodings where ASCII is not a subset (as they did before UTF-8 became common). So IMHO, UTF-8 and Unicode are the only widely understood encoding and character set.
UTF-32 grapheme clusters can easily be six or seven code-points long. Examples: modifier characters, flag/regional-indicator sequences, (and to a lesser extent, combining characters).
Why does it matter? Because you can't insert a linebreak in the middle of a grapheme cluster.
So even after converting all of your UTF-8 text to UTF-32, you STILL have to deal with multi-character code-point sequences.
Madness!
I wish he'd spent a little more time on the fundamental issue of CR vs CR/LF vs LF. That's been a "plain" text nightmare since before Unicode and many of the other complexities existed.
I created Brashtag [1]. It is simpler than markdown.
See how easy it is to write programs to process brashtag [1].
I can chat with llm, interact with jupyter kernel and do literate programming from any text editor all in the same document [3].
It took me like 10 minutes to build this mermaid like thing [2].
List of frustrations:
- One day Anki crashed. --- You can just write #card{....} in a text file. Some program would read it and show flashcards in browser.
- Another day JabRef crashed. --- You can just write #bib{....} put link to a paper and its citation info there and have a program download the paper and copy the citation to bib file.
- Markdown parser processed mathjax wrong. Why keep trying different markdown parsers? Just write #h1{}, #b{} #i{} and just convert them to html tags.
- Literate programming tool I was using failed at some corner case.
- My notes were scattered across files and tools. I couldn't find anything. I tried to have a text file where sections were separated by four dashes. But that didn't allow nesting. Also, the parser I wrote for that hit corner cases. Why not just put #note{} #JIRA-334{} etc. in text file.
[1] https://github.com/PratikDeoghare/brashtag#some-programs [2] https://github.com/PratikDeoghare/brashtag/tree/master/cmd/m... [3] https://www.youtube.com/watch?v=IMXgIE0Vljg
> 'this is a bag named code with a blob, therefore it is code."
This would make #code{} a special bag. Right now no bag is special. Also it would be hard to find the closing } if somebody wrote unbalanced paren code in the bag like #code{ func main() { }.
You could assume that { and } as a structural delimiter only applies to node types that can nest. Since blobs do not nest, then a tag followed by a backtick (or multiple backticks) could be a valid delimiter in lieu of braces. In this case, tags would extend to blobs.
In my opinion, pairing minimal node types with differing delimiters helps legibility and is not overly mentally taxing to learn when node types are few.
#foo`fmt.Println("Hello world!")`
instead of current
#foo{`fmt.Println("Hello world!")`} .
This would mean code can have tags too just like bags. Right now only bags can have tags and children. If you want to tag something, you have to put it in a bag. This works for tagging blogs, code, other bags.
I think the extra braces are worth suffering for uniformity like the trailing commas.
It's that it can be a Microsoft Word, etc., killer.
Because it stops after making you a tree. Markdown takes your document all the way to html. Making it do anything else requires much work.
file.md -> markdown -> file.html
file.bt -> brashtag -> tree -> to-html -> file.html
file.bt -> brashtag -> tree -> todo-list-maker -> todo list
...
file.bt -> brashtag -> tree -> ... -> ...
Brashtag is not about code at all. Most text is already a valid brashtag. It is optimized for human writing. For example, this comment is a valid brashtag document if you surround it with #{} which can be done by a program easily.
As you have noticed it is poor substitute for json. It would be cumbersome to write config file in brashtag.
> This isn't a competitor to markdown.
You can write blog post in markdown. You can write it just as elegantly in brashtag. Only brashtag does not convert the doc to html for you. It just gives you tree. You write a program that processes the tree and produces html.
Example,
Markdown: # Lorem ipsum dolor sit amet
Consectetur _adipiscing_ elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam.
---
Brashtag:
#h1{Lorem ipsum dolor sit amet}
Consectetur #u{adipiscing} elit, sed do eiusmod tempor incididunt #b{ut labore et dolore} magna aliqua. Ut enim ad minim veniam.
OR you could write
#really big title{Lorem ipsum dolor sit amet}
Consectetur #super duper underline{adipiscing} elit, sed do eiusmod tempor incididunt #emphsize{ut labore et dolore} magna aliqua. Ut enim ad minim veniam.
All of these use slightly incompatible formats, the irony of which will not escape the educated reader.
imagine the world wide web but markdown (and decentralized). gno.land is that.
And then ISO, which owns a bunch of extended ASCII encodings that provide essential support for "plain text" in languages other than English (and an ugly historical hostility toward non-European languages).
And then Unicode Consortium, which owns Unicode, its ever-expanding list of complexities and exceptions and an ugly history of hostility toward Asian languages.
Nope. Skintones are part of Unicode. It also has 33 control characters from ASCII including one that rings a bell... There are also numerous characters added by Unicode that are literally called "layout controls".
If you want a format that provides purely semantic information, then "plaintext" doesn't fit the bill.
Just the other day I had opus read a 10 year old proprietary file format. The idea that we will lose the ability to use file formats is not consistent with reality.
My understanding is that HN has started incorporating AI tooling in the back-end to scan for LLM-generated submissions and source content, in an effort to discourage it and encourage human-made content (with human discussions, one would hope). Why do we keep having these kinds of articles every day? Entire "apps" are generated - see the lighthouse one - and submitted, and are plain-as-day LLM-generated nonsense.
Anyway, flagged.
In the end we can't even tell what's what, it's exhausting.
Worse the tools for detecting it have an insane false detection rate. ESL = FP. Good writer = FP. Garbage human slop writer = A'ok.
And, well, neither of us really knows the ground truth. Maybe the guy who has a repo filled with articles co-authored by Claude is actually not posting slop this time and HN users are wrong and Pangram flagging it is just a false positive. But I can't blame anyone for being tired of giving this the benefit of the doubt.
I'm ESL. The mistakes I make are part of what makes me human. If HN'ers are having an allergic reaction at the first hint of slop, it's because we've been drowning in it for months.
To my ears, I find accusations of "it was written by AI" to be a bit MAGA-ish. It's a far too convenient way to say that I don't like what someone is saying without actually having to think for yourself (or heaven forbid, actually read it).
The author's picture clearly has an LLM-generated background, (and is a little tacky, IMO.) That being said, I wouldn't dismiss an article merely because someone used an LLM while setting up their blog.
Even has a simple an/a mistake that I doubt a chatbot is going to make.
* All the posts are walls of text. In many of the posts, all of the paragraphs are just about exactly the same length.
* None of the posts contain personal stories, experiences, or anecdotes.
* All of the posts are hyper-focused around advocacy of "old tech" and decentralization. Which is great, but real blogs with this many posts have at least _some_ variation in subject matter and quality. This one's surprisingly uniform.
* The blog is relatively new, posts began about two months ago and there are 17 posts, that's a little over 2 posts a week. Sure, there are bloggers who post that often, sometimes daily, but they are usually much shorter posts, tend to be shorter, or do it for their job.
* The picture of the author on the About page is very obviously AI generated.
* The author has a GitHub account that had virtually no activity until March of this year, and basically all of the commits were co-authored by Claude.
To be clear, I'm sure a real human is behind the site/posts (as opposed to a fully autonomous agent), but I'd bet dollars to donuts that each article started out its life as a very short prompt.
I vibed TrustSpark for HN plugin some months ago, and I get true positive flags from this account all the time. Unfortunately my RSS feed doesn't see the flags, and I have to click the link to find out.