BNF was here: What have we done about unnecessary notation diversity (2011) [pdf]
grammarware.net
grammarware.net
Some of you probably work within SAP, and it's the same kind of problem - you have to audition for the part of some random German business logic nerd until you maybe get it. And it's still effortful every time.
It's worth mentioning that standing on the shoulders of giants is strange when so many of those giants are still alive and working on their monoliths. The whole experience is as much a cultural and historical and borderline religious pilgrimage as it is a syntactical exercise.
A bit like getting a vague answer to an Old Testament God prayer, I understand C++, Python, etc. on a deeper level not by programming, but watching the creators talk about their languages. If you hate C++, go watch Bjarne get interviewed. The language will look like his handwriting the next time you open your IDE. It's so weird.
If you struggle with Linux, go watch YouTube videos where Linus Torvalds talks for a couple hours. I swear the next time you open a command line or work within a slipshod hideous UI that I guess just works, or something, you'll somehow feel this uncanny relief because you know Linus.
Doesn't make any sense , Linus Torvalds doesn't work on making command line or UI, I can understand if this was about writing kernel drivers or writing programs calling kernel API , The command line and other UI are much abstracted away from the kernel no?
When I know it's "January" for whatever reason, and not "First Month", without knowing HOW and WHY this is the convention, it sucks. But then you see what Janus looks like - which, in this case, makes it even better for some reason than "First Month." In the case of programming languages and systems, many of these mythological figures are still alive and we can ask them questions.
So, with Linus, I'm approaching it from the individual differences angle. Some people will call IT if "Smooth edges of screen fonts" is turned off. "Why is this like this?? I can't work like this!"
Whereas I don't think Linus would notice, and if he did I doubt he would care. There's something about understanding where a creator of any given thing falls on that continuum that I find really helpful in overcoming friction with the "why is bad thing that could be good, by my likely ignorant definition, so bad!!" when it comes to new languages or systems.
Tangential, but I want to go back in time and ask some ancient Danes why #70 and #90 in Danish simply had to be absolutely ridiculous, given any possible alternatives. It still wouldn't be my preference, but something about understanding their perspective would be helpful.
That's interesting, especially because I can't really relate to that. Don't get me wrong, I love this kind of trivia, but not knowing something like that has never been a problem for me at all.
Is this perhaps related to learning a language as a child versus as an adult?
As a child in an English-speaking household, "January" is just one of a million other things to pick up, and you just pick it up. As an adult learner, I can imagine it's very difficult because it seems so random; it's yet another thing that doesn't fit into a system so you have to memorise it and try to internalize it.
As a native English speaker, I find it very difficult to remember weekday names and month names in other languages. Mnemonics help, and knowing the derivation of the word can make for a good mnemonic.
Edit: re-reading the GP comment, sounds like it's not so much about mnemonics, more that knowing the historical reasons for weird design decisions can make it easier to accept them. Like knowing why "October" isn't actually the eighth month.
But for relatively recent programming language things? Oh man.
std::cout << "why"; std::cout << "why"; std::cout << "why";
// It just rolls off the tongue.
using namespace std;
It's fine, I won't tell anyone.Who are we? There is no we. There is no we against them. That's the thing about diversity. My background and world view is different than yours.
But yeah I can see that understanding the roots of a technology, of the creator's views of the technology and the problems its solving would certainly alter our perception of it and make it more enjoyable.
Not sure if the some people view of the world should impact the way programming languages are constructed.
It's like saying you dislike math proofs because you have different views on the world from Gauss or Hilbert. That a triangle might be a rectangle because Pitagora had an old view of the world which doesn't fit today's fashion and politics.
From a more positive angle, you shouldn't worry about understanding "a very different model of the world" because, being a mathematical model of the world it cannot be much different from yours.
That's why you shouldn't watch interviews. If they're good speakers they'll trick you into adopting their views.
Quick summary:
1. It is unable to indicate International/Unicode characters, code points, or byte values. ISO/IEC 14977:1996 only supports ISO/IEC 646:1991 characters.
2. It is unable to indicate character ranges.
3. It requires a sea of commas, so using it produces hard-to-read grammars.
4. It does not build on widely-used regex notation
5. It has a bizarre, difficult-to-understand, and easily-misunderstood “one or more” notation.
6. It is challenging to understand and many key terms are undefined.
I suggest considering the EBNF notation from the W3C Extensible Markup Language (XML) 1.0 (Fifth Edition). That's much closer to the regex notation used by many software developers today, and since the point of a BNF is to communicate, using a language similar to a language most developers already know is a big advantage.
For me, the breaking point came in trying to describe binary formats. I eventually had to come up with my own: https://github.com/kstenerud/dogma
Will it solve the BNF problem? No, but it does help the binary folks.
There are two ways to resolve this. 1) In a way that is backwards-compatible with BNF, and 2) in a way that isn't.
If you way is backwards-compatible, it's going to be much more easily understandable to anyone who already knows BNF. They don't have to go back and double-check whether some construct means concatenation or an alternative form.
It also means that if someone is working on a grammar that doesn't need one of your new fundamentals and is expressible in plain BNF, they're going to write something that anyone who already knows BNF can read without any extra research.
Also, if your additions are good enough, they provide a good draft of a new way to extend "official" BNF in the future. If you extend in an incompatible way, that's not even a possibility.
What you end up with is a bunch of small innovations in each project, but no incentive to consolidate these gains (because it's so niche to begin with, and there's no significant business advantage to be had).
And of course anyone who tries will have massive pushback ala https://xkcd.com/927
So everyone just does their own thing and we pretend that it's all fine.
From the abstract: "instead of adding another syntactic notation and arguing about its excellence, we propose to retain the diversity and to cope with it by formally defining syntactic notations and using such definitions to import existing grammars to grammar engineering frameworks and to export (pretty-print) existing grammars to any desired syntactic notation."
Recently I read a paper "A Translational BNF Grammar Notation (TBNF)" By Paul B Mann, Parsetec.
"EBNF is powerful, however, it describes only the recognition phase, which is only 1/3 of the process of language translation. The second phase is the construction of an abstract-syntax tree and the third phase is the creation of an instruction code sequence. Some parser generators automate the construction of an AST, but none, that I know of, automate the output of instruction codes."
I'm not entirley sold on the idea but it is interesting to think of what could be different in this space.
IETF's RFC works, but it has its own problems. In particular, most people use regexes constantly, and that RFC is unnecessarily incompatible with widely-used regex format.
I suggest looking at the W3C Extensible Markup Language (XML) 1.0 (Fifth Edition) as a plausible starting point: a href="https://www.w3.org/TR/xml/#sec-notation
IntelliJ Grammar-Kit also uses sort of EBNF (https://plugins.jetbrains.com/plugin/6606-grammar-kit), the file extension is even `.bnf`.
Even though parser generators are often much different than EBNF, I still wish more people followed this standard, and parser generators could create EBNF grammar definitions. Because it's not just for parsing directly: it's a good starting point for creating parsers; a reference to validate and regression test; and it can be used to generate a (sometimes ambiguous) parser which is good for prototypes and toy languages.
Speaking of graphical representations, here's a less compact one, but more intuitive: https://sqlite.org/syntaxdiagrams.html
BNF will probably always be a useful communication tool, but it has always been deficient for describing grammar parsing and it certainly will never be useful for grammar comparison like this paper seems to be looking for.
It works as ABNF in IETF contexts. It works in ASN.1. Therefore it works for SNMP mibs. It can define to the bit level PDUs "on the wire".
It can also in some sense notate the semantics if you intuit above the syntax of the choices of one, any, many as to WHY you are proferred the alternates and what they mean. The problem is that most of the logic about WHY is encoded in comments, or out of scope. It only denotes the forms.
The intersection of BNF and UNIX man page commandline arguments remains probably its most visible context. cat [ -v ]
I learned some variants BNF/EBNF in 1979 to learn DEC-10 instruction set semantics. Pascal, Wirth had written diagrams which morally encoded the same information but in a visual form. It irked me that there was no connect/uplift from (E)BNF into the lex part of lex/yacc functionality on the compiler course. We dived into lexical analysis without discussing a restricted grammer for representation of things which syntax we'd all been taught. I guess its what you know times who you learn it from times what relevance it has.
There is some venn diagram collision in my mind with BNF variants and regex. Perhaps the inevitability of [ and { representative forms, with | for alternates.
I don't write DSL so I don't play in this space but I use BNF almost unconsciously every day, especially if its a 'read that IETF draft' day.
They didn't mention what IMHO might be the most important one: ISO EBNF is different enough from traditional regex notation that it's uncomfortable to read for those familiar with the latter. (At least it's not ABNF, which I think is even worse.)
That said, it seems the majority of grammar notations I've seen use | for choice, ? for zero or one, * for zero or more, + for one or more, and ( ) for grouping, with implicit space-concatenation (higher precedence than choice, by analogy with multiplication and addition). This is essentially standard regex metacharacters, and context-free-grammars are a step above regular grammars, so it's not surprising that such a notation developed.