Rust for tokenising and parsing
xnacly.me
xnacly.me
With that realisation I started looking for another more suitable language - I knew the FP aspects of Rust are what I was looking for so at first I considered something like F# but I didn't like that it's tied to microsoft/.NET. Looking a bit further I could have gone with something like Zig/C but then I lose the FP niceness I'm looking for. I also spent a fair amount of time looking at Go, but eventually decided that 1. I wanted a fair amount of syntax sugar, and 2. golang is a server side language, a lot of its features and library are geared towards this use case.
Finally I found OCaml, what really convinced me was seeing the syntax was like a friendly version of Haskell, or like Rust without lifetimes. In fact the first Rust compiler was written in OCaml, and OCaml is well known in the programming language space. I'm still learning OCaml so I'm not sure I can give a fair review yet, but so far it's exactly what I was looking for.
The rest of the golang ecosystem I found really nice actually, and imo it had a really great set of tools for reading/writing to files - and also I like that everything is apart of the go binary, it certainly is easier than juggling between opam and dune (used for OCaml for example).
The ecosystem and tooling are great, probably the best I've worked with. But the main reason I reach for Go is that it's got tiny mental overhead. There's a handful of language features so it becomes obvious what to use, so you can focus on the actual goal of the project.
There are some warts of course. Heavy IO code can be riddled with err checks (actually, why I find it a bit awkward for servers). Similarly the stdlib is quite verbose when doing file system manipulation, I may try https://github.com/chigopher/pathlib because Python's pathlib is by far my favourite interface.
Do you know how they avoid the GC in the Go implementation of the Go compiler? If I understand correctly they need to implement the Go garbage collector in their Go implementation of the Go compiler. But Go already has a garbage collector. So how do they avoid invoking Go's garbage collector so that they can implement the garbage collector of the Go language they are implementing?
Not sure if I'm making sense but I'd like to know more about this from those who understand this more than I do.
But the shape of the question feels like you're asking about whether an interpreter (which the compiler is not) uses the GC of the host language?
I think the relevant code is https://github.com/golang/go/blob/master/src/runtime/mgc.go and adjacent files. I see some annotations like //go:systemstack, //go:nosplit, //go:nowritebarrier that are probably relevant but I wouldn't know if there's any other specific requirements for that code.
First, that is a problem only for the very first version of X. Then you use X for version X+1.
Second, building from source usually doesn't mean having to build every single dependency. Some .so or .dll are already in the system. Only when one has to build everything from scratch the first step would have to solve the original X from X problem but I think that even a Gentoo full system build doesn't start with a user setting in bytes in RAM with switches (?), setting the program counter of the CPU and its registers to eventually start the bootstrap process.
All this to say: the output of a compiler is by necessity not tied to the language the compiler is written in, instead it is tied to the machine the executable should run on. A compiler "merely" translates instructions from a high level language to a machine executable one. So stuff like a GC must be coded, compiled and then "injected" into the binary so the user's code can interact with it. In an interpreted language this isn't necessary, since the host language is already running and contains these tools which would otherwise have to be injected into the binary.
The compiler executable itself is running in a compilation process P which uses memory and has its own garbage collection. (The compiler executable was itself generated by a compilation, using a compiler written in Go itself(self-hosting) or initially, in another language).
But the compilation process P is unrelated to the process Q in which the generated code, LLC, will run when first executed. The OS which runs LLC doesn't even know about the compiler - LLC is just another binary file. The garbage collection in P doesn't affect garbage collection in Q.
Indeed, it should be easy for the compiler to generate an assembly program which constantly keeps allocating more memory until the system runs out, while compiling say a loop which allocates a struct within a loop running a billion times. Unless, of course, you explicitly also generate a garbage collector as part of the low level code.
Your question does become very interesting in the realm of security, there is a famous paper called "Trusting Trust" where a compiled compiler can still have backdoors even if the compiled code is trustworthy and the compiler code is trustworthy but the code which compiled the compiler had backdoors.
char heap[100000000];
int heap_end;
void *alloc(int n_bytes) {
void *out = &heap[heap_end]
heap_end += n_bytes;
return out;
}
As you can see, it doesn't need to allocate any memory to do this.The garbage collector isn’t part of the compiler, it’s part of the runtime. It’s worth being clear about this distinction because I think it’s the root of the OP’s confusion.
So at some point, someone wrote enough of a Go GC in C to support enough of Go to compile itself.
Why wouldn’t it be able to?
I don’t understand how your question specifically relates to garbage collection, or why the compiler would need to avoid it. The Go compiler is a normal Go program and garbage collection works in it the same way it does in any other Go program.
Go programs do have kinda slower startup times compared to regular dynamically linked C/C++ programs.
I've written parsers professionally with Rust for two companies now. I have to say the issues you had with the borrow checker are just in the beginning. After working with Rust a bit you realize it works miracles for parsers. Especially if you need to do runtime parsing in a network service serving large traffic. There are some good strategies we've found out to keep the borrow checker happy and at the same time writing the fastest possible code to serve our customers.
I highly recommend taking a look how flat vectors for the AST and using typed vector indices work. E.g. you have vector for types as `Vec<Type>` and fields in types as `Vec<(TypeId, Field)>`. Keep these sorted, so you can implement lookups with a binary search, which works quite well with CPU caches and is definitely faster than a hashmap lookup.
The other cool thing with writing parsers with Rust is how there are great high level libraries for things like lexing:
https://crates.io/crates/logos
The cool thing with Logos is it keeps the source data as a string under the surface, and just refers to a specific locations in it. Now use these tokens as a basis for your AST tree, which is all flat data structures and IDs. Simplify the usage with a type:
#[Clone, Copy]
struct Walker<'a, Id> {
pub id: Id,
pub ast: &'a Ast,
}
impl<'a, Id> Walker<'a, Id> {
pub fn walk<T>(self, other_id: T) -> Walker<'a, T> {
Walker { id: other_id, ast: self.ast }
}
}
Now you can specialize these with type aliases: type TypeWalker<'a> = Walker<'a, TypeId>;
And implement methods: impl<'a> TypeWalker<'a> {
fn as_ref(&self) -> &'a Type {
&self.ast[self.id]
}
fn name(&self) -> &'a str {
&self.as_ref().name
}
}
From here you can introduce string interning if needed, it's easy to extend. What I like about this design is how all the IDs and Walkers are Copy, so you can pass them around as you like. There's also no reference counting needed anywhere, so you don't need to play the dance with Arc/Weak.I understand Rust feels hard especially in the beginning. You need to program more like you write C++, but with Rust you are enforced to play safe. I would say an amazing strategy is to first write a prototype with Ocaml, it's really good for that. Then, if you need to be faster, do a rewrite in Rust.
> I've written parsers professionally with Rust for two companies now
If you don't mind me asking, which companies? Or how do you get into this industry within an industry? I'd really love to work on some programming language implementations professionally (although maybe that's just because I've built them non-professionally until now),
> Especially if you need to do runtime parsing in a network service serving large traffic.
I almost expected something like this, it just makes sense with how the language is positioned. I'm not sure if you've been following cloudflare's pingora blogs but I've found them very interesting because of how they are able to really optimise parts of their networking without looking like a fast-inverse-sqrt.
> There's also no reference counting needed anywhere, so you don't need to play the dance with Arc/Weak.
I really like the sound of this, it wasn't necessarily confusing to work with Rc and Weak but more I had to put in a lot of extra thought up front (which is also valuable don't get me wrong).
> I would say an amazing strategy is to first write a prototype with Ocaml, it's really good for that.
Thanks! Maybe then the Rust code I have so far won't be thrown in the bin just yet.
You do not need to write programming languages to need parsers and lexers. My last company was Prisma (https://prisma.io) where we had our own schema definition language, which needed a parser. The first implementation was nested structures and reference counting, which was very buggy and hard to fix. We rewrote it with the index/walker strategy described in my previous comment and got a significant speed boost and the whole codebase became much more stable.
The company I'm working for now is called Grafbase (https://grafbase.com). We aim to be the fastest GraphQL federation platform, which we are in many cases already due to the same design principles. We need to be able to parse GraphQL schemas, and one of our devs wrote a pretty fast library for that (also uses Logos):
https://crates.io/crates/cynic-parser
And we also need to parse and plan the operation for every request. Here, again, the ID-based model works miracles. It's fast and easy to work with.
> I really like the sound of this, it wasn't necessarily confusing to work with Rc and Weak but more I had to put in a lot of extra thought up front (which is also valuable don't get me wrong).
These are suddenly _very annoying_ to work with. If you come from the `Weak` side to a model, you need to upgrade it first (and unwrap), which makes passing references either hard or impossible depending on what you want to do. It's also not great for CPU caches if your data is too nested. Keep everything flat and sorted. In the beginning it's a bit more work and thinking, but it scales much better when your project grows.
> Thanks! Maybe then the Rust code I have so far won't be thrown in the bin just yet.
You're already on the right path if you're interested in Ocaml. Keep going.
> If you come from the `Weak` side to a model, you need to upgrade it first (and unwrap), which makes passing references either hard or impossible depending on what you want to do.
You're literally describing my variable environment, eventually I just said fuggit and added a bunch of unsafe code to the core of it just to move past these issues.
Yeah, that's the focus of it, and the thing you can use Rust well.
All the popular Rust parsing libraries aren't even focused on the use that most people use "parser" to name. They can't support language-parsing at all, but you only discover that after you spent weeks fighting with the type-system to get to the real PL problems.
Rust itself is parsed by a set of specialized libraries that won't generalize to other languages. Everything else is aimed at parsing data structures.
F# or even the latest version of C# are what I would recommend. Yes Microsoft are involved but if you're going to live in a world where you won't touch anything created by evil corporations then you're going to have a hard time. Java, Golang, Python, TypeScript/Javascript and Swift all suffer from this. That leaves you with very little choice.
I'd be interested in hearing your thoughts over OCaml after a year or so of using it. The Haskell-likes are very interesting but Haskell itself has a poor learning curve / benefit ratio for me (Rust is similar there actually; I mastered C# and made heavy use of the type system but that involved going very very deep into some holes and I don't have the time to do that with Rust).
Have I missed something GvR or his team did?
But, I can always just write Rust and be happy where I am. Or, to be honest, would not be very unhappy with F#, Haskell or Ocaml either.
What do you mean, exactly?
Anecdotally I would say where a lot of companies would have used Java in the past they are now turning to go for their server-side/backend service implementations.
and they also will avoid you! A monthly go-lang meetup in San Francisco impressed me as the only meetup I have ever been to where no one (in a crowded venue) seemed to want to talk to anyone outside their clique
I only advocate for it on the scenarios where a garbage collected C is more than enough, regardless of the anti-GC naysayers, e.g. see TamaGo Unikernel.
The purpose of creating yet another language with Go was to break from what everyone else was doing, to see if a "simple" language would stop developers from playing with fun language toys all day to instead focus on actual engineering.
Arguably it was successful in that.
https://commandcenter.blogspot.com/2012/06/less-is-exponenti...
They also got lucky with Docker and Kubernetes, pivoting from Python and Java respectively into Go.
Legend has it that Go was conceived while waiting for a C++ program to compile, if that's what you are thinking of?
The term you are looking for is sum types (albeit in a gimped form in the case of Pascal). Enumerations refer to the value applied to the type, quite literally, and is identical in Pascal as every other language with enumerations, including Go. There is only so much you can do with what is little more than a counter.
Pascal doesn't require case matching of enumerations to be exhaustive, but this can be turned on as a compiler warning in modern Pascal environments, FreePascal / Lazarus and such.
Go only has enums by convention, hence the "iota dance" referred to. I've argued before that this does qualify as "having enums" but just barely.
It wouldn't have been difficult to do a much better job of it, is the thing.
Normally in Pascal you would not match on the enumeration at all, but rather on the sum types.
type
Foo = (Bar, Baz)
case foo:
Bar: ... // Matches type Bar
Baz: ... // Matches type Baz
The only reason for enumerations in Pascal (and other languages with similar constructs) is because under the hood the computer needs a binary representation to identify the type, and an incrementing number (an enum) is a convenient source for an identifier. In a theoretical world where the machine is magic you could have the sum types without enums, but in this reality...Thus, yes, in practice it is possible to go around the type system and get the enumerated value out with Ord(foo), but at that point its just an integer and your chance at exhaustive matching is out the window. It is the type system that allows more flexibility in what the compiler can tell you, not the values generated by the enumeration.
> Go only has enums by convention
"Enums by convention" would be manually typing 1, 2, 3, 4, etc. into the code. Indeed, that too is an enumeration, but not as provided by the language. Go actually has enums as a first-class feature of the language[1]. You even say so yourself later on, so this statement is rather curious. I expect you are confusing enums with sum types again.
[1] Arguably Pascal doesn't even have that, only using enums as an implementation detail to support its sum types. Granted, the difference is inconsequential in practice.
Sum types in Pascal are called variant records:
type
FooKind = (Foo, Bar, Baz); (* An enum *)
FooOrBaz = record (* This is the sum type *)
case foo: FooKind of
Foo: (quux: Double);
Bar: (zot, zap: Double);
Baz: (xyzzy: String);
end
Rust conflates the 'enum' keyword with sum types. Pascal does not do this. One of us is confused about what a sum type is. It isn't me.As for Go, my full opinion on that subject may be found here. If you're... curious. Let's say.
This is the exact same thing, except in addition to the tag there is also a data component. Yes, this is the more traditional representation of sum types, but having an "undefined" data component is still identifiable as a sum type. It is the tag that makes the union a sum type.
In Typescript terms, which I think illustrates this well, it is conceptually the difference between:
{ kind: 0 } | { kind: 1 }
and { kind: 0; data: T } | { kind: 1; data: U }
Which is to say that there is no difference with respect to the discussion here.> They are an enumeration of discrete values
Yes, the "tag" is populated with an enumerator. There is an enumerator involved, that it is certain, but it is outside of the type system as the user sees it. It's just an integer generated to serve as an identifier – an identifier like seen in the above examples – but provided automatically. The additional information you can gain from it, like exhaustive matching, comes at the type level, not the number itself.
> Rust conflates the 'enum' keyword with sum types.
Right, because it too uses an enumerator to generate the tag value. Like Ord(foo) before, you can access the enumerated in Rust with something like
mem::discriminant(&foo)
The spoken usage of 'enum' in Rust is ultimately misplaced, I agree. An enumerator is not a type! But it is not wrong in identifying that an enumerator is involved. It conflates 'enum' only in the very same way you have here.There are innumerable ways you can rectify this lack of understanding if you want to.
1. If you have a pet sum type definition, let it be known. But traditionally, a sum type is better known as a discriminated or tagged union. So far, this is what we understand Pascal offers: A union type that discriminates its subtypes by an enumerated value.
2. The tag that discriminates the type within the union (or whatever your explanation above ends ups calling it) is, in implementation, generated by a process that assigns a number to each entity; something also true of Rust. This is undeniable, as proven by the use of Ord and mem::discriminant. If you have another name for that numberer, if not an enumerator, let it be known.
And even though it has them now, the generics syntax is pretty clunky IMO.
Nothing wrong with that, but it will probably never work for me. Newer versions of Java are much more enjoyable to work with versus Go.
I am one of those. I grok abstractions just fine (have commercially written idiomatically obtuse Scala and C#, some Haskell for fun, etc.), but I don't enjoy them.
I use them, of course (writing everything in raw asm is unproductive for most tasks), but rather than getting that warm fuzzy feeling most programmers seem to get when they finish writing a fancy clever abstraction and it works on the first try, I get it when I look at a piece of code I've written and realize there is nothing extraneous to take away, that it is efficient and readable in the sense of being explicit and clear, rather than hiding all the complexity away in order to look pretty or maximize more abstract concerns (reusability, DRY, etc.).
This mindset is a very good fit for writing compute-heavy numerical code, GPU stuff and lots of systems level code, not so much for being a cog in a large team on enterprise web backends, so I mostly write numerical code for physics simulations. You can write many other things this way and get very fast and bloatfree websites or anything else, but it doesn't work well in large teams or people using "industry best practices". It also makes me prefer C to Rust.
"Perfection is achieved, not when there is nothing more to add, but when there is nothing left to take away."
- Antoine de Saint-Exupery
https://www.brainyquote.com/quotes/antoine_de_saintexupery_1...
The vast majority understand abstractions just fine, though each takes time to understand. However most people like their own abstractions best, and those of other people less. To me hell is living in a world of bad abstractions created by someone else.
Every abstraction created adds to cognitive load when reading the code and to the maintenance burden of that code. So you have an abstraction budget, which is usually in overspent IME and needs to be carefully controlled. Most of the most horrible codebases are horrible because they have too many of the wrong sort of abstraction.
Personally, I don't want to write any new code in something that doesn't have ADTs, or the moral equivalent (Java's sealed classes). I've already written a lifetime of code without them, so I suppose part of that is not wanting to write another 20 years of the same code. :)
I am one of them. I don't like Go, though. Enums and tagged unions aren't abstractions but fundamental features in my book. It's pretty transparent how they look in memory and there's nothing hidden about them.
What does confuse me are things like macros or annotations that magically insert something and make the code incomprehensible. I'm sure it's convenient to use, but it makes my brain try to manually translate it to simple instructions like a foreign language.
In my free time I like using Rust without custom traits (except a few iterators), that's close to the sweet spot for me.
The standard library is unimpressive (to be generous), it has plenty of footguns like C but none of its flexibility.
Also for some reason parenthesis AND \n are required. So you get the worse of C and python there.
I really like this about go - that it formats code for you, and miss it in other languages where we have linters but not formatters, which is a terrible idea IMO.
So being built-in, idiomatic and expected cannot be over-appreciated.
If I write
if x
{
bla
}
it will not compile, because the { needs to be on the same line as the X (for no reason whatsoever).Coming from Python, this is one of the major things that I just can't get past with golang (despite having to use it for work). The standard library has a lot of really interesting/impressive/useful things to cover niche cases, but is missing a lot of what I would consider basic functionality that I keep running into requiring me to go get an external module to solve the problem.
Then, on top of that, the documentation for external modules is extremely terrible. In many cases the best you can get is API documentation in the form of "these are the functions, this is what they take and return" with no explanation of what those values need to be, what the function does with them, and so on; a simple list of functions. In others, there is that plus example code which doesn't work because it hasn't been updated since the last time backwards-incompatible changes were made so you end up down a rabbit hole of trying to debug someone else's wrong code.
The only thing letting me write effective golang at this point is that VSCode can autocomplete a lot of method calls, API calls, and so on, and then tell me what parameters they need, but even then I'm just guessing about what function might exist and what it might be called.
The language itself is okay and the more I use it the more I understand why they implemented all the stuff I hate (like a lack of proper error handling leading to half of my lines of code being boilerplate `if err != nil` blocks), but if the tooling around it wasn't so good no one would take it remotely seriously.
Embarrassing that developers are still forgetting nil pointer checks in 2024.
That said, I discovered that Go has the ability to basically encapsulate one error inside of another with a message; for example, if you get an err because your HTTP call returned a 404, you can pass that up and say "Unable to contact login server: <404 error here>". But then the caller takes that error and says "Could not authenticate user: <login error here>", and _their_ caller returns "Could not complete upload: <authentication error here>" and you end up with a four-line string of breadcrumbs that is ostensibly useful but not very readable.
Python's `raise from` produces a much more readable output since it amounts to much more than just a bunch of strings that force you to follow the stack yourself to figure out where the error was.
This is called (>>=)[0], but most of the industry is ignorant enough to call it impractical and non-pragmatic (as opposed to their pragmatic `if err`)
[0] https://hackage.haskell.org/package/base-4.20.0.1/docs/Prelu...
Is there a term equivalent to "armchair quarterback" in programming? Most programmers are already in armchairs.
It's the equivalent of yelling at the TV that the ultra-successful mega-athlete sucks. I can't imagine the thought process that goes into thinking Ken Thompson, Rob Pike and Robert Griesemers are complete idiots that have no clue of what they were doing.
They made a deliberate decision to design a language that did not take many developments in PL design since the 70's into account.
They had their reasons, which make sense in the context of their employer and their backgrounds.
Many people, myself included, prefer to program with languages that do not focus so much on simplicity
All they knew was C and that they wanted to create a language that compiles faster than C++. That's all.
You're not doing yourself any favors.
I've really tried giving Go a go, but it's truly the only language I find physically revolting. For a language that's supposed to be easy to learn, it made sooooo many weird decisions, seemingly just to be quirky. Every single other C-ish language declares types either as "String thing" or "thing: String". For no sane reason at all, Go went with "thing String". etc. etc.
I GENUINELY believe that 80% of Gos success has nothing to do with the language itself, and everything to do with it being bankrolled and used by a huge company like Google.
Personally I also think that if you removed memory safety overnight from Rust, people will still use it. Rust is appealing not because it's memory safe. For some uses it is, but most people flock to Rust because it offers an alternative to C++ without fifty years of accumulated cruft. Rust is a modern language, with a well working package manager/build tool and a wide ecosystem of libraries for every usecase. Memory safety and other features are just a cherry on the top. If Rust used garbage collection I am sure it would also be very popular just because of those other things.
Other languages like D or Nim tried to fit into that space also, but they don't have the budget to really make it. Most of work on those languages is done by unpaid volunteers, so there's little direction and there's a lot of one man projects.
One thing I've learned over the years is that if you go with the grain — not against it — of a language (or any system, really), the design tends to become apparent quicker. "When in Rome," and so forth. Cultural displeasure tends to disappear if you give the native way an earnest chance rather than resisting it. For example, in the beginning, marking identifiers as public by giving them a capital letter struck me as the ugliest thing ever. I don't mind it now. It's never going to be something I love looking at, but it does have the benefit of making declarations' visibility extremely obvious.
I don't think Go's popularity is due to Google at all. Google the company has never really promoted Go (unlike Microsoft with C# and Sun with Java, for example). Go is still treated as a bastard stepchild in many Google projects such as Protobuf/gRPC, Beam, and Google Cloud. The Go team has never seemed very enthusiastic about PR, either. There was that one big redesign of the Go site, but relatively little after that.
I think Go grew by word of mouth more than anything. Projects like Kubernetes, Prometheus, Traefik, etc. helped a lot. Don't forget that it took years for Go to become popular. It wasn't taken very seriously by many in the beginning. Go was not popular within Google until relatively recently. For many years the only serious thing written in Go internally at Google, as I understand it, was the dl.google.com backend.
let x: usize = 42;And complaining about "thing string" vs. "string thing" seems high on pedantry.
Yes, there are aspects of Go I really dislike, but I find fewer things to dislike in Go compared to things I dislike in all the other languages I have programmed in.
It's a very simple and straightforward language, which I think is why people like it, but it's just a pain to use. It feels like it fights any attempt at using it to do things optimally or quickly.
Do people actually care that much about languages? I mean, we're here writing English, which is a complete dumpster fire. Go is undeniable perfection compared to the horror that is English. Clearly you and I don't care that much about languages.
I expect people like Go because of its tooling (what also saves English), which was a million miles ahead of the pack when it first came out. Granted, everyone else took notice, so the gap has started to narrow.
> It feels like it fights any attempt at using it to do things optimally or quickly.
Serious question: Is that because you are trying to write code in another language with Go syntax? Go unquestionably requires a unique mental model that doesn't transfer from other languages; even those that appear similar on the surface. Because of that, I posit that it is a really hard language to learn. It is easy to get something working, but I mean truly learn it.
While every programming language requires its own mental model, Go seems to take it to another level (before reaching a completely different paradigm). I expect that is because its lack of features prevents you from papering over "misuse" like is possible in other, more featureful languages, so you feel it right away instead of gradually being able build the right mental model.
These are not directly comparable just because we use the same term to describe them.
To your second point, I wholeheartedly disagree. Go is not a difficult language to learn, nor is it particularly unique compared to the type of language it attempts to emulate. In fact I think it’s one of the easiest languages to learn if you already have experience in C-likes because of how obvious it is what it’s trying to do.
I think it is just a bad language. It’s simple to figure out how you need to use it, but it is obnoxious and tedious to do it in that way.
Toggle switches are for giving instructions to a machine. Programming languages are a higher level abstraction over the toggle switches so that the intent of the toggling can be communicated with other people. You don't just write code, run it through the machine, and then throw it away. Other people, and probably even yourself, will read what was written again and again and again. The language is very much for people first and foremost, with the side effect of also being understandable by machine.
> I think it is just a bad language.
It is – nobody is suggesting otherwise – but you didn't answer the question. Are you writing Go with Go syntax, or another language with Go syntax? Perhaps the best way to answer, if it is that you just didn't know how, is to post some sample code that you find to be obnoxious and tedious and we can see if it is that way because of Go, or if it is because you are trying to use patterns from other languages that don't fit the language.
Someone can say the German language is a bad language. But it is not the language that is bad, it is the person's perception.
When they try to evaluate German while thinking in English it is no surprise they consider it sub-standard. Germans, OTOH, are much better equipped to evaluate the German language than those who only know how to think in English.
(Full disclosure; my grandfather was German but I only know how to think in English.)
There would be no logical reason to make up some imagined situation. I am not understanding your reason for mentioning it. What is the thought process behind your thinking there to help me better understand your intent?
Or, if this is just your subtle way of saying that you've never actually written any Go code in your life and are just making this up because you read someone like it elsewhere, then fine. So be it. But that would well and truly not be useful. Given that you are seemingly here in good faith we can be sure that is not what is going on here and look forward to your reasoned response.
I don’t know what your goal is here, but it’s clearly not anything that involves a conversation.
I will say that the error propagation is a pain a lot of the time, but I can appreciate being forced to handle errors everywhere they pop up, explicitly.
Java -> Python -> C++ -> Rust -> Go
I have to say, given this progression going to Rust from C++ was wonderful, and going to Go from Rust was disappointing. I run into serious language issues almost daily. The one I ran into yesterday was that defer's function arguments are evaluated immediately (even if the underlying type is a reference!).
https://go.dev/play/p/zEQ77TIP8Iy
Perhaps with a progression Java -> Go -> Rust moving to rust could feel slow and painful.
Turbo Pascal for education, C as professional lingua franca in mid-90s (manual memory management). C++ was all the rage in late 90s (OOP,STL) . Java got hot around 2003 (GC, canonical concurrency library and memory model). Scala grew in popularity around 2010-2012 (FP for the masses, much less verbosity, mainstream ADTs and pattern matching). Kotlin was cobbled together to have the Scala syntactic sugar without the Haskell-on-the-JVM complexity later.
And then they came up with golang which completely broke with any intellectual tradition and went back to before the Java heyday.
Rust feels like a Scala with pointers so the "C++ => Rust" transition looks analogous to the "Java => Scala" one.
they are all actively in-use.. if gp is earlier in their career, it could all be in last 10 years.
I'd never do a project in go.
Go suitable for networking? Really? With no packed structs and no way to set the endianness?
I remember that famous rant about how Go’s stdlib file api assumes Unix, and doesn’t handle Windows very well.
If you are against “worse is better” like the author, that’s a show stopping design flaw.
If you are for it, you would slap a windows if statement and add a unit test when your product crosses that bridge.
Until you want to do unix things like mmap() and madvise() of course. In which case it assumes an OS without those really basic features.
That's the point. It's a rejection of the keyboard jockeys who become more concerned with the code itself than the problem being solved.
"The key point here is our programmers are Googlers, they’re not researchers. They’re typically, fairly young, fresh out of school, probably learned Java, maybe learned C or C++, probably learned Python. They’re not capable of understanding a brilliant language but we want to use them to build good software. So, the language that we give them has to be easy for them to understand and easy to adopt." - Rob Pike
I suppose in a sense this is rejecting the "keyboard jockeys", but probably not in the way you mean.
You cannot separate the tool used to solve a problem from the problem itself. The choice of tool is a critical and central consideration.
Really I think it's more useful to view it as a better C in the less is more tradition, compared to say C++ and Java, which at the time were pretty horrible. That's my understanding of its origin. It makes sense in that context; it doesn't have pretensions to be a super advanced state of the art language, but more of a workhorse, which as Pike pointed out here could be useful for onboarding junior programmers.
Certain things about it I think have proven really quite useful and I wish other languages would adopt them:
* It's easy to read precisely because the base language is so boring * Programs almost never break on upgrade - this is wonderful * Fewer dependencies, not more * Formatters for code
Lots of little things (struct tags for example) I'm not so keen on but I think it's pretty successful in meeting its rather modest goals.
But Go is nothing at all like C, and it's completely unsuitable for most of the situations where C is used. I'm having trouble even imagining what you're getting at with this comparison. The largest areas of overlap I can think of are "vaguely similar syntax style" and "equally bad and outdated type system". Pretty much everything else of substance is different. Go is GC'd, Go has a runtime, etc.
Your response as it stands is not doing much to fight the "Go enthusiasts are Blub Paradox victims" perception
I'm not so sure about that. Python is a DSL for connecting C functions together. Whereas the biggest criticism of Go (gc, at least, but fair to say it has become synonymous with Go) is around its poor C-interop.
In fact, because of that limitation, Go has developed a bit of a "rewrite those C functions in Go" attitude. While that doesn't indicate that Go is a better C (that's subjective anyway), it does indicate that they are found on the same playground. You're probably not going to rewrite Linux in Go, which might be what you are grasping at, but that's not where C ends. Not even close.
You may have a point about Java. Griesemer was a founding member, after all.
Just as C++ intended to improve C but went in a very different direction and added a lot more. Crucially, it doesn't add inheritance as C++/Java did, which I think is an interesting choice which I find quite pleasing and avoids a lot of horrible architectural decisions and vast inheritance trees.
There are certainly lots of bits I would change having used it a while, but I find it quite useful to work in and far more like C than say Java, Ruby or Python which it has supplanted more than usage of actual C. Not sure one of the goals was to supplant C usage and that was not at all what I meant to imply.
Influenced by the same person who influenced C, at least, but it is nearly a straight up clone of Newsqueak, with a dash of Oberon, combined with a will to be more Zen than Python.
If you follow the bouncing ball of influence through Newsqueak, Limbo, et al. I'm sure you make it to C. But then why stop there? What about B, BCPL, etc.?
But I'm not sure it matters. Go was created to test the theory, not because a theory was proven. It didn't have to be successful. It may be that the studies didn't happen even within Google, although that is our greatest chance. We do know Google actually cares about data, unlike programmers.
That said, since Go was released it seems every other language has tried to copy it with their own twist, so while that may not come from a place of evidence, it would appear that the feeling of increased productivity[1] was felt.
[1] Or something adjacent. Focusing on engineering isn't necessarily about productivity. You can't discount productivity, but it is not the top engineering concern, especially in a place like Google.
What are you referring to here? I would consider myself quite well-apprised of recent developments in PL theory (and practice) and I am struggling to come up with examples matching this description.
Go's main selling point, at least as of 10 years ago, was its green threading system, and even at the time it was substantially inferior to the green threading systems available in BEAM (Erlang, Elixir) or GHC Haskell
- Rust and Zig's built-in language tooling, including dependency management and an opinionated autoformatter.
- Java's ZGC was an obvious response to Go's GC. There's also the realization that the fewer GC knobs there are to tune the better.
- Possibly Java's quest to add value types as well, not 100% sure on the timeline.
- Broad industry trend towards providing easy cross compilation into static binaries.
or gevent, or many other greenlet implementations across languages, 100%
Why do you say it was lifted from Go and not from one of the places Go lifted it from, like the two I mentioned?
> Rust and Zig's built-in language tooling
Maybe my memory is failing here, but I don't recall rust ecosystem tooling ever lagging behind Go's. The opposite, if anything.
> Java's ZGC/value types
Can't comment on this one, could be right, although are any of these things clearly sourced from Go? Not disagreeing, I just don't know the history here.
I'm willing to believe that Java is copying from Go, because it's one of the most lumbering languages in popular use today.
> Broad industry trend towards providing easy cross compilation into static binaries.
I'm pretty sure Go does't get credit for this one either. This has been happening independently across many ecosystems as a natural response to increasing containerization and falling storage costs.
Maybe it is, to the extend that playing with lego blocks authorised by the lego corporation is considered engineering. That, of course, is a silly line to draw for defining the discipline.
> we all have way more fun diving deep into crafting the perfect type in Haskell
What Haskell allows you to do in practice is defining required and missing control flows (for solving your problems) as types[0], and constraining those types with rules that define your business domains the control flows must operate in. That is the actual engineering, as opposed to joining plastic bricks.
[0] https://hackage.haskell.org/package/foldl-1.4.17/docs/Contro...
It allows that in theory, at least, certainly. That's the whole reason for its existence. But in practice, developers become enamoured with the language and start to forget about the problem. If that weren't a problem we'd all be using Haskell, but back in the real world.
Anyway, why do Go fans always reach for the Haskell strawman in discussions like this? Most mainstream languages are not nearly as exotic as Haskell, while also not being intentionally crippled like Go. But for some reason Go fans always want to compare it to Haskell.
Even JavaScript, Python and Java are not allergic to adding modern features like iterator map/filter/etc., do you think those are esoteric ivory tower languages too?
Exactly. Thanks for reiterating.
> Anyway, why do Go fans always reach for the Haskell strawman in discussions like this?
What's a Go fan? Someone who thinks that Go blows? That is as bizarre as becoming enamoured by a language. What leads one to have feelings about a language anyway? It is an impossible to understand concept for me.
> Even JavaScript, Python and Java are not allergic to adding modern features like iterator map/filter/etc.
In what world are patterns from the 1960s "modern"? Do you consider selt belts in cars to also be a modern feature?
Not a single person reading this thread believes that you are a coolly detached rational observer
Responding to an argument by not only pretending not to have an opinion, but pretending not to understand the concept of having an opinion, is a very interesting rhetorical strategy. Not sure it works
All one of them?
> Responding to an argument by not only pretending not to have an opinion
The "argument" has only ever been about sharing of that which we understand as fact. One can present a case for why an apparent fact is not factual – mistakes and misinformation can, indeed, slip in – but what is fact would not rest on one's feelings towards it. The story of Go is the same whether you love it, hate it, ascribe no emotion towards it, or even if you have never heard of it before.
> but pretending not to understand the concept of having an opinion
Feeling and opinion traditionally do not imply the exact same thing – they are different words for a reason – so it is not clear if you misspoke, don't recognize a difference, or if you are trying to change the subject, but I will, for the sake of quality, assume the former. While I can understand feelings in some contexts, like feelings towards people, I have no idea why an inanimate programming language would conjure feelings? That is like developing feelings towards a grain of sand you found on the beach, which I don't understand either.
I am not about to claim nobody develops feelings for inanimate programming languages. Humans are varieitied and can do all sorts of weird and wonderful things, but that does not imply I understand it. Feel free to explain it, though. That's the beauty of not understanding something: You get to learn!
I mean, that's why I also asked last time. Surely you're not one of those anti-education types?
> Not sure it works
Okay, cool. Is this some kind of problem, or why are you mentioning it?
Not sure which languages you define as brilliant but I’ve never seen this pattern anywhere in my career, in any language. Including Rust, which is often offered by Go fans as an example of an esoteric exotic language.
In everywhere I’ve seen, people are focused on solving some engineering problem, not on showing off their brilliance by solving language puzzles.
This defense of Go is REALLY common but as far as I can tell is purely based on myth.
Me neither, but I've never actually tired to look. In honesty, have you? It is not exactly to the faint of heart to study. I suspect only Google is willing to go to the necessary depts. But, if you know of something otherwise, I'm sure we're all interested in the details.
> This defense of Go is REALLY common but as far as I can tell is purely based on myth.
1. Defence? Are you in some kind of fight...? That doesn't make any sense.
2. Myth? Are you confusing the hypothesis on which Go was built with results of experimentation? It very well may be that the hypothesis didn't stand up to experimentation, but what does that have to with our discussion? Furthermore, if you really want to change the subject to talk about that, why have you failed to tell us anything about the details of experimentation?
However, in hindsight after I quit using it I realized that Clipper developers often focused on perfecting code — myself included — whereas FoxPro developers just got shit done for the client/end-user. #fwiw
Unless you take a good look at the language itself, you will remain ignorant.
As ugly and ad-hoc as the language feels, it’s hard to deny that what a lot of people want is just good built-in tooling.
I was going to say that maybe the initial lack of generics helped keep compile times low for go, but OCaml manages to have good compile times and generics, so maybe that depends on the implementation of generics (would love to hear from someone with a better understanding of this).
Rust is designed with the philosophy of zero-cost abstractions. (I don’t like the name, because the cost is never zero, but it is what it is.) The abstractions usually involve a lot of function calls and you need a compiler with aggressive inlining in order to get reasonable performance out of Rust. Usage of generics still results in the same non-virtual calls which can be inlined. But the compiler then has to do a lot of work to evaluate inlining for every instantiation of every generic.
Go is designed with the philosophy of simple abstractions, which may come with a cost. Generics are implemented in a way that means you are still doing a lot of dynamic dispatch. If you need speed in Go, you should be writing the monomorphic code yourself. Generics don’t get instantiated for every single type you use them with. They only get instantiated for every “shape” of type.
So when the generated asm is the same between the abstraction and the non-abstraction version, wheres the cost?
A Rust project's cognitive cost budget comes out of what's left over after the language is done spending. This is true of any language, but many language designers do not discount cognitive costs to zero, which, with the "zero cost abstraction" slogan, Rust explicitly does.
The generated asm isn’t the same.
There’s also a presupposition here that you know what the non-abstracted version would look like. If you don’t know what the non-abstracted version looks like, you can’t do a comparison.
OCaml types are complex enough that monomorphization like Rust or C++ is impossible, so everything is boxed.
Not really, no. It doesn't even contain wrappers for the most common system calls.
Looking at Primeys' comment he actually gave some really interesting suggestions on how to manage this without needing Rc / weak pointers or copying loads of dynamic memory all over the place. Instead you have a flat structure of copy-able elements, giving you better cache locality and a really easy way to work with them.
The more your stuff is held in things that are cheap/free to clone, the less you have to fight the borrow checker… since you can get away with clones!
And for actual interpretation there’s these libraries that can help a lot for memory management with arenas and the like. It’s super specialized stuff but helps to give you perf + usability. Projects like Ruffle use this heavily and it’s nice when you figure out the patterns
Having said that OCaml and Haskell are both languages that will do all of this “for free” with their built in reference counting and GC… I just like the idea of going very fast in Rust
But, it does still run on .NET.
At this point, isn't every major language controlled by one main corporate entity?
Except Python? But Python doesn't have algebraic types, or very complete pattern matching.
My effort has been in adding these features to a front end language that transpiles to an underlying FP language, including but not limited to Rust.
{Runs, ducks, hides. :-D}
Could you explain your thought process when deciding to not use F# because it runs on top of .NET? (both of which are open-source, and .NET is what makes F# fast and usable in almost every domain)
As for the hate - my pet theory is that developers need something like a sacrificial lamb to blame their misfortunes on, and a banner to rally under which often happens to be "against that other group" or "against that competing language", and because .NET is platform that happens to be made by microsoft and is a host for two very powerful multi-paradigm languages causes it to be a point of contention for many. From what I've seen, other languages do not receive so much undeserved hate and here on HN some like Go, Ruby or BEAM family receive copious amount of unjustified praise not rooted in technical merits.
I agree on F#. It changed my C && OO perspective in fantastic ways, but I too can't support anything Microsoft anymore.
But, seeing as OCaml was the basis for F#, I have a question, though:
Does OCaml allow the use of specifically sized integer types?
I seem to remember in my various explorations that OCaml just has a kind of "number" type. If I want a floating point variable, I want a specific 32- or 64- or 128-bit version; same with my ints. I did very much like F# having the ability to specify the size and signedness of my int vars.
Thanks in advance, OCaml folks.
Other options have worse support and weaker tooling, and often not even more open development process (e.g. you can see and contribute to ongoing F# work on Github).
This tired opinion ".net bad because microsoft bad" has zero practical relevance to actually using C# itself and even more so F# and it honestly needs to die out because it borders on mental illness. You can hate microsoft products, I do so too, and still judge a particular piece of techology and the people that work on it on their merits.
The F# folks (including Don Syme) did a fantastic job on the early versions of the language (I used it up to ver 2 or early 3), but I am tired of the corporate engine that funds that ecosystem. I now construct my software within an operating system of a different pedigree. Such considerations are important to me, but thanks for sharing your preference. As for me, I hate nothing or no one, but I am as picky as a poor man can be about whom I choose to rely on for my tools.
As for your opinion on the borderlands of mental illness, I'll contact you at your outlook email address should I seek your opinion about such differently-technical topics. But I was only asking if the OCaml compiler can target specific varieties of ints, as the .NET compiler does.
int is pointer-sized minus one bit, e.g. 31 on 32-bit, 63 on 64-bit
nativeint is pointer-sized, but boxed.
float is 64-bit and boxed.
There's limited support for different number sizes or specifying signedness.
F#, being based upon the .NET compiler and runtime, allows specific targeting of both int size and signedness. I'm surprised that OCaml doesn't also allow that, seeing as its compiler has a long pedigree with respect to both programming contests and real-world hardcore software systems (such as trading desk software). I remember first learning of it around two decades ago as a team used it to win the ICFP contest. I guess that such specificity is just outside of the concerns of the OCaml folks.
First off, the only way to express union types is with runtime reflection. You might as well be coding in Python (but without the convenient syntax sugar).
Second off, “if err != nil” is really terrible in parsers. I’m actually somewhat of a defender of Go’s error handling approach in servers. Sure, it could have used a more convenient syntax. But in servers, I almost never return an error without handling it or adding additional context. The same isn’t true in parser’s though. Almost half of my parser code was error checks that simply wouldn’t exist in other languages.
For Rust, I think the value proposition is if you are also writing a virtual machine or an interpreter, your compiler front end can be written in the same language as your backend. Your other alternatives are C and C++, but then you don’t have sum types. You could write the front end in Ocaml, but then you would have to write the backend and runtime in some other language anyways.
Also, regarding F#. It runs on .NET, and indeed, since the ecosystem and community are very small, you need to rely on .NET (basically C#) libraries. But it's really not "tied" to Microsoft and is open source.
A few notes:
* The AST would, I believe, be much simpler defined as an algebraic data types. It's not like the sqlite grammar is going to randomly grow new nodes that requires the extensibility their convoluted encoding requires. The encoding they uses looks like what someone familiar with OO, but not algebraic data types, would come up with.
* "Macros work different in most languages. However they are used for mostly the same reasons: code deduplication and less repetition." That could be said for any abstraction mechanism. E.g. functions. The defining features of macros is they run at compile-time.
* The work on parser combinators would be a good place to start to see how to structure parsing in a clean way.
The author never claimed to be an experienced programmer. The title of the blog is "Why I love ...". Your notes look fair to me, but calling out inexperience is unnecessary IMO. I love it if someone loves programming. I think that's great. Experience will come.
I don't know the author, so it's useful for me to see in the comments that some people think they are not so experienced.
Doesn't mean I won't respect the author at all, it's great that they write about what they do!
"unnecessary" is the same. Who defines what's necessary? Is Hacker News necessary?
I wrote a compiler in school many years ago, but besides thinking "this project is only one a world class expert or an enthusiastic amateur would attempt", I wasn't immediately sure which I was dealing with.
> calling out inexperience is unnecessary IMO. I love it if someone loves programming. I think that's great.
I'll observe that the commenter did not make the value judgement about inexperience that you appear to think they did.
In the context of the blog post, he wants to generate structure definitions. This is not possible with functions.
[0] https://github.com/ryandv/chesskell/blob/master/src/Chess/Fa...
[1] https://en.wikipedia.org/wiki/Forsyth%E2%80%93Edwards_Notati...
Do you see an obvious reason why a similar approach won't work in Rust? E.g. winnow [1] seems to offer declarative enough style, and there are several more parser combinator libraries in Rust.
data Color = Color
{ r :: Word8
, b :: Word8
, c :: Word8
} deriving Show
hex_primary :: Parser Word8
hex_primary = toWord8 <$> sat isHexDigit <*> sat isHexDigit
where toWord8 a b = read ['0', 'x', a, b]
hex_color :: Parser Color
hex_color = do
_ <- char '#'
Color <$> hex_primary <*> hex_primary <*> hex_primary
Sure, it works in Rust, but it's a pretty far cry from being as simple or legible - there's a lot of extra boilerplate in the Rust.Haskell demonstrates the use of parser combinators very well, but I'd still use parser combinators in another language. Parser combinators are implemented in plenty of languages, including Rust, and actually doing anything with the parsed output becomes a lot easier once you leave the Haskell domain.
Parser combinators in Haskell are exactly simple and legible.
When you think of them as an embedded DSL, Haskell is just a well-suited medium because it allows for very clean syntax: Function application is whitespace, and currying and partial application allows for a high degree of composability simply with parentheses and value bindings.
They're simple and legible, if you just want to know what a combinator does, and you don't need to understand how they work underneath. And you could argue that you don't need to to make practical use of them, just like you don't need to understand how LINQ works in C# to use and value them.
The catch comes with the incomprehensible error messages that are a consequence of the DSL being embedded: Once you forget a partial application, or you put a parenthesis the wrong way, you'll not understand why or where it goes wrong. So no, they're not as easy to work with as they are simple and legible.
It’s slightly longer, but more legible.
You don't need switching to Stack (as other commenters suggest) to have isolated builds and project sandboxes etc. If you want to bootstrap a specific compiler version, a-la nvm/pyenv/opam, use GHCup with Cabal project setup: https://www.haskell.org/ghcup/
Has-kill
$
Has-skill
$
{-# LANGUAGE OverloadedStrings #-}
import Prelude hiding (putStrLn)
import Data.Text (Text, replace)
import Data.Text.IO (putStrLn)
transform :: Text -> Text
transform = replace "k" "sk" . replace "ke" "-ki"
main :: IO ()
main = putStrLn $ transform "Haskell"What is the reason for hiding the putStrLn of Prelude and importing that of Data.Text.IO?
{-# LANGUAGE OverloadedStrings #-}
import qualified Data.Text as T
import qualified Data.Text.IO as T
main :: IO ()
main = T.interact (T.replace "k" "sk" . T.replace "ke" "-ki") sed: -e expression #1, char 5: unterminated `s' command
It's like this: sed 's/find/replace/'Seems like many of them have nothing better to do, probably because of layoffs and statistics and linear algebra and 'predicting the next token' (hee hee) in the input, based on gigantic corpuses of data, masquerading as "AI", and many of those bros were/are worthless anyway.
Just a few days ago, I wrote a FEN "parser" for an experimental quad-bitboard impelementation. It almost wrote itself.
P.S.: I am the author of chessIO on Hackage
Again: not throwing shade. I think this is a place where Rust is genuinely quite strong.
E.g., a context-free rule S ::= abc|aabbcc|aaabbbccc|... can effectively parse a^Nb^Nc^N which is an example of context-sensitive grammar.
This is a simple example, but something like that can be seen in practice. One example is when language allows definition of operators.
So, how does Rust handle that?
{-# LANGUAGE OverloadedStrings #-}
import Data.Attoparsec.Text
import qualified Data.Text as T
type ParseError = String
csgParse :: T.Text -> Either ParseError Int
csgParse = eitherResult . parse parser where
parser = do
as <- many' $ char 'a'
let n = length as
count n $ char 'b'
count n $ char 'c'
char '\n'
return n
ghci> csgParse "aaabbbccc\n"
Right 3You used monadic parser, monadic parsers are known to be able to parse context-sensitive grammars. But, they hide the fact that they are combiinators, implemented with closures beneath them. For example, that "count n $ char 'b'" can be as complex as parsing a set of statements containing expressions with an operator specified (symbol, fixity, precedence) earlier in code.
In Haskell, it is easy - parameterize your expression grammar with operators, apply them, parse text. This will work even with Applicative parsers, even unextended.
But in Rust? I haven't seen how it can be done.
use winnow::combinator::{ repeat };
use winnow::token::take_while;
use winnow::prelude::*;
pub fn csg_parse(input: &mut &str) -> PResult<usize> {
let a: &str = take_while(0.., 'a').parse_next(input)?;
let n: usize = a.len();
repeat(n, "b").parse_next(input)?;
repeat(n, "c").parse_next(input)?;
'\n'.parse_next(input)?;
return Ok(n);
}
#[cfg(test)]
mod test {
use super::*;
#[test]
fn parses_examples() {
assert_eq!(Ok(0), csg_parse(&mut "\n"));
assert_eq!(Ok(3), csg_parse(&mut "aaabbbccc\n"));
assert!(csg_parse(&mut "abcc\n").is_err());
assert!(csg_parse(&mut "abbcc\n").is_err());
assert!(csg_parse(&mut "aabc\n").is_err());
assert!(csg_parse(&mut "aabbc\n").is_err());
assert!(csg_parse(&mut "def\n").is_err());
}
}Okay, good enough.
fn parse_abc(input: &str, n: usize) -> IResult<&str, (Vec<char>, Vec<char>, Vec<char>)> {
let (input, result) = tuple(( many_m_n(n, n, char('a')),
many_m_n(n, n, char('b')),
many_m_n(n, n, char('c'))
))(input)?;
Ok((input, result))
}
It parses (the beginning of) the input, ensuring `n` repetitions of 'a', 'b', and 'c'. Parse errors are reported through the return type, and the remaining characters are returned for the application to deal with as it sees fit.https://play.rust-lang.org/?version=stable&mode=debug&editio...
If you have to specify N, no, it doesn't
In fact, that eBNF only produces the lexer. The parser part is not that impressive either, 120 LoC and quite repetitive https://github.com/gritzko/librdx/blob/master/JSON.c
So, I believe, a parser infrastructure evolves till it only needs eBNF to make a parser. That is the saturation point.
Though I agree that a little code generation and/or macro magic can make C significantly more workable.
Won't the code here:
https://github.com/gritzko/librdx/blob/master/JSON.lex
accept "[" as valid json?
delimiter = OpenObject | CloseObject | OpenArray | CloseArray | Comma | Colon;
primitive = Number | String | Literal;
JSON = ws* ( primitive? ( ws* delimiter ws* primitive? )* ) ws*;
Root = JSON;
(pick zero of everything in JSON except one delimiter...)I usually begin with the RFCs:
https://datatracker.ietf.org/doc/html/rfc4627#autoid-3
I'm not sure one can implement JSON with ragel... I believe ragel can only handle regular languages and JSON is context free.
Educational and elegant approach.
Another great talk about making efficient lexers and parsers is Andrew Kelley's "Practical Data Oriented Design" [2]. Summary: "it explains various strategies one can use to reduce memory footprint of programs while also making the program cache friendly which increase throughput".
--
Coroutines for Go - https://research.swtch.com/coro
The parallelism provided by the goroutines caused races and eventually led to abandoning the design in favor of the lexer storing state in an object, which was a more faithful simulation of a coroutine. Proper coroutines would have avoided the races and been more efficient than goroutines.
I can take my parser combinator library that I use for high-level compiler parsers, and use that same library in a no-std setting and compile it to a micro-controller, and deploy that as a high-performance protocol parser in an embedded environment. Exact same library! Just with fewer String and more &'static str.
So toying around with compilers translates my skill-set rather well into doing embedded protocol parsers.
Curious what the rest of the prior art looks like
I wrote it long time ago and it’s not fully implemented tho
[0] https://github.com/salsa-rs/salsa/blob/e4d36daf2dc4a09600975...
[trying to remind myself how this works because it's been a while]
So it's got macros for defining "union types", which combine a bunch of individual structs into an enum with same-name variants, and implement From and TryFrom to box/unbox the structs in their group's enum
ASTInner is a struct that holds the Any (all possible AST nodes) enum in its `details` field, alongside some other info we want all AST nodes to have
And then AST<TKind> is a struct that holds (1) an RC<ASTInner>, and (2) a PhantomData<TKind>, where TKind is the (hierarchical) type of AST struct that it's known to contain
AST<TKind> can then be:
1. Downcast to a TKind (basically just unboxing it)
2. Upcast to an AST<Any>
3. Recast to a different AST<TKind> (changing the box's PhantomData type but not actually transforming the value). This uses trait implementations (implemented by the macros) to automatically know which parent types it can be "upwardly casted to", and which more-specific types it can try and be casted to
The above three methods also have try_ versions
What this means then is you can write functions against, eg, AST<Expression>. You will have to pass an AST<Expression>, but eg. an AST<BooleanLiteral> can be infallibly recast to an AST<Expression>, but an AST<Any> can only try_recast to AST<Expression> (returning an Option<AST<Expression>>)
Another cool property of this is that there are no dynamic traits, and the only heap pointers are the Rc's between AST nodes (and at the root node). Everything else is enums and concrete structs; the re-casting happens solely with that PhantomType, at the type level, without actually changing any data or even cloning the Rc unless you unbox the details (in downcast())
I worked in this codebase for a while and the dev experience was actually quite nice once I got all this set up. But figuring it out in the first place was a nightmare
I'm wondering now if it would be possible/worthwhile to extract it into a crate
I’m imagining seeing the node! macro used, and seeing the macro definition, but still having a tough time knowing exactly what code is produced.
Do I just use the Example and see what type hints I get from it? Can I hover over it in my IDE and see an expanded version? Do I need to reference the compiled code to be sure?
(I do all my work in JS/TS so I don’t touch any macros; just curious about the workflow here!)
$ cargo expand
And you’ll see the resulting code.Rust is really several languages, ”vanilla” rust, declarative macros and proc macros. Each have a slightly different capability set and different dialect. You get used to working with each in turn over time.
Also unit tests is generally a good playground area to understand the impacts of modifying a macro.
it isn't too bad, although the fewer proc macros in a code base, the better. declarative macros are slightly easier to grok, but much easier to maintain and test. (i feel the same way about opaque codegen in other languages.)
The railroad diagrams are tremendously useful:
https://www.sqlite.org/syntaxdiagrams.html
I don't think the lemon parser generator gets enough credit:
https://sqlite.org/src/doc/trunk/doc/lemon.html
With respect of the choice of the language, any language with Algebraic Data Types would work great. Even Typescript would be great for this.
FWIW I wrote a small introduction to writing parsers by hand in Rust a while ago:
https://www.nhatcher.com/post/a-rustic-invitation-to-parsing...
I was struggling though with the lack of strong typing in the returned parse tree, though I think some improvements have beenade there which I did not have a chance to look into yet
I do still like declarative parsing over imperative, so I wrote https://docs.rs/inpt on top of the regex crate. But Andrew Gallant gets all the credit, the regex crate is overpowered.
this time it got traction. funny how HN works.
https://lrparsing.sourceforge.net/doc/examples/lrparsing-sql...
The grammer contains things you won't have seen before, like Prio(). Think of them as macros. It all gets translated to LR(1) productions which you can ask it to print out. LR(1) productions are simpler than EBNF. They look like:
symbol1 := symbol2 symbol3
symbol1 := symbol4 symbol3
symbol3 := token1 symbol2 token2
...
Documentation on what the macros do, and how to get it to spit out the LR1(1) productions is here:https://lrparsing.sourceforge.net/doc/html/
It was used to do a similar task the OP is attempting.
Edit: Never mind. I see it right there under the parser. Thanks!
--
1: https://github.com/antlr/grammars-v4/tree/master/sql/sqlite
Also I'm collecting several LALR(1) grammars here https://mingodad.github.io/parsertl-playground/playground/ that is an Yacc/Lex compatible online editor/interpreter that can generate EBNF for railroad diagram, SQL, C++ from the grammars, select "SQLite3 parser (partially working)" from "Examples" then click "Parse" to see the parse tree for the content in "Input source".
I also created https://mingodad.github.io/plgh/json2ebnf.html to have a unified view of tree-sitter grammars and https://mingodad.github.io/lua-wasm-playground/ where there is an Lua script to generate an alternative EBNF to write tree-sitter grammars that can later be converted to the standard "grammar.js".
Me: "How can a programming language be so damn complex? Am I just dumb?"
For example, the Korean alphabet is pretty simple, and the spelling is simpler than English, but to English speakers it can look intimidating, like hundreds of blocks of unrecognizable squiggles.
The article (over)uses macros, which make the code look more convoluted than it really is. The macro-by-example syntax is conceptually simple, but it has a punctuation-heavy syntax that may look alien if you don't know what you're looking at. BTW, this macro system has been originally designed for JavaScript, which is why even in Rust it looks odd by Rust's standards.
Much as you might anticipate (although perhaps its designer Sean Baxter did not) this was not kindly looked upon by many C++ programmers and members of WG21 (the C++ committee)
The larger thing that "Safe C++" and the reaction to it misses is that Rust's boon is its Culture. The "Safe C++" proposal gives C++ a potential safety technology but does not and cannot gift it the accompanying Safety Culture. Government programmes to demand safety will be most effective - just as with other types of safety - if they deliver an improved culture not just technological change.
But more importantly, Safe C++ is just not a thing yet. People seem to discount the herculean effort that was required to properly implement the borrow checker, the thousands of little problems that needed to be solved for it to be sound, not to mention a few really, really hard problems, like variance, lifetimes in higher-kinded trait bounds, generic associated types, and how lifetimes interact with a Hindley-Milner type system in general.
Not trying to discount Safe C++'s efforts of course. I really hope they, too, succeed. I also hope they manage to find a syntax that's less... what it is now.
In K&R C this very spartan type system makes some sense, there's no resources, you're on a tiny Unix machine, you'd otherwise be grateful for an assembler. In C++ it does look kinda silly, like an SUV with a lawnmower engine. Or one of those very complicated looking board games which turns out to just be Snakes and Ladders with more steps.
But I don't think Safe C++ fixes that anyhow.
† Technically maybe the C pointer types are not just the integers wearing a funny hat. That's one of many unresolved soundness bugs in the language, hence ISO/IEC DTS 6010 (which will some day become a TR)
For C++, it'll be about cramming lifetimes into diamond-inheritance OOP, which... feels even harder.
Safe C sounds like a much, much more believable project, if such a proposal were to exist.