A decade of developing a programming language
yorickpeterse.com
yorickpeterse.com
I bought _Practical Foundations for Programming Languages_ by Harper and _Types and Programming Languages_ by Pierce and I just can't get through the first few pages of either of them. I would love to see a book as gentle and fun as _Crafting Interpreters_ but about making a static ML like language without starting with hardcore theory. (Bob, if you're listening, please make a sequel!)
https://mukulrathi.com/create-your-own-programming-language/...
https://jaked.org/blog/2021-09-07-Reconstructing-TypeScript-...
I read through the first 10 chapters of TAPL, and skimmed the rest. The first 10 chapters were good to remind myself of the framing. But as far as I can tell, all the stuff I care about is stuffed into one chapter (chapter 11 I think), and the rest isn't that relevant (type inference stuff that is not mainstream AFAIK)
This is also good:
https://github.com/golang/example/blob/master/gotypes/README...
And yeah some of us had the same conversation on Reddit -- somebody needs to make a Crafting Interpreters for type checking :) Preferably with OOP and functional and nominal/structural systems.
---
Also, it dawned on me that what makes TAPL incredibly difficult to read is that it lacks example PROGRAMS.
It has the type checkers for languages, but no programs that pass and fail the type checker. You are left to kind of imagine what the language looks like from the definition of its type checker !! Look at chapter 10 for example.
I mean I get that this is a math book, but there does seem to be a big hole in PL textbooks / literature.
Also I was kinda shocked that the Dragon Book doesn't contain a type checker. For some reason I thought it would -- doesn't everyone say it's the authoritative compiler textbook? And IIRC there are like 10 pages on type checking out of ~500 or more.
It might not have the same lingo as the modern type checking specification but most of the ideas are there. It talks about type systems, type expressions, type rules, constructed types like array/struct, type variables for unknown types, etc. It puts the type rules in the grammar productions. It stores the type info in the AST nodes. It uses the chained symbol tables as the type environments.
It has sections on type conversion, operator/function overloading, polymorphic functions (parameterized types), and type unification.
It has all the pieces and steps to build a type checker.
https://www.amazon.com/Compilers-Principles-Techniques-Tools...
It has 11 sub-sections, 2 of which are about type checking
6.3 - Types and Declarations, pages 370-378
6.5 - Type Checking - pages 386 - 398
So that's about 20 pages out of 993 pages. A fraction of a chapter!
---
What I was referring to is the appendix "A Complete Front End", which contains a lexer and parser, but not a type checker! Doesn't sound complete to me.
I guess they have a pass to create the three-address code, and type checking is only in service of that. But they don't give you the source code!
---
Also weirdly, when you look at the intro chapter "The structure of a compiler", it goes
1.2.2 Syntax Analysis
1.2.3 Semantic Analysis (3 paragraphs that mentions type checking and coercions)
1.2.4 Intermediate Code Generation
But when you look at the chapters, there's NO semantic analysis chapter! It goes from (4) Syntax analysis, (5) Syntax-directed translation, to (6) intermediate code generation.
I think there's a missing chapter! (or two)
WTF. The 2nd edition is really butchered.
What the heck happened ...
The second edition cut a bunch of C++ material out.
And yeah I don't think I missed anything -- I think it even had better diagrams.
So yeah I believe that general rule! I didn't realize that at the time I bought the Dragon Book.
Pretty sad how you can lose knowledge over time!
---
The books also respond to fashion -- e.g. Java was a popular teaching language, so they had to add more Java. I think that is reasonable in some ways but can backfire in the long term.
Although I guess I am not too excited about MIX assembly in TAOCP and so forth, so it's a hard problem
Books also tend to become more sparse on the page, and with more pointless “illustrative” stock photos that add nothing and take up page real estate (leading to more advanced topics being dropped to reclaim space), but make the book seem more approachable to undergraduates.
Generally applies more to physics, chemistry, and math books than computer science, but it is a general phenomenon.
Simon Peyton Jones together with others wrote a relatively easy to read book about the writing of Miranda (Haskell), which even includes a simple explanation of lambda calculus, so it includes everything you need to know about the theory, more or less: "The Implementation of Functional Programming Languages" 1987, https://www.microsoft.com/en-us/research/wp-content/uploads/...
I don't think an ML-like language is the best option. I would start with type-checking something simple like C or Pascal: no overloading, generics, type inference, anonymous functions. Just plain old subroutines and declared types on everything, arrays and structs as the only composite data structures.
Then you could add implicit conversions, type inference for locals (not that hard), subtyping, target type inference for anonymous functions, generics and maybe overloads if you're brave enough.
EDIT: ChatGPT could actually demonstrate the whole process for type checking small chunks of code, including: assigning them type variables, collecting constraints and then unification. It would even then point out if there was a type error in the code snippet!
https://gilmi.me/blog/tags/type%20inference
If you find that helpful, I've made more content on compilers (including live coding a compiler, without narration) which should be easily reachable from my website.
Edit: by "screwing up semantics" I mostly mean the combination of overloading and implicit conversions, which is known to cause issues in Java, C++, etc.
Basically, the existing material is a bunch of existing incredibly complicated implementations, the odd blog post that just throws a bunch of code your way without really explaining the why/thought process behind it, and books that aren't worth the money.
The result is that you can of course piece things together (as I did), but it leaves you forever wondering whether you did it in a sensible way, or if you constructed some weird monstrosity.
To put it differently: you can probably build a garden by digging some holes and throwing a few plants around, but without the right resources it can be difficult to determine what the impact of your approach may be, and whether there are better ways of going about it. Oh and I'm aware there are resources on gardening, it's just a metaphor :)
The Dragon book has a chapter on type checking. It gives explanations on many topics. It has plenty of examples. It has type treatments on different areas of a language, like expression, array, struct, function, pointer, etc.
Despite being really old, its ideas and explanation on the topic are still relevant.
So that leaves blog posts for most developers actually implementing this stuff.
But the solution here is that when you finally figure out a good strategy to deal with something muddy like this is to write that better blog post :)
As for _when_ I'll do that, that depends on when I can convince my inner critic to actually commit to the idea :)
[1]: https://github.com/yorickpeterse/pattern-matching-in-rust
- Generics, especially inferring generic type params
- Recursive types
- Deeply mutable vs immutable types
- Type refinement based on checks/usage
- Giving good error messages for deep type mismatches (though I'm not sure TypeScript itself has figured this one out yet, lol)
etc. And even the stuff I've figured out, I figured out almost all from scratch (sometimes having to bail and take a totally new approach midway through). I would love to read a book giving me a proper bottom-up walkthrough of the current landscape of type system implementation
But that's an exception, and deliberately decided to be complex.
If you want ML-style total type inference when you expect the compiler to deduce the parameter and the return types of your functions, then yes.
Local variable type inference isn't hard at all, unless you want to do stuff like Rust: declare a variable without an initial value and infer its type from the first assignment downcode.
So this is actually harder to do that it might seem. The best advice I have is to contribute to a compiler or two and understand their semantic checking model.
The book Applicative Order Programming: The Standard ML Perspective has you implement an ML-like language at the end (the book is mostly about programming in Standard ML itself). Chapter 10 covers type checking and type inference. Chapter 11 covers interpretation via graph reduction. Chapter 12 covers compilation of the ML-like language to an abstract stack machine. The code is all in Standard ML. It's very direct and practical without getting bogged down in theory, although it does talk about theory, and it's not as gentle and fun as Crafting Interpreters.
It was published in 1991. It's on libgen; I'd recommend downloading a PDF and reading those chapters in order to judge for yourself.
I left some links/references to books that helped me out.
I think the most important part is figuring out who is going to use the PL and what problems they are solving. As precisely as possible, best with a list of concrete Personas and Projects. This creates a "value system", which makes it easier to answer all questions during implementation.
Absolutely, just look at Haskell as an example.
> most important part is figuring out who is going to use the PL and what problems they are solving.
Completely agree. Rust and Go have received a lot of criticism over the years, but I think that is often down to them not being as general purpose as many people try to claim. The creators had particular users and use cases in mind.
Is that the claim actually that they are general purpose, or do people not understand the claims made? Go has been very explicit that it is designed for systems programming. The Rust camp seems to hold a similar position. "Systems" implies that it is not designed for all tasks. But I'm not sure "systems" is properly understood by a lot of people.
It is. Language features generally aren't isolated modules that you can add and remove freely: they interact, and "the good parts" are only good because they play nicely with everything else. You can't have Lisp macros without Lisp (or otherwise homoiconic) syntax, or Haskell's concurrency without its effect tracking, or global type inference with Java-style inheritance.
Type hints don't provide any safety, though. That was never the goal, given that they're strictly optional and don't really exist at runtime anyway (though I have written some experimental monstrosities that used annotations for code generation at runtime). They're documentation in a standard form that static analysis tools can leverage.
I really can't imagine a situation where having type hints in Python is worse than simply not having them. They're not the worst of both worlds, they're a compromise with some of the benefits and drawbacks of each.
Can’t they lie / go out of sync with what’s actually happening? IMO, an unenforced type hint is very dangerous because of this - it can give you a false sense of confidence in a scenario where you would otherwise write more defensive code.
Defensive code is good, but I'm absolutely sick of writing code that has to be defensive about the type of absolutely every variable everywhere. It's ridiculous, verbose, and nearly impossible to do universally. THAT is the worst of both worlds. Having to manually fret about the types that every variable may hold. At the point that you're having to write defensive code due to dynamic types, you've already lost the advantages of dynamic types entirely.
I use type hints to say "use and expect these types or get Undefined Behavior", and that contract is good enough for me.
This is demonstrably untrue with the default configuration of Mypy. It will silently ignore explicit hints, inline or not, in certain situations.
If it's demonstrably untrue, you should be able to demonstrate it, right?
https://paste.sr.ht/~chiefnoah/7e07a961cf266fa620d1fd2d31ba2...
This particular issue will get picked up with `--strict`, but it's nearly impossible to do that on a large codebase with typing added post-hoc.
Pyright has saner defaults, it catches this particular issue: https://paste.sr.ht/~chiefnoah/80816fded2a08a03ca80804d524ee...
I agree, though. Pyright has better defaults. It's not great that mypy just completely skips all functions without typing information given. It makes a lot more sense for it to just check everything and assume Any on all unspecified arguments and returns. It's still valuable to type-check the contents of functions with unspecified types.
I used to think this, but based on experience I'm now less convinced. Finding most bugs, like real static typing does, is great; you can significantly reduce your test coverage and iterate with more confidence. Finding a few bugs is pretty useless if you're not finding enough to actually change your workflow.
It's not about workflow or finding enough bugs, but finding bugs that you might not have otherwise seen can be monumentally beneficial.
There's a kind of excluded middle here though. Either that kind of bug hits production often enough to matter - in which case a checker that catches it sometimes isn't good enough, you need a checker that eliminates it completely. Or it doesn't hit production often enough to matter, in which case a checker is of limited use.
Writing type annotations in comments for an optional preprocessor to pass judgement on before the actual implementation ignores them is a bad thing. But that's on python, not on gradual typing.
(An any type + static is also fine if you have a way to retrieve type information at runtime, and there's a sense in which template instantiations of a generic function and a function taking arguments of type any are the same thing as each other)
oh, clojure also is a little bit weird being a sort of grab-bag collection of data structures that it inherits from java and then turns lisp-ish. doesn't make it bad, each thing they add is nice, just feels a little motley
also, i refuse to call it "closure", i pronounce the j
Wait, it's not "clodger"?
i say it's the largest lisp like community out there tbh.
Wow so edgy
There's a certain dearth of pluggable code generators and linkers. Well, not on GNU/Linux, where you get both as(1) and ld(1) practically out of the box, but making your compiler emit a PE/COFF on Windows is a pain. You either bring your own homemade codegen and linker or use LLVM, using Microsoft's ml64.exe and link.exe is incredibly impractical.
I've drafted up a language called "Moth" that has a lot of "meta power" to program block scope any way you want, as most languages hard-wire scoping rules, which I find limiting. Things like "classes", "functions", while-loops etc. would be defined by libraries, NOT the language. It's like lambda's on steroids and without bloated arrow syntax. But it may run slow as molasses, as scoping meta power adds lots of compiler/interpreter indirection.
However, it may turn out that only a few scoping rules are practical in most cases, and the compiler could then optimize for those. It would then only be slow if you are doing something "weird" with scope. Stick to a fixed known set, and things zip along. But finding that set requires R&D and road testing.
Yeah. There's always one little thing or another that bothers me in every language I've learned. Guess I've just come full circle now that I've finally made my own. No doubt it will bother someone else too. If anyone ever uses it.
Essentially, you have a lambda-like combiner called an operative which does not reduce its operands, and implicitly receives a reference to the caller's dynamic environment as an additional argument. The operative body can then optionally evaluate anything as if it were the caller, with even the ability to mutate the caller's local scope. The parent scopes of the caller are accessible for evaluation, but not mutation. You can only mutate an environment for which you have a direct reference - and the environments themselves are first-class values which you can pass around and store in other environments. It is only possible to mutate the parent environment of a caller if the caller's caller has passed a reference to its own environment, and the caller forwards that reference explicitly.
You can create empty environments or environments from an initial set of bindings, and build on them from there, and you can also isolate the static environment from a callee, ensuring that only a limited set of bindings are accessible to the callee. In effect, this enables a kind of "sandboxing" approach - where you can easily do things like load an external plugin which can be written using a limited subset of Kernel features, and only access the specific functions of your host program which you allow. It provides a lot of abstractive power, with the ability to have things like types and classes defined as libraries. The operatives can introduce "sub-languages", which can simulate the behavior of other language semantics.
As you suspect, this runs pretty slow because it's almost impossible to optimize ahead-of-time. A given piece of code is just a tree of symbols and self-evaluating atoms, and the symbols in some code are only given any functional meaning by the environment it is evaluated in. Kernel busts the myth that "compilation vs interpretation is just an implementation choice," and explores what can happen when we go all-in on interpretation. The author has written about the nature of interpreted programming languages and how this differs from compiled languages.[2]
There are opportunities to have compilation and performance without much loss of abstractive power, but this is not a focus of Kernel - whose design goals are described at detail in the report. I've done a lot of R&D on making performance acceptable for a Kernel-like language. The main insight is that you should be able to make some assumptions about the environment passed to an operative without knowing anything else about the environment. I do this by modelling environments as row polymorphic types, where the operative specifies a set of bindings it expects the caller's dynamic environment to contain, complete with type signatures, and assumes these bindings behave in some specific way. (Eg, that the binding `+` means addition, which is in no way guaranteed by Kernel, and that the type signature of `+` is `Number, Number -> Number`).
---
A mostly complete implementation of Kernel: https://github.com/dbohdan/klisp (Cloned from Andres Navarro's bitbucket repo which is no longer available).
Some performance improvements of klisp, with a lot of hand-written 32-bit x86 assembly: https://github.com/ghosthamlet/bronze-age-lisp (Cloned from Oto Havle's bitbucket repo, also no longer available).
Another implementation of Kernel, with some examples of defining records, classes, objects, generators etc as libraries: https://github.com/vito/hummus/tree/master
---
[1]:http://web.cs.wpi.edu/%7Ejshutt/kernel.html
[2]:https://fexpr.blogspot.com/2016/08/interpreted-programming-l...
> This meant I was able to experiment with the semantics and virtual machine of the language, instead of worrying over what keyword to use for function definitions.
Yeah. I see a lot of people asking why people are so fascinated by lisp and why there are so many lisps out there. I think this is a huge reason. It certainly was for me.
I just wanted to get some ideas working as soon as possible. Lisp is the easiest language to parse that I've ever seen, managed to write a parser by hand. And yet it's a fully featured programming language. It just gives you huge power for very low effort. It took a single bit to add metaprogramming to my lisp:
if (function.flags.evaluate_arguments) {
arguments = evaluate_all(interpreter, environment, arguments);
}
I see it as a little frontend for my data structures. I get to work on them endlessly and everything I do improves something.I've written a dozen or so parsers for different languages by hand: the parser is easy regardless of the syntax you choose. The complicated bit is always compilation/interpretation.
The biggest value I see to using S-expressions isn't the time saved in writing a parser, it's reducing the amount of bikeshedding you'll be tempted to do fiddling with syntax.
All programming languages with identifiers are context-sensitive if you put well-formedness into the grammar. C is not exactly difficult to parse, rather it just needs some specific approach compared to most other languages. The biggest (and technically the only [1]) ambiguity is a confusion between type-specifier and primary-expression, which is a roundabout way to say that `(A) * B` is either a multiplication or a dereference-then-cast expression depending on what `A` is in the current scope. But it also has a typical solution that works for most parsers: parser keeps a stack of scopes with contained identifiers, and lexer will distinguish type identifiers from other identifiers with that information. This approach is so popular that even has a name "semantic feedback".
[1] As an example, the notorious pointer and array declaration syntax is actually an unambiguous context-free grammar. It is notorious only because we humans can't easily read one.
> IIRC much of Go's syntax was designed to avoid parsing complexity.
Go's semantics (in particular, name resolution and module system) was designed to avoid complexity. Its syntax is quite normal, modulo personal and subjective bits which all languages have. I can't see any particular mention of syntax decision to performance in the initial spec document [2].
[2] https://github.com/golang/go/blob/18c5b488a3b2e218c0e0cf2a7d...
It will not be obvious that what the C code is doing is C interpretation, because all those data structures don't resemble the C syntax.
The Lisp meta-circular interpreter side-steps that; the issue is settled elsewhere. It exists against a backdrop where the easy correspondence between the printed syntax and the data structure being handled by the interpreter is taken for granted.
The C would need the same kind of backdrop; and that would be a lot of documentation.
The real problem is the preprocessor, it's very hard to implement, and implement in a standards-compliant way. (haven't finished one either, but everybody says so). And without preprocessor and proper include processing, you simply can't parse correctly because you lack the context of which types are defined (besides missing preprocessor symbols).
Another problem that comes with that is that you always have to parse everything from beginning to end -- including the include files, which can easily be 10s or 100s of thousands of lines of code. I haven't seen a real performant and always correct parser for C IDEs, not sure if one exists.
Not sure this is so problematic anymore.
The Zig folks targeted WASM so that they can bootstrap without needing a second compiler implementation. Compilers don't need a lot of POSIX in order to be functional.
The fact that Zig did a boatload of extra work due to path dependence is true.
However, this in no way contradicts that self-hosting via WASM/WASI is likely a far better idea than maintaining two complete compilers in different codebases and languages solely in order to bootstrap.
In fact, nowadays, someone building a new language and compiler is probably better off targeting WASM/WASI before any other architecture.
Welp, video: https://www.youtube.com/watch?v=MCfD7aIl-_E
The Elixir team didn’t get that memo because they are actively in the process of researching and working on a gradual type implementation.[1]
[1] https://elixir-lang.org/blog/2023/09/20/strong-arrows-gradua...
To put it in other terms: my informal impression is that the dynamic "features" of TypeScript are used grudgingly; the community strongly pushes towards strictness, eliminating anys, preferring well-typed libraries, and so on. There's little appetite to - in a single project, unless absolutely required - mix-and-match dynamism with staticness, which is the thing that gradual typing gets you. Rather, it feels like we're migrating from dynamic to static and gradual typing is just how we're doing the migration. But in the case of a new language, why not just start at the destination?
I feel for the people trying to develop these gradual systems but it's truly a herculean task, especially in the Python community that is now understandably so extremely shy about major breaking changes.
This approach is comfortable to me both in Erlang and in Common Lisp, I see it as a balance between safety/performances and development speed (and I'm saying that as someone using Go for all professional development and being really happy with its full static typing).
Most static languages don't make you type every single variable anymore. Java, C++, Rust, C#, and many others let you make the compiler infer types where reasonably possible. That's still full static typing.
My Python and Rust have about the same kinds of explicit type annotations in roughly the same places. My C++ has a little bit more, just because `Foo obj{a}` is more idiomatic than `auto foo = Foo{a}`.
> Recommendation: either make your language statically typed or dynamically typed (preferably statically typed, but that's a different topic), as gradual typing just doesn't make sense for new languages.
The "for new languages" part is really important. Gradual typing makes a lot of sense when you are trying to retrofit some amount of static checking onto an existing enormous corpus of dynamically typed code. That's the case for TypeScript with JavaScript and Elixir with Elixir and Erlang.
IOW, Raku is a gradual type of language.
I think the macros make it even harder. Elixir appears to be done almost all with macros -- there are lots of little "compilers" to Erlang/BEAM.
I think gradual typing is a nice concept, and his arguments against it amount to "gradual typing is not stating typing", which is like the whole point. E.g. he goes on how the compiler can't do some optimizations on functions using gradual typing, but, well, it isn't supposed to anyway.
The benefit of gradual typing is that you can make your program fully dynamic (e.g. for quick exploration), and if you want more assurances and optimizations make it fully static, or if you just want that for specific parts, do them static. And you have the option to go for any of those 3 things from the start.
Having worked with both dynamically and statically typed languages extensively, I never felt I was _more_ productive in a dynamically (or gradually) typed language compared to one that was just statically typed. For very basic programs you may spend a bit more time typing in a statically typed language due to having to add type annotations, but typing isn't what I spend most of my time on, so it's not a big deal.
In addition, that work you need to (potentially) pay upfront will help you a lot in the long term, so I suspect that for anything but the most basic programs static typing leads to better productivity over time.
Neither has the opposite, so there's that.
There's an ACM paper too ("An Experiment About Static and Dynamic Type Systems", 2010) which found higher productivity for the same code quality with dynamic typing. Of course a few papers here and there, pro or against, are as good as none. It's hardly a well studied area.
Besides, one or the other proven better doesn't mean much, just like you liking eggs over easy vs scrambled eggs doesn't depend on some study. If it works for you, and you're more productive with dynamic typing or static typing, use that.
Even if a study "proved" that one kind is "more productive" based on some statistics from measuring some group, or that it has "less bugs" most people when given a choice would still use what they prefer and makes them, as individuals, more productive and happy coding.
>Having worked with both dynamically and statically typed languages extensively, I never felt I was _more_ productive in a dynamically (or gradually) typed language compared to one that was just statically typed.
Depends on the type of program, the type of programming (e.g. imperative/declarive/functional/logical/OO or some combination and so on), the program's scale, the team size, and other aspects, including individual aptitude and preference, not to mention the language and its semantics beyond dynamic/static (and the ecosystem too). I'd certainly be way more producting using dynamic numpy than some C equivalent, even if as a lib it had feature parity.
>In addition, that work you need to (potentially) pay upfront will help you a lot in the long term, so I suspect that for anything but the most basic programs static typing leads to better productivity over time.
There are problems where upfront work is not a benefit, e.g. if it means getting behind in building your MVP or getting behind a competitor adding new features faster while you "perfect it", and your early stage startup loses steam. Also for things where the overhead of upfront might put you off from even attempting them. It can also be a problem to have big upfront costs for exploratory programming and searching into the problem space for your design/solution.
This lets you use dynamic typing when you want, and enables more seamless interoperability with dynamically-typed languages.
I have opinions about Elixir, mostly around aesthetics (for context, I started programming in Python, not Ruby), but that doesn’t mean my opinions are objectively more valid than those of a genius like José Valim.
I’ve actually found Elixir to be very well-designed and internally consistent, and after two years using it, it’s obvious to me that José and team are very thoughtful and deliberative. I don’t think they want to introduce gradual typing because it’s trendy.
Depends on how you implement gradual typing. There's a spectrum [1] of gradual languages, and the differences between, say TypeScript and Reticulated Python matter a great deal.
It is true that you can have some serious performance hits on the migration from dynamically typed -> fully typed code; my advisor Ben Greenman is doing work examining that space of performance en route to fully-typed systems.
In practice, there are several gradually typed languages that give you very good performance when all the code is typed—i.e. you don't have to pay a cost just for having dynamic in your language. You might have to pay a cost when you use it, and the cost can vary dramatically.
Note that's just performance—sometimes performance matters less than being able to prototype something fast in a language, which is a feature I find valuable.
> In fact, the few places where dynamic typing was used in the standard library was due to the type system not being powerful enough to provide a better alternative.
Hey, that sounds like either a win for gradual typing, or a sign that your type system needs some serious work! ;-)
[1]: https://prl.khoury.northeastern.edu/blog/2018/10/06/a-spectr...
I think one way to simplify language creation is to use an existing language with somewhat similar operational semantics as a compilation target. This simplifies the backend a lot and leaves more time to explore what the language (frontend) should look like. The backend can always be rewritten at a later time. My personal choice is usually JavaScript[1].
Regarding type checkers/type inference, I've also ran into difficulties with this topic, and I've written several articles trying to make it more approachable[2].
But on the other hand, we are currently facing certain challenges. KCL is somewhat ambiguous in grammar and semantics e.g. the record type, and we are working hard to improve it.
Is that because of type inference?
Is not a correct statement. You can use gradual typing with explicit type annotations or via inference, it makes no difference to the concept of gradual typing. Gradual typing itself is a way of handling the case of a language being both statically and dynamically typed and handling the interaction between the two portions.
Ok, if we're taking "1 is an int" as type inferencing then, yes, every statically typed language, at least every mainstream one, has at least a small amount of type inference since we don't, in C for instance, have to annotate values. But C does not infer the types of its variables or functions, nor do you have to annotate them in every location where they are used. But that's not inference either, that's using the information determined by annotating the variables.
----------
My main point was that it's weird to say the entirety of gradual typing works by type inference. It works by permitting mixing statically typed and dynamically typed code together in one language. Whether the statically typed portion uses type inference or more "classical" type annotations is orthogonal to the way it works under the hood.
There are always language specific features that may reduce a given problems implementation complexity, but it often depends how much time people are willing to commit to "reinventing the wheel". If you find yourself struggling with support scaffolding issues instead of the core challenge, than chances are you are using the wrong tool.
I am not suggesting Erlang/Elixir, Lisp and Julia are perfect... but at least one is not trying to build a castle out of grains of sand, The only groups I see freeing themselves of C/C++ library inertia is the Go and Julia communities to a lesser extent.
Have a wonderful day, =)
Nonetheless, it is a very welcome addition to the low-level landscape with some cool ideas.
Even if a language doesn't take off, its ideas can make a difference and influence future designs. The creator of Zig certainly made an impression on me with this talk:
https://youtube.com/watch?t=120&v=Gv2I7qTux7g
Even if Zig hadn't been successful, the ideas it represents would've enriched the world. Gives me hope I'll be able to make something out of my own ideas too one day.
C/C++ tend to be CPU bound languages that are tightly coupled to the Von Neumann architectures. Anything that is not directly isomorphic tends to not survive very long unless its for a VM, and supports wrapper libraries.
Best of luck, =)
Avoid gradual typing: Absolutely. I had gradual typing and abandoned it. Gradual typing is not as good as simple Go-like inference, and that simple inference is easy to implement and takes care of 95% of the "problem."
However, you also need one or more dynamic typing escape hatches. When I implemented a config file format by tweaking JSON, I had to implement dynamic typing in C.
Avoid self-hosting your compiler: It depends. You probably should avoid it by default, unless your language is just so much better than the original. Rust is an example of this (compared to C++). I'm writing in C, and I want memory safety [1], so bootstrapping makes sense.
Avoid writing your own code generator, linker, etc.: This is excellent advice! Use the standard tools. Of course, me being me, I'm breaking it. :) I am trying a different distribution method, where you distribute an LLVM-like IR, and the end machine then generates the machine code on first run. In that case, there's no linker necessary, but I do have to write my own code generator. Fun.
Avoid bike shedding about syntax: Yes, absolutely. I did this, but the syntax still changed enormously!
Cross-platform support is a challenge: Yes, in more ways than one. First, you have to somehow generate code for all of the platforms, then you have to make sure your library works on all of the platforms too.
Compiler books aren't worth the money: Yes, but please do read Crafting Interpreters. Anyway, getting a simple parser and bytecode generator (or LLVM codegen) is the simple part of making a language. Then you need to make it robust, and no one talks about that. Maybe I should write a blogpost about that once my language stabilizes.
Growing a language is hard: Yes, absolutely. He mentioned two ways it needs to grow: libraries and users.
You can design a language to be easy to grow via libraries. See "Growing a Language" by Guy Steele. [2] I went the extra mile with this, and user code can add its own keywords and lexing code. So growing my language is "easy."
But growing the userbase? That's hard. You need to have a plan, and the best plan is to solve a massive pain point, or multiple. I'm targeting multiple.
First, I'm targeting shell; my language can shell out as easily as shells, or even more easily, but it's a proper language with strong static type checking. If there are people who want that instead of bash, they'll get it. And there's a lot of bash people might want to replace.
Second, I'm targeting build systems. My language is so easy to grow, I've implemented a build system DSL, and then I put a build system on top.
Shell and build systems are both things people hate but use a lot. These are good targets.
The best test suite is a real application: Yes, absolutely. Except, the best test suite is actually a bunch of real applications.
My language's own build scripts are the first real program written in it. I'm also going to replace every shell script on my machine with my language, and most of them will make it into the formal test suite.
Don't prioritize performance over functionality: Yes, absolutely. This is why I would suggest making an interpreter first; you don't depend on LLVM (shudder), and you can easily add functionality.
Building a language takes time: I've taken 11 years, and there's no release yet. Yes, this is true.
cries in the corner
Anyway, I wish the author luck, even though I'll be a competitor. :)
[1]: https://gavinhoward.com/2023/02/why-i-use-c-when-i-believe-i...
Well, that's like, your opinion, man...