An Overview of the Starlark Language
laurent.le-brun.eu
laurent.le-brun.eu
One challenge with Starlark was not having exceptions, which by itself is a good thing. But not having go-style multi-value returns in Starlark makes error handling verbose. Since the errors originate in plugin calls to go code, the solution I implemented was a thread-local to store the last error and handle that on the next plugin call if the error was not explicitly handled in the code [3]. Worked out pretty well in my test apps. For example, this app [4] implements a bookmark manager, with minimal error handling required in the code.
Starlark is great for adding dynamic behavior for go applications. I was worried about performance when I started but perf has not been a concern in the context of a web app server.
[1] https://github.com/google/starlark-go
[2] https://github.com/claceio/clace
[3] https://clace.io/docs/plugins/overview/#automatic-error-hand...
[4] https://github.com/claceio/apps/blob/main/utils/bookmarks/ap...
My biggest frustration when using Bazel is not even Bazel—it’s the fact that when I’m looking at code like this[1], everything comes from a single, unannotated `rctx` variable that has no type annotation. So when I’m trying to read the code, it’s a matter of constantly grepping the repo, grepping the bazel docs, and grepping the bazel source code to get any new code written.
Why can’t I just hover and see the docs? Why can’t I press `rctx.` and see all the attributes available?
I frequently hear “why do you need types? the language is Turing complete just run it and see if it fails” but that misses the point that I have want to use types to guide what code to write in the first place.
[1] https://github.com/bazel-contrib/toolchains_llvm/blob/master...
Inferred types, understandable type errors, expressive language: choose two out of three. If you sacrifice good type inference and require users to specify type annotations themselves they will not like it especially considering users' mental model of this language comes from Python; if you sacrifice understandable type errors, users will reject the whole system as soon as they encounter a type error; if you sacrifice expressiveness you are taking away features users want like subtyping.
If you want something that's ignored by Bazel and only used by some other tools you already have it: docstrings. You can enforce docstrings to be of a certain format and contain type annotations. Be the solution you want to see: be that hobbyist that writes a type checker and language server.
How do you type the omega combinator in Python?
lambda x: x(x)
What about the application of the omega combinator to itself? (lambda x: x(x))(lambda x: x(x))
How do you type the Y combinator in Python? lambda f: (lambda x: f(x(x)))(lambda x: f(x(x)))
Depending on the choice of the type system you can either answer: this cannot be assigned a type in which case your type system is too weak to express many real-world programs, or it does have a type but then your type system is so sophisticated that people using this language (just for writing BUILD rules) won't be able to understand type errors in this type system. I have yet to find a middle ground.In case you claim this is impractical functional programming, bear in mind you can write the same thing using classes in object-oriented programming.
---
Okay let's not even talk about these weird-looking "combinators" even though they have a rich history. Consider this function:
lambda x: {'this': x, 'next': x + 1}
What is its type? Again the answer depends on plenty of choices that need to be made by the type system designer. Do you force dictionaries to have a single type for values? If so, many users used to Python will reject your system for being too inflexible. If not, will you now introduce row polymorphism in your type system? Will you now introduce depth subtyping and width subtyping in your type system? (For example if a function only needs field 'x' in a dictionary but the caller passes a dictionary with both fields 'x' and 'y' should not result in an error; that's why you need subtyping.) Now let's consider the lowly plus operator. In real Python the plus operator can work on integers, floats, strings, sets, etc. Let's say your type inference algorithm uses the RHS to find that x must be an integer. But that's wrong; it could still be a float. Can you now write a type for that expression? Hint: it involves record types, intersection types, type variables and other things that would not be comprehensible.Also, from a brief investigation, according to [1] Haskell also isn’t able to express the omega combinator in its type system (& indeed the answer says very few type systems are able to do so). This suggests this is a very bad litmus test if almost no languages can actually pass it even though a) languages exist with explainable type errors b) the languages are general purpose and fairly expressive c) they support type inference. An obvious example of this is Rust.
[1] https://stackoverflow.com/questions/33546004/is-it-possible-...
Why not? Because you are adding typing retroactively to a dynamically typed language like Python. If a language was built with typing, you can do this because your initial choice of a type system already constrains the set of valid programs. But a Python-like language where users are already used to having no restrictions on runtime types? Absolutely not. Especially not in a language that encourages duck typing. Let's go back to the omega combinator example. The Stack Overflow link is correct: Haskell cannot assign a type to it. That means the type system is purposefully designed weak enough that the omega combinator cannot be written in the first place! No valid Haskell program has it! Therefore you can totally disregard it. Python is different. You can already write the omega combinator in it. I wrote it three times in this thread. So a bolt-on type system needs to assign a type to it. And you can't.
Python does not have a type system that "works well" and will never have one.
You are correct that the holy grail you’re trying to reach is impossible, but relaxing some constraints still yields a heck of a lot of practical benefit.
When weird type theory people insist you want global inference and that you should sacrifice something you actually value to get it, treat them the same as if Hannibal Lecter was explaining why you should let him kill and eat your best friend. Maybe Hannibal sincerely believes this is a good idea, that doesn't make it a good idea, it just further underscores that Hannibal is crazy.
Well, now that I think more about it, except when you are Google or Meta and you have a giant monorepo full of code that users of this language has actually written. So if Starlark were to gain a type system, the type system designer would probably just run it through the entire Starlark code at Google.
Speaking of Meta, this suddenly reminds me of Flow https://flow.org/ an effort at Meta to add a bolt-on type system to JavaScript. Naturally it doesn't aim to "work well" 100% of the time but maybe working 80% of the time is already valuable enough for them to use and announce it.
Except we have lots of easily accessible OSS code available via GitHub/GitLab etc in addition to self-reporting from closed repos. Can you find any instance of an omega combinator in use? I'm not even sure omega combinator is relevant to Starlark given that Starlark is a more limited language derived from Python. For example, it doesn't typically allow for recursion which I believe would be required to express an Omega combinator within it. Yes I know the definition of the combinator says it's not recursive, but that's in a pure CS sense. In a practical sense, there's not much difference between a function being passed itself as an argument and then invoking it vs invoking itself directly, particularly from an enforcement perspective. Indeed, you can verify it'll fail the recursion prevention code in Starlark [1]
[1] https://github.com/bazelbuild/bazel/blob/e1a73e6e89a082e3d80...
In Python, all your examples can be typed as "Any" https://docs.python.org/3/library/typing.html#the-any-type
It's an extremely simple type that never leads to hard-to-understand type errors, because it never produces type errors. And you can use it for any real-world program, including malformed ones.
Of course that means the type system is unsound and accepts programs that will fail at runtime. But that is perfectly fine as long as the type checker can spot some issues and gets out of your way otherwise.
Python typing is so basic that it didn't even support a proper JSON type until a year ago (!), and yet it brings value to millions of Python programmers around the world.
A bit expensive, perhaps. And it breaks down if the code is in an intermediate state. Like, "I just deleted a comma, where did my type annotations go?"
Unless that is true, the end result might not be ideal. You could certainly record the information that is there at every line of code at runtime, and you could probably calculate the union of the parameter types to find only the fields that are always there, which would be fairly okay I guess.
One difference is that Python specifies the type system in PEPs (basically Python's RFCs), while OP didn't ask for that here. This would mean whatever tool is implemented would define the type system.
My idea was to do type checking in a separate tool (e.g. built on top of Buildifier) and let the interpreter ignore the types. So it could be completely optional. The type system could be gradual, like in Typescript.
I don't know when/if it will happen (I'm no longer working at Google, so it's harder to make large contributions like this).
EDIT: as a parting shot, dhall is another non-Turing-complete language in common usage, but its claim to fame is that it gets used in places that arguably shouldn't be doing any computation at all.
sudo give me my files.
> if it has race conditions
Does this mean starlark cannot generate a race condition? If so, how does threading work?
In short: Evaluating a file has no visible side-effect. The evaluation returns a set of frozen values. These values that can accessed by other threads become immutable.
Basic answer: there is no threading, at least not for the user.
The build system runs a thread per BUILD file (in bazel). Since the BUILD files are written in starlark and starlark has no threads -- there is no problem.
Bazel is massively parallel. It has to evaluate a potentially large number of BUILD files; each file can have multiple load statements. You end up with a graph of dependencies and Bazel will evaluate as many files as possible in parallel.
So are there languages that do that ?
Parallelism: Easier with sandboxing, but the build declares all inputs/outputs so it builds a merkle tree and strictly inforces a DAG. This way you can parallelize independent subgraphs and cache at a fine grained level.
At the end of the file evaluation, Starlark exports the values and functions defined in the file. All values are frozen, meaning they won't ever be mutated.
Since no shared object can be modified, you can have any number of threads reading the object. So you can safely evaluate many files in parallel.
Caveat: this assumes the interpreter doesn't expose functions like reading/writing on a disk, etc.
ye gods
I worked on the original Buck and kind of caught the “principled and performant polyglot build system” bug and I’ve been diving down that rabbit hole since.
There are still usability and adoption challenges, but it’s the future and between buck2 and bazel7/bzlmod the gap to slow, non-deterministic alternatives is closing fast.
I wouldn’t write it off.
What I hate about it is that not every bazel ruleset I use has yet been ported to it. But it's getting there.
For example in my monorepo I'm only missing rule_distroless and I migrated all my deps to bzlmod.
The modules lock file situation has been greatly improved since 7.2.0-rc1. The lock file contains much less digests in it and its less likely to cause spurious merge conflicts
With bazelisk and .bazelversion we managed to have everybody use the same version of bazel and thus we could cheaply try out the rc1 and so far is working fine.
The Bazel team flags things off by default for years, and then things on for years, because they’re engineers and try hard not to break people’s shit.
Are feature flags and migration paths and interim solutions more work? Yeah, no doubt.
But for all their recent fumbles in AI and steadily downward trajectory on search and all the things that happens when a big company is run by people who were never all that technical?
Google still has a hard core of the baddest hackers you don’t want to mess with.