HNHacker News
TopNewBestAskShowJobs

typesanitizer

103 karma · joined April 30, 2022

https://typesanitizer.com
submissionscomments
typesanitizer··on Object storage is all you need
> [..] so isolation isn't a WHERE clause somebody has to remember, [..] Two things the post doesn't cover and I'm happy to get into. What closing the cross-region lost update actually took: [..]

Set of my Claude alarm bell, and lo-and-behold, Pangram judges this comment to be 100% AI-generated, albeit with limited confidence.

typesanitizer··on After 7 years in production, Scarf has reluctantly moved away from Haskell
Go compiles things at package-level granularity. You only need to recompile your reverse dependencies on making changes. Also there's build caching available out-of-the-box, as well as some support for test caching.
typesanitizer··on An epic treatise on error models for systems programming languages
I have removed the mention of "Let it crash" from that section, and added a clarification for my original intent. I did not mean it as criticism of Erlang or Joe Armstrong, although I 100% understand how it could've been interpreted as such.

Thanks for the critical feedback.

typesanitizer··on An epic treatise on error models for systems programming languages
Hi Gavin, I've seen your blog before, including some posts about Yao.

After feedback on Lobste.rs, I plan on adding a section on conditions and restarts, hopefully sometime later today if time permits. :)

I'd be happy to add some information about Yao as well alongside Common Lisp. It would be helpful for me to have some more details about Yao before writing about it, so I have some follow-up questions below. Please feel free to link me to existing writing; I may be misremembering details, but I don't think the answers to these have been covered somewhere.

I looked at the docs on the master branch as of mid Jan 2025 (could you confirm if these are up-to-date), particularly design.md, and I noticed these points:

> That means that Yao will have unsafe code, like Rust's unsafe. However, unlike Rust, Yao's way of doing unsafe code will be harder to use,

So Yao has a delineation between safe and unsafe code, correct? Does "safe" in Yao have the same (or stronger) set of guarantees as Rust (i.e. no memory safety if all necessary invariants are upheld by unsafe code, and boundary of safe-unsafe code)?

> Yao's memory management will be like C++'s RAII [..]

Does Yao guarantee memory safety in the absence of unsafe blocks, and does the condition+restart system in Yao fall under the safe subset of the language? If so, I'm curious how lifetimes/ownership/regions are represented at runtime (if they are), and how they interact with restarts. Specifically:

1. Are the types of restart functions passed down to functions that are called? 2. If the type information is not passed, and conditions+restarts are part of the safe subset of Yao, then how is resumption logic checked for type safety and lifetime safety? Can resumptions cause run-time type errors and/or memory unsafety, e.g. by escaping a value beyond its intended lifetime?

---

For reading the docs, I used this archive.org link (I saw you're aware of the Gitea instance being down in another comment): https://web.archive.org/web/20250114231213/https://git.yzena...

typesanitizer··on An epic treatise on error models for systems programming languages
Hi, blog post author here (unrelated to paper authors).

> So in a laps of 15 years (2010-1025),

The paper was published in 2014, so the period is 2010-2014, not 2010-2025.

> they hand picked 20 bugs from 5 open source filesystem projects (198 total)

The bugs were randomly chosen; "hand picked" would imply that the authors investigated the contents of the bug reports deeply before deciding whether to include them (which would certainly fall under "bad science"). The paper states the following:

> We studied 198 randomly sampled, real world fail- ures reported on five popular distributed data-analytic and storage systems, including HDFS, a distributed file system [27]; Hadoop MapReduce, a distributed data- analytic framework [28]; HBase and Cassandra, two NoSQL distributed databases [2, 3]; and Redis, an in- memory key-value store supporting master/slave replica- tion [54]

So only 1 out of 5 projects is a file system.

> and extrapolated this result. That is not science.

The authors also provide a measure of statistical confidence in the 'Limitations section'.

> (3) Size of our sample set. Modern statistics suggests that a random sample set of size 30 or more is large enough to represent the entire population [57]. More rigorously, under standard assumptions, the Central Limit Theorem predicts a 6.9% margin of error at the 95% confidence level for our 198 random samples. Obviously, one can study more samples to further reduce the margin of error

Do you believe that this is insufficient or that the reasoning in this section is wrong?

typesanitizer··on An epic treatise on error models for systems programming languages
Hi, author here. I'm a big fan of Armstrong's work, I've watched several of his talks multiple times and always get something new out of them even if I don't agree entirely. :)

I do mention Erlang near the start of the post, around the 360 word mark:

> Joe Armstrong’s talk The Do’s and Don’ts of Error Handling: Armstrong covers the key requirements for handling and recovering from errors in distributed systems, based on his PhD thesis from 2003 (PDF) [sidenote 3].

> [sidenote 3] By this point in time, Armstrong was about 52 years old, and had 10+ years of experience working on Erlang at Ericsson.

> Out of the above, Armstrong’s thesis is probably the most holistic, but it’s grounding in Erlang means that it also does not take into account one of the most widespread forms of static analysis we have today – type systems

typesanitizer··on An epic treatise on error models for systems programming languages
Copying my comment from the Lobste.rs thread (https://lobste.rs/s/az2qlz/epic_treatise_on_error_models_for...)

> Hi, author here, the title also does say “for systems programming languages” :)

> For continuations to work in a systems programming language, you can probably only allow one-shot delimited continuations. It’s unclear to me as to how one-shot continuations can be integrated into a systems language where you want to ensure careful control over lifetimes. Perhaps you (or someone else here) knows of some research integrating ownership/borrowing with continuations/algebraic effects that I’m unfamiliar with?

> The closest exception to this that I know of is Haskell, which has support for both linear types and a primitive for continuations. However, I haven’t seen anyone integrate the two, and I’ve definitely seen some soundness-related issues in various effect systems libraries in Haskell (which doesn’t inspire confidence), but it’s also possible I missed some developments there as I haven’t written much Haskell in a while.

typesanitizer··on The Evolution of SRE at Google
Thanks for writing the summary notes and sharing those here. After reading the Usenix article, I was thinking that we could apply some of the ideas at $WORK, but the exact "How" was still not super clear. Your notes offer a compact and accessible starting point without having to ask colleagues to dive in to a 100+ page PDF. :D
typesanitizer··on Optimizers need a rethink
> If you have any language where it is "semantially correct" to execute it with a simple interpetter, than all optimizations in that language are not semantically important by definition, right?

Technically, yes. :)

But I think this should perhaps be treated as a bug in how we define/design languages, rather than as an immutable truth.

- We already have time-based versioning for languages. - We also have "tiers" of support for different platforms in language implementations (e.g. rarer architectures might have Tier 2 or Tier 3 support where the debugging tooling might not quite work)

One idea would be to introduce "tiers" into a language's definition. A smaller implementation could implement the language at Tier 1 (perhaps this would even be within reach for a university course project). An industrial-strength implementation could implement the language at Tier 3.

(Yes, this would also introduce more complications, such as making sure that the dynamic semantics at different tiers are equivalent. At that point, it becomes a matter of tradeoffs -- does introducing tiers help reduce complexity overall?)

typesanitizer··on Optimizers need a rethink
> In my view, if a compiler optimization is so critical that users rely on it reliably “hitting” then what you really want is for that optimization to be something guaranteed by the language using syntax or types. The way tail calls work in functional languages comes to mind. Also, the way value types work in C#, Rust, C++, etc - you’re guaranteed that passing them around won’t call into the allocator. Basically, relying on the compiler to deliver an optimization whose speedup from hitting is enormous (like order of magnitude, as in the escape analysis to remove GC allocations case) and whose probability of hitting is not 100% is sort of a language design bug. > > This is sort of what the article is saying, I guess.

I agree with this to some extent but not fully. I think there are shades of grey to this -- adding language features is a fairly complex and time-consuming process, especially for mainstream languages. Even for properties which many people would like to have, such as "no GC", there are complex tradeoffs (e.g. https://em-tg.github.io/csborrow/)

My position is that language users need to empowered in different ways depending on the requirements. If you look at the Haskell example involving inspection testing/fusion, there are certain guarantees around some type conversions (A -> B, B -> A) being eliminated -- these are somewhat specific to the library at hand. Trying to formalize each and every performance-sensitive library's needs using language features is likely not practical.

Rather, I think it makes sense instead focus on a more bottoms-up approach, where you give somewhat general tools to the language users (doesn't need to expose a full IR), and see what common patterns emerge before deciding whether to "bless" some of them as first-class language features.

typesanitizer··on Optimizers need a rethink
I'm guessing you've tried these flags mentioned in the blog post but haven't had luck with them?

> LLVM supports an interesting feature called Optimization Remarks – these remarks track whether an optimization was performed or missed. Clang support recording remarks using -fsave-optimization-record and Rustc supports -Zremark-dir=<blah>. There are also some tools (opt-viewer.py, optview2) to help view and understand the output.

typesanitizer··on Optimizers need a rethink
I've added a clarification in the post to make my position explicit:

> This is not to imply that we should get rid of SQL or get rid of query planning entirely. Rather, more explicit planning would be an additional tool in database user’s toolbelt.

I'm not sure if there was some specific part of the blog post that made you think I'm against automatic query planning altogether; if there was, please share that so that I can tweak the wording to remove that implication.

typesanitizer··on Optimizers need a rethink
Thanks for the feedback.

The preceding paragraph had "and occasionally language features" so I thought it would be understood that I didn't mean it as an optimizer-specific thing, but on re-reading the post, I totally see how the other wording "The knobs to steer the optimizer are limited. Usually, these [...]" implies the wrong thing.

I've changed the wording to be clearer and put the D example into a different bucket.

> In some cases, languages have features which enforce performance-related properties at the semantic checking layer, hence, granting more control that integrates with semantic checks instead of relying on the optimizer: > > - D has first-class support for marking functions as “no GC”.

typesanitizer··on The case of a leaky goroutine
Structured concurrency is part of the Swift standard library, and was added at the same time when first-class support for concurrency was added.

TaskGroup in the standard library - https://developer.apple.com/documentation/swift/taskgroup

Explore structured concurrency in Swift (WWDC 21) - https://developer.apple.com/videos/play/wwdc2021/10134/

typesanitizer··on Sourcegraph is no longer open source
Based on your use of "campaign" (the older name for Batch Changes), it sounds like you were looking into Sourcegraph about 2.5 years ago or before that. Lots has changed since then.

We recently released a new indexer scip-clang (https://about.sourcegraph.com/blog/announcing-scip-clang), which we've used to successfully index large codebases like Chromium. The indexer relies on a JSON compilation database (same as our older indexer lsif-clang) which is easy to produce from CMake, Bazel, Meson, Make etc.

We've also added support for cross-repo code navigation for C++ recently. (https://about.sourcegraph.com/blog/c-cpp-cross-repo)

typesanitizer··on Overview of C++ language support in Apple Clang
I don't know the answer to your original question, but based on this StackOverflow answer (https://stackoverflow.com/a/60564952/2682729), it seems like a small compiler wrapper should do the trick to get the same experience across platforms without having to install another compiler.

Depending on the platforms you need to support, a more heavyweight but more reproducible solution would be to use a fixed LLVM toolchain, say using Bazel.

typesanitizer··on The technology behind GitHub’s new code search
All of our SCIP indexers are open-source: scip-java (for Java, Kotlin and Scala), scip-typescript (for TypeScript and JavaScript), scip-python, scip-ruby, scip-go and scip-clang (for C and C++).

There are also some community-maintained OSS SCIP indexers. - rust-analyzer: https://sourcegraph.com/github.com/rust-lang/rust-analyzer/-... - scip-zig: https://github.com/zigtools/scip-zig

typesanitizer··on The technology behind GitHub’s new code search
(I work on C++ indexing at Sourcegraph.)

As my colleague mentioned in a sibling comment, we have an existing indexer lsif-clang which supports C++. I just added a Chromium example to the lsif-clang README right now: (direct link) https://sourcegraph.com/github.com/chromium/chromium@cab0660...

We are also actively working on a new SCIP indexer which should support features like cross-repo references in the future. https://github.com/sourcegraph/scip-clang

Right now, Abseil doesn't have precise code navigation because no one has uploaded an index for it. In an ideal world, we would automatically have precise indexes for all the C++ code on Sourcegraph, but that's a hard problem because of the large variety in build systems, build configurations, and system dependencies that are often specified outside the build system.

typesanitizer··on Zig-style generics are not well-suited for most languages
Fixed, thanks.
typesanitizer··on Steve Yegge Joins as Head of Engineering of Sourcegraph
For C++, we do support Bazel via compile_commands.json; we have customers who have used it successfully. Depending on the editor you're using, you probably need to get Bazel to generate a compile_commands.json anyways.

So core code navigation functionality ought to work. There are some issues with different language features (macros, newer C++17 and C++20 features), as well as robustness (crashes on indexing certain code) but we're looking into bringing C++ support up to par with other languages.

typesanitizer··on Steve Yegge Joins as Head of Engineering of Sourcegraph
Created a PR to mention tools using SCIP in the README. https://github.com/sourcegraph/scip/pull/101
typesanitizer··on Ferrocene: Rust toolchain to safety-critical environments
I don't think this is entirely accurate. For example, the Translation Validation section on this Wikipedia page mentions (https://en.wikipedia.org/wiki/Compiler_correctness)

> Translation validation can be used even with a compiler that sometimes generates incorrect code, as long as this incorrect does not manifest itself for a given program. Depending on the input program the translation validation can fail (because the generated code is wrong or the translation validation technique is too weak to show correctness). However, if translation validation succeeds, then the compiled program is guaranteed to be correct for all inputs.

I don't remember where I read this, but I think there are some examples in practice where instead of proving the correctness of some optimizations of existing C compilers (which is a codebase that keeps evolving), there have been situations where the correctness was ensured using program equivalence checking. So you'd "freeze" the reference compiler and the equivalence checker (which would be qualified), and pair that with an upstream compiler with sophisticated optimizations. So long as the equivalence checker keeps passing, you're golden.

typesanitizer··on Google is 2B lines of code and it's all in one place (2015)
Hey! I work on code intelligence at Sourcegraph, so I figured I'd chime in here.

> why don't you have decent crossrefs

We do have compiler-accurate cross references for many repos. Some examples:

- TypeScript: https://sourcegraph.com/github.com/sindresorhus/got/-/blob/s...

- C: https://sourcegraph.com/github.com/neovim/neovim/-/blob/src/...

- C++: https://sourcegraph.com/github.com/KhronosGroup/Vulkan-Sampl...

Why are not all repos covered?

Because different languages have different build systems, so inferring the right build commands, dependencies etc. is not so straightforward; these are necessary pre-requisites for compiler-accurate cross references. We're working on fixing this with auto-indexing: https://docs.sourcegraph.com/code_intelligence/explanations/...

For C and C++ specifically, auto-indexing is challenging because of the large variety in build systems, informal specification of dependencies (such as in a README instead of a machine-readable format), and platform-specific code.

Outside of auto-indexing, we do have an indexer for C and C++ right now (https://github.com/sourcegraph/lsif-clang) which can be run in CI; that way one can generate an index and upload it to Sourcegraph on a regular basis. It is 'Partially available' (https://docs.sourcegraph.com/code_intelligence/references/in...) right now. We're keenly aware of the interest in C++, and are working our way through different languages based on usage.

typesanitizer··on Experience Report: 6 months of Go
I sympathize with your point that being able to use simpler tools (such as grep) is often a good thing instead of having to rely on heavy-duty functionality (such as a full IDE/language server). That said:

- Ctrl+F in a file specifically isn't guaranteed to always return a hit with Go since a package can span multiple files, and you can have methods in other files.

- Increasingly many tools (like Sourcegraph and GitHub) make rich code navigation available on the web, without having to use an IDE or check out the code locally. Yes, it's not perfect for all languages, but it's improving every day. Similarly for documentation tools in different ecosystems -- I think both Haddock (Haskell) and rustdoc support cross-linking references to definitions.

- If you have textual code search for dependencies, you can still grep through the code with a regex like `func \(.* *?MyType\) myMethod`, and it will give you hits in other packages too.

typesanitizer··on Experience Report: 6 months of Go
Ah, making it non-exported was an unintentional mistake. I've pushed a fix marking it as exported with an EDIT note.
typesanitizer··on Experience Report: 6 months of Go
> There's no perfect language, yet people are always trying to find one. It's OK if you don't like Go and prefer another language, the real devil / tradeoff is in the fact that conformance to a single language (or set of languages) is such a strong social phenomena. I think that's why people end up so angry with viewed-as-subpar languages like Go gaining so much traction

I hope my post didn't come across this way. To be clear, I think Rust and Swift (as two examples that I mention multiple times) have a lot of problems (slow compilation being a very big one), and are by no means perfect. I'm not angry at Go's popularity as much as wanting improvements to the developer experience.

typesanitizer··on Experience Report: 6 months of Go
> Some complaints also seem like a lack of experience using the language.

> In my experience, the two-value (explicit capacity) form of "make" is significantly _less_ common than the single-value form. Indeed, gripping through the stdlib shows "make([]T, n) is much more common than "make([]T, n, m)".

I've written a fair bit of C++, where this pattern is very common. IME in 95%+ of the cases, what one wants is a vector with a capacity without initializing it, because it will be filled up right away.

I'd argue that make([]T, n) is more common in actual Go code precisely because it has the shorter spelling, not because it has the exact desired semantics.

typesanitizer··on Experience Report: 6 months of Go
> I would be curious to know what he thought of Swift (having worked on the compiler). The language seems to be the opposite of Go in including every language feature under the sun.

Are you asking from a language design perspective? Or from an implementation perspective?

From a language design perspective, yes, Swift has a lot of features. Most of these features exist for good reasons.

- First-class Objective-C interop: Needed for initial adoption and migration, since Apple's SDKs were all Objective-C. (See also: Kotlin and Java etc.)

- Protocols with associated types: Writing generic code with constraints makes surfacing type errors easier compared to templates. Associated types enable many natural patterns of programming.

- Library evolution: Being able to evolve APIs without breaking ABI is super important for a platform.

- Support for DSLs: Swift is used heavily for UI programming, and there is a convergence across languages in terms of having DSLs for making building UIs easier.

- Use of weak pointers (vs having a tracing GC for cycles): Better for Objective-C compatibility.

- Async/await + actors: Trying to balance usage of Dispatch (which is the platform API) with newer programming patterns, while still being able to compile in a way with low resource usage.

- Upcoming C++ interop: Many big iOS applications use large amounts of C++, so better interop would make Swift usage easier for them.

Does that mean I think every feature of Swift is perfect? No. For one thing, I think method overloading is way too flexible, which is what causes exponential time for type inference in a bunch of cases.

typesanitizer··on Experience Report: 6 months of Go
I didn't mention it because I didn't run into date handling code myself over the past 6 months. I've focused on the positives and negatives that I've run into in practice.
typesanitizer··on Experience Report: 6 months of Go
If you look at the GopherCon talk I gave (linked at the beginning of the post), it is about reading the spec. So yes, I did read the spec, and I realize that this is not syntactically valid (it would be a very basic compiler bug if this were valid syntax and it was diagnosed as a syntax error).

However, the spec only states what _is_ but not _why_ it is that way. Sure, I could look at the git blame of the spec for every odd thing I run into, but there is only so much time in the day...

Page 1 of 2Next →