JavaScript macros in Bun
bun.sh
bun.sh
IMO these aren’t macros in the Lisp-sense of the word (or Rust, or even C); yeah they run code at compile time, but that’s where the common ends.
Macros should be able to apply syntactic transformation on the code. Lisp is famous for allowing that by representing code as lists. Rust has a compiler-level API to give tokens and run arbitrary code, then spit new tokens out. C macros operate on the tokens level, so with enough magic you can transform code to the shape you want.
This… isn’t any of that.
A pretty good example (and something I’m still sad that it didn’t take off) of macros in JS is Sweet.js[0]. Babel macros[1] are a bit higher level, where macros require the input to already be a valid AST, but that’s also something I’d call macros.
This… I’d say it’s more of a compile time code execution feature, not a macro feature.
[0]: https://www.sweetjs.org/ [1]: https://babeljs.io/blog/2017/09/11/zero-config-with-babel-ma...
https://gist.github.com/Jarred-Sumner/454da846d614f7bb4bcceb...
Ideas are cheap, and naming is hard, but I like Zig’s “comptime” term, which is closer to this than “macro”. “comptime” distinguishes it from C preprocessor stuff, but also describes what it actually does more clearly.
> Another failure I see is the fact that the "type" property is not meant to be used for this type of language augmentation, it's meant to describe the type of the file without inferring from the file extension
It sounds like you’re confusing Import Attributes with Import Assertions, which was the previous iteration of Import Attributes.
Interpretation of import attributes is host defined. For Import Assertions, that wasn’t true - they were intended never to impact runtime evaluation. That’s the difference with Import Attributes. Import Attributes do impact runtime evaluation.
This turns what looks like a function call in the importer into something like macro expansion (it doesn't look like actual macros though).
The only inconsistency is that Bun is front loading some of these effects to the server/bundler runtime and effectively memoizing the equivalent behavior before it reaches a client. But it’s not doing that of its own volition, it’s doing it to address an explicit attribute in the source code.
The only way this would be a meaningful problem is if the explicit value has a chance of colliding with either existing code in the ecosystem (highly unlikely, no one is really using this syntax yet except perhaps in an equally experimental context), or some plausible pending standard (unless one was proposed in the last couple weeks, I’m pretty sure I can rule that out too).
I share other commenters’ lament that this is not a true macro solution (and I think it should actually just be renamed to something like comptime). But I don’t think this deserves the deviation-from-standards challenge it’s getting in this thread. And lest I come off as a Bun fanatic, I think I’m one of the people who more frequently questions potential Bun spec deviations when they come up on Twitter.
https://bun.sh/blog/bun-macros#how-it-works
So the worry about having legacy anti-spec things still seems pretty valid.
Specification and implementation are chickend-and-egg.
B. If you blur the line between "runtime" and "compiler", you'll realize that a lot of JS ecosystem (e.g. Babel) run ahead of that.
---
The way we do this kind of thing at Notion is very simple. We have normal CLI commands that generate code and write it to disk, and we check in those outputs. Then in CI, we run all the generation commands and verify the codegen is up-to-date. Checking in generated code means it's really easy to understand the behavior of codegen and how it changes, and everyone gets excellent typechecking and auto-completion of codegen'd artifacts.
The downside to static codegen is that the "templates" we use for codegen need to be valid Typescript files, or we risk breakage if imports or types change.
Bun macros have some advantages like execution at build time for stuff like Git SHA interpolation, but because these macros can't actually emit code, they feel less powerful that straight-up codegen. I could replace a lot of the Notion CLI commands with Bun macros if:
1. There's a return type for macros that emits code, like:
// src/macros/prism.ts
export function emitPrismImports() {
const importOrder = toposortLangagues(PRISM_LANGUAGES)
const imports = importOrder
.map(lang => `await import(/* webpackChunkName: "prism" */ "${importPath}"`)
.join('\n')
return Bun.code`${imports}`
}
// src/client/syntaxHighlighting.ts
import { emitPrismImports } from '@notionhq/macros/prism' with { type: 'macro' }
async function highlight(text, lang) {
emitPrismImports()
return Prism.highlight(text, lang)
}
2. There's a way to run `bun build` and just do macro evaluation. I need to continue using Typescript, Webpack, Jest, etc for now, so we need eval'd macros on disk so they can be typechecked, tab-completed, and bundled by other tools.The uncompiled versions still wouldn't typecheck nicely though :thinking_face:
Any tips on how to do this? I’ve been running into similar problems with code generation in TS, and I haven’t been able to come up with a good technique for how to solve this problem yet. (Anything I’ve been able to find is either an AST that’s awkward to use, or string templating that can’t be valid code since it has extra syntax in it.)
The scripts that consume our templates are mostly very simple ~15 line affairs but there are a few very complex, 1000+ line affairs that use the Typescript type checker API to codegen model classes and test data generators from hand-written interface types. Those things read ASTs but writing ASTs seems like a huge waste of time to me. AST is only nice if you need to very carefully morph existing code, and even then I usually end up doing a string replacement of an AST node’s position in its original file. In all that code I think the only thing that writes using an AST is import statement stuff.
Luckily that super complex stuff runs on every commit so it can’t break. It’s the one-off codegen commands like “make me a new component” or “make me an integration test” that are in danger of going stale, and even those we could write some CI tests for like “make a new thingy, check that it typechecks”.
Anyways here’s a full example template file:
/*__FILE_METADATA__*/
/* ================================================================================
__TEMPLATE_CLASS_NAME__.
Docs: https://dev.notion.so/notion/Record-Framework-1f41f97e1cc746628a6af57ba75d75ad
THIS IS FILE IS PARTIALLY GENERATED. ONLY EDIT THE EDITABLE REGION BELOW.
__GENERATED_BY__
__GENERATED_FROM__
================================================================================ */
import type {
/*__VALUE_IMPORT__*/ BlockValue as __VALUE_TYPE__,
/*__TABLE_IMPORT__*/ BlockTable as __TABLE__,
} from /*__SCHEMA_IMPORT__*/ "../schemas/Block"
import { Model } from "./Model"
/*__REMOVE_LINE__*/ // : Value <Value> as __MODEL__<Value> | undefined <- Conform to template args in RecordStore.
/**
* This class is generated from {@link __VALUE_TYPE__}.
* To customize, edit the section in {@link __TEMPLATE_CLASS_NAME__} below.
*/
abstract class Generated__TEMPLATE_CLASS_NAME__<
Value extends __VALUE_TYPE__ = __VALUE_TYPE__
> extends Model<Value, __TABLE__> {
/*__REMOVE_LINE__*/
__BODY__: undefined // Put the goods here.
}
/**
* Read and interpret the data of a __RECORD__ record.
*/
export class __TEMPLATE_CLASS_NAME__<
Value extends __VALUE_TYPE__ = __VALUE_TYPE__
> extends Generated__TEMPLATE_CLASS_NAME__<Value> {
/*__USER_EDITABLE_SECTION__*/
}What are the commented bits like /* __REMOVE_LINE__ / and / __SCHEMA_IMPORT__ */ used for?
__REMOVE_LINE__ is a common hack with our template system to remove lines in the template that are needed to make the template valid, but aren't wanted in the resulting code. We replace /^.__REMOVE_LINE__.$/gm with "". In this case we need to accept the __MODEL__ template var because a different template called with the same arguments needs it, but we don't care in this one.
__SCHEMA_IMPORT__ is another marker we use to replace a whole line, here we replace it with `} from "${generateData.schemaImportPath}"`
> The result of the macro must be serializable!
> Functions and instances of most classes (except those mentioned above) are not serializable.
It's surprising to see effort spent coming up with and developing niche new features rather than on bridging the gaps (within reason) with node.
I've never worked on a language project like this before though so I'm not in a place to cast any judgement, just echoing the sentiment that this seems strange from the outside. Maybe this is just what it takes to keep the project interesting for Jarred? I can relate to getting bored with a project once I've proved out the difficult parts and most of what remains just feels like chores
Our focus is very much on Node.js compatibility
The "killer app" is the app that necessitates adoption of the platform.
E.g. Lotus 1-2-3 is the killer app of the IBM PC. Lots of users want Lotus 1-2-3, therefore they need to adopt the IBC PC platform.